I experimented with something similar with LLMs. They can kinda do it for some stuff.
More interesting, you can ask a vision model to color parts of speech in img2img and it works OK for frontier models.