This seems like something that could be solved by asking an LLM to write the prompt for the image model. You can also feed in the output of an image model into an LLM and ask it to check it/make improvements.
There's definitely more dogfooding that needs to be done. And id argue that if your purpose is truly to learn or to teach, the process of describing that image will do wonders for retention.
This is the way.
There's definitely more dogfooding that needs to be done. And id argue that if your purpose is truly to learn or to teach, the process of describing that image will do wonders for retention.