logoalt Hacker News

vonneumannstantoday at 2:54 PM0 repliesview on HN

Short answer yes.

Slightly longer answer for older text to image models you teach them how to encode images and text into the same latent space. Then you simply do a conversion, take a text input, put it into latent space and then extract the image that latent space represents.