logoalt Hacker News

E-Reverancetoday at 3:49 AM1 replyview on HN

I strictly meant using the embeddings for training a reward model, not the generator. By good for generation I just meant the reward model might find more visual cues for aesthetic preference and avoid some of the spurious semantic correlation CLIP has


Replies

schopra909today at 4:21 AM

That might work! Off the dome, it’s not clear to me whether spatial/depth priors are better/worse than an LLM for this type of task.

Only reason I can think why the LLM might still work better here is that it’s trained to solve a bunch of different image/video related questions, so it’s perceptual modules may be more robust adaptive for this aesthetic grading task versus something like LingBot

show 1 reply