That might work! Off the dome, it’s not clear to me whether spatial/depth priors are better/worse than an LLM for this type of task.
Only reason I can think why the LLM might still work better here is that it’s trained to solve a bunch of different image/video related questions, so it’s perceptual modules may be more robust adaptive for this aesthetic grading task versus something like LingBot
Apologies for making the reply chain so long but I think a video like this somewhat proves how a lot of aesthetic preferences can be *ultra* sensitive to small visual details : https://youtu.be/twcMra_67-w?t=88
The video is timestamped to open at the comparison frame. I don't think an LLM can tell the quality difference without direct reference for comparison