logoalt Hacker News

gpm • today at 3:13 PM • 1 reply • view on HN

All the alignment issues seem likely to be solved as soon as LLMs can actually do vision well IMO.


Replies

minimaxir • today at 3:27 PM

All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.