All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.
All modern multimodal models can do vision sufficiently well for web design, with the exception of respecting negative space and seeing poor padding/margins on text. The issue is that prompts are often egregiously underspecified so the design -> vision loop doesn't know how to refine.