logoalt Hacker News

jjcm • today at 6:30 PM • 4 replies • view on HN

Ran image -> html tests for this. I was curious if this smaller model was good enough for complex UI. It was not.

Haiku 5.5: https://html.non.io/lcars-haiku-5.5/

Opus 5.5 for comparison: https://html.non.io/lcars-opus-5.5

Designs it was building from: https://diffui.ai/app/canvas/5093e689-1e74-4f26-b632-2a4500f...

One interesting thing is it took a look at the job at hand, and immediately delegated it to Opus 5.5. It at least knows what it isn't good at. Very fast though, and likely best used for small subagent tasks / tightly scoped work.


Replies

thefourthchime • today at 7:07 PM

Pac-Man Bench:

Considering the price, no model comes close to being as good as this. However, it did take an extremely long time.

TIME 19m COST $0.16 https://jonclegg.github.io/pacman-bakeoff/#claude-haiku-5-5

All results: https://jonclegg.github.io/pacman-bakeoff/

➕ show 3 replies
sparklingmango • today at 7:37 PM

> likely best used for small subagent tasks / tightly scoped work.

Hasn't this always been the case with Haiku?

saretup • today at 7:08 PM

To be fair, you're making it compete with the best public LLM right now that's 2 size/price tiers above it.

➕ show 1 reply
BrokenCogs • today at 6:34 PM

Neither of these look "good" to me. There is so much visual noise on the page, like someone turned the "AI Slop" dial to 11. In fact I prefer the simpler design Haiku made.

➕ show 3 replies