logoalt Hacker News

d1l • today at 7:08 PM • 2 replies • view on HN

At work we use haiku 4.5 for a handful of latency sensitive tasks that are fairly simple. It performs well. Just started testing 5.5 as I’ve been anticipating a nice improvement since it was teased. Results so far are trash. Prompt leakage even. And it’s slower. I guess it’s cheap but I think they got the balance wrong on this.


Replies

HyperL0gi • today at 9:41 PM

Exact same thing here.

Both evals and Human pairwise tests for our use case are giving Haiku 4.5 first place in pretty much all tests.

No we'll try understand if we need to change our prompts to match performance ...

edit: maybe this will help: https://platform.claude.com/docs/en/build-with-claude/prompt...

➕ show 1 reply
saretup • today at 7:18 PM

Curious as to why. Haiku 4.5 has been far away from pareto frontier for a long time. Maybe you need to update your prompt for the newer model in your workflow.