logoalt Hacker News

eli • today at 6:19 PM • 2 replies • view on HN

Why isn't that worth reading into? I care about the experience of actually using the model, not hypothetically what it could achieve without overactive guardrails


Replies

abejora • today at 6:25 PM

You're right about its real world performance, and I worded my original comment wrongly.

I was merely thinking of the theoretical aspect of it: performance of opus 5.5 is better than sonnet 5.5 across the board, with the exception of Terminal-Bench. So I was curious why this one stood out. Was it because they focused on it during training? Did sonnet 5.5 had access to more references for this benchmark? But based on my first reading, I concluded that it might just be the safety constraints that made the difference here, and I wanted to share that.

➕ show 3 replies
chis • today at 8:26 PM

Well presumably now it’ll fall back to Sonnet 5.5 lol