Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.
https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
Here's how the thinking effort levels compare:
low
27 input, 1,623 output, thinking_tokens: 0
1.6284
Duration: 10138ms (10s)
medium
27 input, 1,796 output, thinking_tokens: 0
1.7914 cents
Duration: 11266ms (11s)
high
27 input, 2,334 output, thinking_tokens: 745
2.3394 cents
Duration: 17376ms (17s)
xhigh
27 input, 5,730 output, thinking_tokens: 2535
5.7354 cents
Duration: 41882ms (41s)
max (failed to return response)
27 input, 128,000 output, thinking_tokens: 128000
$1.28
Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.Does anyone really still care about these pelicans?
Any model release it’s the top comment, I do not understand why.
what's most surprising is the difference between high and xhigh
This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.
This is great news because it means the model has not been benchmaxxed on stupid metrics.
PS: the next human that brings up pelicans on bicycles should try to draw them.
Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.
I’m not sure whether that’s a feature or a bug at this point though.
I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!
The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?
Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?
that pelican one-pedaling
I think the next models will be benchmaxxing on the Pelican benchmark tbh
It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...
2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...
Up until very recently, all models struggled with this.
All results: https://jonclegg.github.io/pacman-bakeoff/