logoalt Hacker News

simonw • today at 6:44 PM • 13 replies • view on HN

Pelicans. Sonnet 5.5 has the same problem as Opus 5.5: on "max" thinking effort it burned through 128,000 thinking tokens (taking 15 minutes to do that) and ran out before it had produced the final SVG.

https://tools.simonwillison.net/markdown-svg-renderer?url=ht...

Here's how the thinking effort levels compare:

  low
  27 input, 1,623 output, thinking_tokens: 0
  1.6284
  Duration: 10138ms (10s)
  
  medium
  27 input, 1,796 output, thinking_tokens: 0
  1.7914 cents
  Duration: 11266ms (11s)

  high
  27 input, 2,334 output, thinking_tokens: 745
  2.3394 cents
  Duration: 17376ms (17s)

  xhigh
  27 input, 5,730 output, thinking_tokens: 2535
  5.7354 cents
  Duration: 41882ms (41s)

  max (failed to return response)
  27 input, 128,000 output, thinking_tokens: 128000
  $1.28
  Duration: 940617ms (15m 40s)
Low and medium both used 0 thinking tokens.

Replies

thefourthchime • today at 7:05 PM

It does vey well at one shotting a PacMan clone, pretty much perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-son...

2nd only to Opus 5.5, which is perfect. https://jonclegg.github.io/pacman-bakeoff/entries/claude-opu...

Up until very recently, all models struggled with this.

All results: https://jonclegg.github.io/pacman-bakeoff/

➕ show 10 replies
dennisy • today at 9:19 PM

Does anyone really still care about these pelicans?

Any model release it’s the top comment, I do not understand why.

➕ show 4 replies
mgaunard • today at 10:05 PM

what's most surprising is the difference between high and xhigh

croemer • today at 6:52 PM

This is evidence that Sonnet 5.5 wasn't yet trained on the HN comments from the Opus 5.5 release. Maybe Pelicanmaxing will lead to 127000 thinking tokens being used on Max.

➕ show 1 reply
amelius • today at 8:50 PM

This is great news because it means the model has not been benchmaxxed on stupid metrics.

PS: the next human that brings up pelicans on bicycles should try to draw them.

codingisfreedom • today at 9:32 PM

Sonnet 5 had the same problem with ‘max’. In a free sub, I would never get an answer back even for very simple prompts. It would just churn on nothing and return max token usage reached.

I’m not sure whether that’s a feature or a bug at this point though.

parkersweb • today at 7:50 PM

I like the one where the pelican is using the non-pedalling leg to control the handlebars because its wings won’t reach!

platinumrad • today at 7:17 PM

The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.

➕ show 2 replies
TomGarden • today at 6:56 PM

Where do you run sonnet/opus where you are limited to 128k, given they are both 1M context window models?

➕ show 3 replies
keeeba • today at 7:24 PM

Thank you for the pelicans sir, how do you think they compare to other models in Sonnet’s pricing/capability range?

pelicanmaxer • today at 8:00 PM

that pelican one-pedaling

aimaxxed • today at 7:08 PM

“Pelicans are solved.”

➕ show 1 reply
heyjstn • today at 7:03 PM

I think the next models will be benchmaxxing on the Pelican benchmark tbh

➕ show 1 reply