logoalt Hacker News

platinumrad • today at 7:17 PM • 2 replies • view on HN

The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.


Replies

usef- • today at 9:06 PM

Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.

I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.

That would make it about so, I assume?

             Score  Tokens  Reason  Cost 
 Kimi K3 Max    44  48k     32k     $2.00
 Half reason    44  32k ?   16k ?    ?

 Opus Med       51  26k     12k     $1.34
 Opus High      54  36k     18k     $1.82
 Opus Max       58  119k    84k     $5.98

 Sonnet Med     41  ?       ?       $0.59
 Sonnet High    47  ?       ?       $1.08
 Sonnet Max     56  193k    142k    $7.60
Medium is Anthropic's default.

Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.