The contrast between Anthropic, who seem to be training their models to output ever-increasing numbers of reasoning tokens, and Fireworks's Ember-1, which was explicitly trained to preserve the quality of a model's responses while cutting down on reasoning, is interesting. Claude Code also uses more many tokens per task per model than any other harness in benchmarks.
Anthropic's "Max" modes seem like a yolo mode: "use 10x the tokens to try to break the hardest possible problems". But their models don't seem less efficient at normal reasoning modes.
I can't see Ember on AA's index yet, but their post claims "half the reasoning tokens for the same answers" as Kimi K3.
That would make it about so, I assume?
Medium is Anthropic's default.Having a less efficient mode isn't necessarily a mistake -- the purpose of configurable effort levels after all is to be able to put more thought into a problem.