logoalt Hacker News

minimaxir • today at 6:05 PM • 18 replies • view on HN

Pricing is...a bit weird.

    Input
    $0.10 / MTok for prompts up to 100,000 tokens
    $0.50 / MTok for prompts over 100,000 tokens

    Output 
    $0.50 / MTok for prompts up to 100,000 tokens
    $2.50 / MTok for prompts over 100,000 tokens
100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents; for typical generation or Jev-like classifiers, it's a good value and as noted in this article, that is apparently the vast majority of Haiku use.

In both cases, still much cheaper than Haiku 4.5's $1 input / $5 output and these prices better compete with GPT-6 Luna. ($0.10 input / $0.50 output, but with no token threshold [EDIT: the threshold for Luna is apparently 272k])


Replies

jeremyjh • today at 10:09 PM

If you can't get any coding done with 100K context that is either a broken model, a broken harness or a skill issue. I would mostly use Haiku in task or explorer subagents. I'm not saying I stay under that on every task, but I do have quite a few sessions that cap out well below that, so that price difference would be very meaningful.

I use Luna for this day in and out and its excellent - if Haiku is that much better I will be changing things up.

➕ show 1 reply
dannyw • today at 6:16 PM

Haiku 5.5 is noticeably smarter than GPT-6 Luna, so I can see their pricing strategy here.

For a while Anthropic has lacked a cost effective “cheap” LLM for summarisation, compacting, RAG helpers, etc.

These ‘ephemeral’ workloads are often under 100k tokens, or can be structured to be under 100k.

In some coding benchmarks, Haiku 5.5 beats Sonnet 5! (Especially implementation; do a well defined Jira ticket; etc), it’s really impressive how much intelligence per dollar has grown in just a few short months.

➕ show 4 replies
Tiberium • today at 6:18 PM

There's also a tokenizer efficiency difference: modern Claude's 100K tokens are about ~60-65K modern GPT tokens, so in reality the Luna cutoff is much further away than the Haiku one.

You can test with Anthropic's count_tokens endpoint or with https://crates.io/crates/tokwc

➕ show 1 reply
Eridrus • today at 6:14 PM

It's actually existing flat per-token pricing that is weird.

Neither encode nor decode are linear in compute, so providers need to price for average expected length.

This is just getting closer to the true cost of generating tokens.

➕ show 3 replies
HarHarVeryFunny • today at 6:24 PM

Notable that one suggested use case for Haiku is "classification requests", i.e. Jev competitor, and the pricing matches GPT-6 Luna which is behind OpenAI's "Decisions API" Jev competitor.

For this application 100K token input is plenty.

Of course Anthropic and OpenAI, both at $0.10/M, are still 2.5x the cost of Jev's $0.04/M.

➕ show 1 reply
tr4656 • today at 6:10 PM

Luna does as well, but just at a higher limit.

From OpenAI's website: Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.

➕ show 2 replies
alexchamberlain • today at 7:27 PM

Isn't it less than a year since Claude models went from 100k token limit to 1M limit? Don't get me wrong - my main agent normally gets to 25% or so before I clear it these days, but as a subagent, doing research or summarisation, I don't think 100k is "absurdly low".

port3000 • today at 6:34 PM

They are targeting businesses/API use for fast decision making and agent integration. Plus they now need to be competitive with Jev-type models in that space.

➕ show 1 reply
mkotlikov • today at 8:41 PM

If you look at how different reasoning levels can easily exceed task cost of sonnet 5.5 you will see that you will basically never fall into that under 100,000 token threshold. I mean maybe you can choose low and do a basic summary task, but then you could choose something much cheaper instead. I don't know what Anthropic is thinking with its dumber models.

mnicky • today at 7:07 PM

You could also use it as a subagent prompted eg by Sonnet/Opus orchestrator agent and for many agentic workflows significant part of the dispatched tasks might be under 100k budget.

giancarlostoro • today at 6:08 PM

I with they'd give Haiku like 400k tokens roughly, I think between 400k or even 600k tokens is a sweet spot, but Haiku is basically designed to be for small edits is my understanding, but it sucks because any time I ask Opus to "try" letting Haiku do the work, it just falls apart and Opus comes back and tells me it switched to Sonnet (even before Sonnet finally jumped up to 5.x).

I will try the new Haiku, but it would be worthwhile if Haiku could take sane instructions and do all file editing for Opus / Sonnet / Fable then it would be worth using.

insanitybit • today at 6:16 PM

I mostly use Haiku for really, really basic stuff, never for actual engaging work. I've used it for first-pass analysis to triage bugs, for example - all it does is related N bugs together to see if any potentially relate. Then I have Sonnet investigate further.

enraged_camel • today at 6:18 PM

>> 100k tokens is an absurdly low cutoff and it is only applicable to Haiku and not Sonnet or Opus. It's a low enough cutoff that it will be quickly exceeded if you are doing anything with Agents

Your vibes don't appear to be supported by facts. From the announcement:

>> Claude Haiku 5.5 is priced 90% lower than Claude Haiku 4.5 for requests up to 100,000 tokens, and 50% lower for requests over 100,000 tokens. On Haiku 4.5, 90% of requests fell into the former category.

➕ show 2 replies
sixtyj • today at 8:23 PM

Chatbot could be < 100k tokens.

AustinDev • today at 6:17 PM

encode and decode tok/s which is ($/s) when it comes to pricing drops heavily above 100k tokens.

There are plenty of workflows like translations where you'd easily be under the cap.

system2 • today at 6:17 PM

Who in their right mind would use haiku while Mimo or GLM cost 10% of what they are charging with much smarter models?

➕ show 9 replies
esafak • today at 6:17 PM

It's their creative way of 'matching' Luna's prices.

j45 • today at 6:12 PM

It could be to incentivize people to not be lazy users of tokens.