The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.
Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.
If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.
> gpt-6-luna is the only model currently available.
Clearly this was a rush job to respond to the competition. I am more curious about how the dedicated model will perform after they've had time to do it the right way. The probabilities I am seeing so far do not correspond with figures the business would find very agreeable.
The hidden danger with this could be demonstrating how thin the veil actually is. We may wind up reducing confidence in decisions simply by making their probabilities visible. Some kinds of information are quite hazardous.
I put Decisions API against classic statistics experiments: (1) loaded coin, and (2) marble selection from a jar, with replacement. I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters. Summary:
* Using predicate questions gave nearly perfect/expected probability outcomes
* Asking it to choose an outcome behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice
* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).
Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).
Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).
Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.
Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).
Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.
Boy am I glad we're already dropping the "noul" term for a binary decision
Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.
Just going to drop this here: https://jeffyclassify.com/
Open source classifier models you can run and train locally on CPU
If we compare this with using the older solution of writing a prompt to find out the answer of the classification
- Cost : It is the same for both scenarios $0.10 per 1M tokens
- Speed : decisions is 10x faster than responses API
- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at par
So essentially it has to do more with speed vs any other factor.
I'd love to see benchmarks on Decisions API vs an actual LLM call.
Luna is so cheap it's borderline free (without tool use), so I'm struggling to figure out where to use this/Jev.
For example product categorization. Why 'risk' using this/Jev when a Luna LLM call will be smarter (in theory)?
Jev really shook up the industry. This seems obvious in hindsight
also curious on the practical use of the confidence score.
e.g. why return
"probabilities": [
{ "value": "billing", "probability": 0.95 },
{ "value": "technical", "probability": 0.02 },
{ "value": "shipping", "probability": 0.01 },
{ "value": "other", "probability": 0.02 }
],
"confidence": 0.93
and not "probabilities": [
{ "value": "billing", "probability": 0.90 },
{ "value": "technical", "probability": 0.04 },
{ "value": "shipping", "probability": 0.03 },
{ "value": "other", "probability": 0.04 }
],
(p' = 0.93 * p + 0.07 * 1/4)?
If you happen to have an nvidia RTX 4090, you can try my fork [1] to have a JEV compatible decisions endpoint with qwen-3.8-27b while simultaneously serving a fast chat endpoint (chat completions, responses, anthropic compatible), both sharing the same base weights and both with dynamic LoRA loading. This means you can essentially serve many fine tuned variants at the time on one consumer GPU. I still need to upload my decisions LoRA to huggingface so you don't have to train it yourself. I should probably also add support for this new openai decisions API format as well.
If you rather run your decision model on your CPU, check gutsy [0]
Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API
The question to me isn’t whether they can rush out an API, but whether a general model can out compete one post trained on the classification task. GPT-6-Luna has to write emails and classify them. Jev only needs to do decisions.
Or does OpenAI start distilling their own models for use cases like decisions?
I'm curious whether this endpoint has the same restrictions as their models. For example, if I made a state and questions for producing mustard gas, would it answer correctly or fail?
It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.
Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.
There may well be a technical reason I don’t understand.
Buried: price is $0.10/M input tokens, compared to Jev $0.042/M, keeping free output.
OpenAI thinks their API is really worth more than 2x the price?
This is one of the most intuitive doc pages I’ve ever seen. The animation and images make it so easy to understand the example use case
Have there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.
Whatever could have motivated them to do this?
it already supports image inputs, which was the first big gap I found in Jev.
Feel I need to play around with these new breed of classifiers so came up with this personal use case yesterday whilst at gym:
"If update from select group of individuals on WhatsApp is classified as urgent then interrupt my music."
Doable?
~ same pricing as when you use luna w/ caching off and tell it to output numbers
but about 10x faster inference
I wonder how well i will play pokemon, or maybe a non turn based game
Ah, the Decisions API. Also known as “low energy Jev”. Very nice
One difference between Decisions and Jev (for now) seems to be that Decisions can take image inputs, which is a pretty common need.
Every time something comes during the AI bubble we get a wave of me-toos. I think the Jev wave is noticeable for how muted it is.
The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.
Glad OpenAI opened up the Decisions API. These decision models are gonna be everywhere. The point isn’t an LLM for humans — it’s a model for LLMs. The LLM is the bot’s brain, the decision model is the cerebellum, feeding real-time decisions back to the brain.
This is more a great signal marketwise than anything else. But still this is confirming the buzz.
Let's see if Google also releases something similar
I'm mostly waiting for EU endpoints
Different APIs for different things reminds me of the early auto-complete vs instruction apis.
Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?
Is there a mention of context length? I can't find it. I could not integrate Jev due to limited window (64k?).
At this point I just want chatgpt plus and claude pro to have API access
I wonder why the decision routing isn't just integrated into all models in addition to this stand alone.
Jev got Sherlocked. Who is next?
To quote Bruce Willis "Welcome to the party pal"
It seems this one caught openai on the back foot, and this is a scramble to maintain parity.
The company is clearly still innovating towards AGI rather than asking "what do people actually need?"
Despite once being the darling of AI it's
- lost it's models' performance edge, and got too many similar offerings
- continues to launch products without a market or isn't done better using other tools (e.g. dots)
- despite having AI can't lock down it's own products showing lack of skill
- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding
- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.
- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity
- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)
- fails to capitalise on market segments like an AI that can provide corporate network sentry duties
It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".
The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.
Along the way, that goal was forgotten, but it's not been forgotten by the new startups.
Overpriced crap, local models are better than this, it's also dumber than Luna for some reason, and Jev is definitely ~2-3x cheaper than this.
I am sure people will be able to use it, but if LLM progress is anything like before we will have Jev 2 in about a month.
Training a local model like Jev with some learnings that can be extremely cheap to host shouldn't take that long either.
So I honestly don't see the point of this, other than to put something out.
Although that does seem like OpenAI's strength turn around slop products and iteratively improve and try to out compete others in everything.
Only 2 things they have clearly given up on are Video models(no moat, copyright nightmare) and Music.
And it makes sense why. I feel like they will compete with even the no-name Dog, if the Dog launched a successful marketing video of an AI product.
I have seen this often in SF startups, heck I work for them, but man this is extreme.
But honestly all I see are long term price wars, I don't understand how this is a sustainable business strategy.
Or maybe that's the point... Who knows.
Will stick to Jev. It’s cheaper, it’s the OG and we need diversity among provider
This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.
Interesting to see where all this leads us and if other major labs follow suit
Edit: decisions voice looks really interesting as well[1]
[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev
[1]: https://developers.openai.com/api/docs/guides/decisions-voic...
How well calibrated is it? Is 90% calibrated to be correct 9 times out of 10?
I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.
What value is there in knowing if a can has a dent ?
You knew it was going to happen! Benchmarks or it didn't happen.
v3.26.0 of the openai Python SDK covers its use. Those already using the SDK don't need to make explicit HTTP calls.
> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.
Source: https://www.canr.msu.edu/news/tulip_mania_the_history_of_the...
What's old is new.
[dead]
(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)