logoalt Hacker News

Decisions API is in public beta

376 points • by chiefstorm • last Tuesday at 8:57 PM • 219 comments • view on HN

Comments

simonw • last Tuesday at 10:27 PM

  curl https://api.openai.com/v1/decisions \
    -H "Authorization: Bearer $(llm keys get openai)" \
    -H "Content-Type: application/json" \
    --data '
  {
    "model": "gpt-6-luna",
    "input": [{
      "role": "user",
      "content": [
        {"type": "input_text", "text": "I am angry about the new product feature"}
      ]
    }],
    "questions": [{
      "type": "predicate",
      "name": "complaint",
      "instructions": "Is this a complaint?"
    }, {
      "type": "predicate",
      "name": "compliment",
      "instructions": "Is this a compliment?"
    }]
  }'
Returned:

  {
    "model": "gpt-6-luna",
    "answers": [
      {
        "type": "predicate",
        "name": "complaint",
        "probability": 0.91
      },
      {
        "type": "predicate",
        "name": "compliment",
        "probability": 0.06
      }
    ],
    "usage": {
      "input_tokens": 310,
      "input_tokens_details": {
        "cached_tokens": 0,
        "cache_write_tokens": 0
      },
      "output_tokens": 0,
      "output_tokens_details": {
        "reasoning_tokens": 0
      },
      "total_tokens": 310
    }
  }
That https://api.openai.com/v1/decisions endpoint is notable because usually when OpenAI define an endpoint like that it ends up as a defecto standard for other providers.

(I turned this all into a new llm plugin: https://github.com/simonw/llm-openai-decisions)

➕ show 4 replies
TSiege • last Tuesday at 10:26 PM

The response to Jev should be the nail in the coffin over whether or not the AI business is a commodity market.

Out of no where Jev appeared as the next round of the price wars. Jev showed the value of System One models. A fast yes/no/confidence score not only is cheaper but also often all people want. Open source versions flood hugging face and now the big players are giving up a potentially big driver of output tokens to keep customers and race to the bottom price wise.

If I were OpenAI or Anthropic I’d be racing to make their products as sticky as possible bc ppl will flock to what’s cheapest otherwise.

➕ show 14 replies
bob1029 • yesterday at 7:43 AM

> gpt-6-luna is the only model currently available.

Clearly this was a rush job to respond to the competition. I am more curious about how the dedicated model will perform after they've had time to do it the right way. The probabilities I am seeing so far do not correspond with figures the business would find very agreeable.

The hidden danger with this could be demonstrating how thin the veil actually is. We may wind up reducing confidence in decisions simply by making their probabilities visible. Some kinds of information are quite hazardous.

➕ show 1 reply
armcat • yesterday at 12:56 PM

I put Decisions API against classic statistics experiments: (1) loaded coin, and (2) marble selection from a jar, with replacement. I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters. Summary:

* Using predicate questions gave nearly perfect/expected probability outcomes

* Asking it to choose an outcome behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice

* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%

➕ show 3 replies
Topfi • last Tuesday at 10:33 PM

Ran my decisions evals (still rudimentary, less than 600 calls (UI component selection, chat charting, tag selection, PKM stuff)) on this via OpenRouter against Jev and Mercury Decide. Jev because it has replaced my mt0 efforts by sheer force of affordability (more importantly, the limits running on a MacBook Neo bring even after vocab pruning and quant insanity) and Mercury Decide because I do like dLLM efforts (and I'd like to use fewer model providers if possible).

Preliminary of course, but seems to be slower than Jev and similar to Mercury Decides latency, though not in growing linearly with the amount of input (346ms p50 and 860ms p95, (Mercury Decide also had some extremes up to 1,3s that were around 800ms today, likely preview related, it scaled far more consistently with size)), less "confidence" concerning my ambiguous UI component and response shape specific tasks (have very specific use cases for these models which Luna often fails to meet at 0.6 and lower), lead to a few failed calls which neither competitor had (4 vs 0 for both) and measured more expensive than Jev to boot by a factor of 3,1 times on average (Mercury Decide pricing I think is still unknown so no numbers there).

Basically slower, more expensive and less capable than Jev, roughly on par with Mercury Decide (provided, in my insane set of use cases and requirements that are a PKM focused Firefox fork with multiple infinite canvas using decision models to improve information synthesis from multiple sources).

Seems a bit undercooked overall and I'd rather frontier-labs don't jump on bandwagons until they can offer something competitive in price, performance or both. In fairness, though, I have yet to test image input, maybe that makes all the difference. Also, again, mine is unlikely to reflect everyones use case, so interested in seeing others results.

Didn't comment at the time, but having read up on Devday after the fact, there seems to have been a lot of that going around. Notion and GDocs, Jev, Muse, most seems to have been cloned from existing competitors (and despite infinite, ultrafast, ultra code tokens with unsandboxed Mega Astra not that amazing to boot).

Prefer less announcements, but focused and at a higher quality. Considering ChatGPT Atlas (their Chromium based browser) and its insanely fast death, I'd be skeptical to put much into any of these even if they were in some way an improvement over what is out there. Maybe focus on a fresh pre-train and some sandboxing improvements.

➕ show 4 replies
isoprophlex • yesterday at 5:11 AM

Boy am I glad we're already dropping the "noul" term for a binary decision

➕ show 5 replies
agentdev001 • yesterday at 3:10 AM

Note that, being that this is gpt-6-luna under the hood, this offers you 1m token input window, and multi-modal (image) input. In my testing so far, I'm seeing 160-175ms end to end. Worst 5% 285ms, worst so far was 743ms.

➕ show 1 reply
nico • yesterday at 1:13 AM

Just going to drop this here: https://jeffyclassify.com/

Open source classifier models you can run and train locally on CPU

ashu1461 • yesterday at 1:31 AM

If we compare this with using the older solution of writing a prompt to find out the answer of the classification

- Cost : It is the same for both scenarios $0.10 per 1M tokens

- Speed : decisions is 10x faster than responses API

- Quality : I guess if we compare with luna which is a pretty good model it itself, both will be at par

So essentially it has to do more with speed vs any other factor.

➕ show 1 reply
EcommerceFlow • yesterday at 12:52 PM

I'd love to see benchmarks on Decisions API vs an actual LLM call.

Luna is so cheap it's borderline free (without tool use), so I'm struggling to figure out where to use this/Jev.

For example product categorization. Why 'risk' using this/Jev when a Luna LLM call will be smarter (in theory)?

➕ show 1 reply
sidcool • last Tuesday at 10:32 PM

Jev really shook up the industry. This seems obvious in hindsight

➕ show 1 reply
wise_blood • yesterday at 12:19 PM

also curious on the practical use of the confidence score.

e.g. why return

      "probabilities": [
        { "value": "billing", "probability": 0.95 },
        { "value": "technical", "probability": 0.02 },
        { "value": "shipping", "probability": 0.01 },
        { "value": "other", "probability": 0.02 }
      ],
      "confidence": 0.93

 
and not

      "probabilities": [
        { "value": "billing", "probability": 0.90 },
        { "value": "technical", "probability": 0.04 },
        { "value": "shipping", "probability": 0.03 },
        { "value": "other", "probability": 0.04 }
      ],

(p' = 0.93 * p + 0.07 * 1/4)

?

➕ show 2 replies
netsroht • yesterday at 7:25 AM

If you happen to have an nvidia RTX 4090, you can try my fork [1] to have a JEV compatible decisions endpoint with qwen-3.8-27b while simultaneously serving a fast chat endpoint (chat completions, responses, anthropic compatible), both sharing the same base weights and both with dynamic LoRA loading. This means you can essentially serve many fine tuned variants at the time on one consumer GPU. I still need to upload my decisions LoRA to huggingface so you don't have to train it yourself. I should probably also add support for this new openai decisions API format as well.

[1] https://github.com/tensorninja/ninfer-4090

➕ show 1 reply
mrkn1 • yesterday at 12:10 AM

If you rather run your decision model on your CPU, check gutsy [0]

[0] - https://news.ycombinator.com/item?id=49976996

mohsen1 • last Tuesday at 10:21 PM

Since it is fast and understand images, I wonder if it can play video games. I have a harness setup for the LLM play EA FC but even the fastest LLMs are too slow for it. I need to try this with Decisions API

➕ show 4 replies
softwaredoug • yesterday at 11:28 AM

The question to me isn’t whether they can rush out an API, but whether a general model can out compete one post trained on the classification task. GPT-6-Luna has to write emails and classify them. Jev only needs to do decisions.

Or does OpenAI start distilling their own models for use cases like decisions?

➕ show 1 reply
ranyume • yesterday at 1:07 PM

I'm curious whether this endpoint has the same restrictions as their models. For example, if I made a state and questions for producing mustard gas, would it answer correctly or fail?

AM1010101 • yesterday at 2:54 AM

It’s interesting they skipped caching. I could see wanting to ask follow up questions so having your first x tokens in cache would be interesting.

Also if you have a long “system prompt” then caching would have saved a considerable amount on bulk data processing.

There may well be a technical reason I don’t understand.

➕ show 1 reply
solatic • yesterday at 12:45 PM

Buried: price is $0.10/M input tokens, compared to Jev $0.042/M, keeping free output.

OpenAI thinks their API is really worth more than 2x the price?

➕ show 1 reply
sajithdilshan • yesterday at 12:44 PM

This is one of the most intuitive doc pages I’ve ever seen. The animation and images make it so easy to understand the example use case

elpakal • yesterday at 2:51 AM

Have there been any signals from Anthropic about matching this? We use AWS bedrock and just switched to Anthropic from OpenAI because of the ZDR guarantee. Would be great to not have to entertain switching back.

➕ show 4 replies
AbstractH24 • yesterday at 12:06 PM

Whatever could have motivated them to do this?

mritchie712 • last Tuesday at 10:21 PM

it already supports image inputs, which was the first big gap I found in Jev.

➕ show 1 reply
monkeydust • yesterday at 8:02 AM

Feel I need to play around with these new breed of classifiers so came up with this personal use case yesterday whilst at gym:

"If update from select group of individuals on WhatsApp is classified as urgent then interrupt my music."

Doable?

➕ show 2 replies
tosh • yesterday at 12:17 PM

~ same pricing as when you use luna w/ caching off and tell it to output numbers

but about 10x faster inference

qwertytyyuu • yesterday at 12:55 PM

I wonder how well i will play pokemon, or maybe a non turn based game

scionaura • yesterday at 1:07 PM

Ah, the Decisions API. Also known as “low energy Jev”. Very nice

stillatit • last Tuesday at 11:55 PM

One difference between Decisions and Jev (for now) seems to be that Decisions can take image inputs, which is a pretty common need.

➕ show 1 reply
Eufrat • yesterday at 4:44 AM

Every time something comes during the AI bubble we get a wave of me-toos. I think the Jev wave is noticeable for how muted it is.

The fact that this keeps happening demonstrates there is no moat. The fact that each wave gets a little less attention demonstrates there is no killer product here.

Sunny_Fung • yesterday at 12:05 PM

Glad OpenAI opened up the Decisions API. These decision models are gonna be everywhere. The point isn’t an LLM for humans — it’s a model for LLMs. The LLM is the bot’s brain, the decision model is the cerebellum, feeding real-time decisions back to the brain.

➕ show 1 reply
etienne_l • yesterday at 8:31 AM

This is more a great signal marketwise than anything else. But still this is confirming the buzz.

wise_blood • yesterday at 9:03 AM

Let's see if Google also releases something similar

I'm mostly waiting for EU endpoints

jasonjmcghee • yesterday at 12:30 AM

Different APIs for different things reminds me of the early auto-complete vs instruction apis.

Will this get folded into models / post training pipelines at some point and make them better at calibrated outputs?

pcwelder • yesterday at 4:55 AM

Is there a mention of context length? I can't find it. I could not integrate Jev due to limited window (64k?).

➕ show 1 reply
a_c • yesterday at 7:48 AM

At this point I just want chatgpt plus and claude pro to have API access

swader999 • yesterday at 12:28 AM

I wonder why the decision routing isn't just integrated into all models in addition to this stand alone.

amelius • yesterday at 10:19 AM

Jev got Sherlocked. Who is next?

SillyUsername • yesterday at 6:24 AM

To quote Bruce Willis "Welcome to the party pal"

It seems this one caught openai on the back foot, and this is a scramble to maintain parity.

The company is clearly still innovating towards AGI rather than asking "what do people actually need?"

Despite once being the darling of AI it's

- lost it's models' performance edge, and got too many similar offerings

- continues to launch products without a market or isn't done better using other tools (e.g. dots)

- despite having AI can't lock down it's own products showing lack of skill

- focuses on solving maths problems humans can do for tests, when real world problems - disease, materials, energy research etc is all outstanding

- abandoned it's open model and open source programmes, despite Google, and multiple successful Chinese, and now European companies make their frontier models open weight.

- pissed off a portion of its non corporate fan base by killing GPT 4o instead of recognising the brand and product attachment as an opportunity

- launched laughing stock projects like being able to actually call chatgpt on a telephone number (wtf!)

- fails to capitalise on market segments like an AI that can provide corporate network sentry duties

It's increasingly looking like the company has jumped the shark and if I was an investor would be asking questions as to why it actually took so long to bring a jev like product to market, and why they are labelling something that is a simplification of existing models as "beta".

The whole point of AI as I see it is to make our life easier and answer the questions we can't. It isn't to make an AGI so powerful that it can replace us.

Along the way, that goal was forgotten, but it's not been forgotten by the new startups.

minraws • yesterday at 5:32 AM

Overpriced crap, local models are better than this, it's also dumber than Luna for some reason, and Jev is definitely ~2-3x cheaper than this.

I am sure people will be able to use it, but if LLM progress is anything like before we will have Jev 2 in about a month.

Training a local model like Jev with some learnings that can be extremely cheap to host shouldn't take that long either.

So I honestly don't see the point of this, other than to put something out.

Although that does seem like OpenAI's strength turn around slop products and iteratively improve and try to out compete others in everything.

Only 2 things they have clearly given up on are Video models(no moat, copyright nightmare) and Music.

And it makes sense why. I feel like they will compete with even the no-name Dog, if the Dog launched a successful marketing video of an AI product.

I have seen this often in SF startups, heck I work for them, but man this is extreme.

But honestly all I see are long term price wars, I don't understand how this is a sustainable business strategy.

Or maybe that's the point... Who knows.

Havoc • yesterday at 10:50 AM

Will stick to Jev. It’s cheaper, it’s the OG and we need diversity among provider

Imustaskforhelp • last Tuesday at 9:01 PM

This rather didn't take long for OAI to create*, I remember people giving opinions and discussions that it won't take too long and that openAI should do it[0], so looks like they were right.

Interesting to see where all this leads us and if other major labs follow suit

Edit: decisions voice looks really interesting as well[1]

[0]: https://news.ycombinator.com/item?id=49802161: OpenAI is well positioned to fast-follow Jev

[1]: https://developers.openai.com/api/docs/guides/decisions-voic...

➕ show 1 reply
AM1010101 • yesterday at 2:51 AM

How well calibrated is it? Is 90% calibrated to be correct 9 times out of 10?

I think Jev had put significant effort here and its not clear if luna will be well calibrated in this way.

➕ show 2 replies
MiroslavPokorny • yesterday at 12:56 AM

What value is there in knowing if a can has a dent ?

➕ show 1 reply
lab14 • last Tuesday at 9:39 PM

How is the pricing vs Jev?

➕ show 2 replies
esafak • last Tuesday at 9:52 PM

You knew it was going to happen! Benchmarks or it didn't happen.

krembo • yesterday at 4:21 AM

But... Can it draw a pelican on bicycles?

➕ show 1 reply
OutOfHere • yesterday at 12:05 AM

v3.26.0 of the openai Python SDK covers its use. Those already using the SDK don't need to make explicit HTTP calls.

waterTanuki • yesterday at 12:37 AM

> The tulip became a luxury item and many varieties were introduced. The varieties were classified and the most sought-after, prized tulips were the streaked tulips, especially yellow or white streaks on a red or purple background. These flame-like tulips were highly sought after. Interestingly, the streaks or “flames” of the tulip petals were caused by a virus. The virus is the tulip breaking virus, or tulip mosaic virus.

Source: https://www.canr.msu.edu/news/tulip_mania_the_history_of_the...

What's old is new.

peterson_lock • last Tuesday at 10:19 PM

Can we use this through subscription?

➕ show 1 reply
hcentelles • yesterday at 12:57 AM

[dead]

🔗 View 8 more comments