I put Decisions API against classic statistics experiments: (1) loaded coin, and (2) marble selection from a jar, with replacement. I tested both predicate and choice questions. I ran thousand trials against each experiment, and I also did an experiment where I change the order of choices, to see if it matters. Summary:
* Using predicate questions gave nearly perfect/expected probability outcomes
* Asking it to choose an outcome behaved differently from drawing randomly - if the true probability of a red marble draw was 50%, using Decisions API produced 86%, i.e. it picked the right marble but gave a significantly more biased weight on its choice
* Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
It's still a predictive model under the hood, so order will always matter. They are also just as open to prompt injection and bias/loaded questions as the underlying model.
I wonder if a better model (and/or higher effort level) than GPT-6-Luna would produce a different result.
> * Changing the choice order changes the probabilities! Moving the red marble from first to last choice changed its probability estimate from 86% to 73%
These models seem to have the same biases and limitations of LLMs minus speed. Outputs and inputs should be treated the same way as LLM prompts.