Agree, except the probabilities for outcomes in the structured output. I don't think you can get those for most frontier LLMs (logprobas). You can get it for open source models but not frontier LLMs.
That number is a big deal, assuming it is well calibrated. Did they talk about calibration?
That number is a big deal, assuming it is well calibrated. Did they talk about calibration?