I'd love to see benchmarks on Decisions API vs an actual LLM call.
Luna is so cheap it's borderline free (without tool use), so I'm struggling to figure out where to use this/Jev.
For example product categorization. Why 'risk' using this/Jev when a Luna LLM call will be smarter (in theory)?
Probabilities are one reason. If the Decisions call returns a decision with a low probability you could route that decision to a smarter model.