This isn't really anything new it just seems like a new API but you could do the exact same thing with just a little bit of prompt engineering all the way back when GPT-3 was first released. Am I missing something?
It's much faster and cheaper (an order of magnitude).
And theoretically will give you better answers statistically as it's calibrated.
No amount of prompt engineering will give you the true probabilities for the model producing a certain response; this is something you can only get by inspecting the internal state at inference time.