In the Jev use case, LLMs are horribly uncalibrated. In general, they will not produce good probability estimates.
Their generality also comes with a latency/computation costs.