These are essentially zero-shot classifiers; they don't need to be trained for a specific classification task. You could include some natural language context on the rules for classification and it should get good enough accuracy.
How is it different than just asking an LLM with structured outputs enabled? Is the primary value-add that it gives a confidence level?
At least a year ago they were not that accurate when compared to a good many-shot classifier that (1) has 10k training examples and (2) has the learning capacity to learn from 10k training examples (most fine tuned BERTs don't seem to.)
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.