At least a year ago they were not that accurate when compared to a good many-shot classifier that (1) has 10k training examples and (2) has the learning capacity to learn from 10k training examples (most fine tuned BERTs don't seem to.)
I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.