Is simply changing the temperature so that the model appears calibrated over a particular benchmark after the fact “allowed”? Feels p-hacking esque.
if it works it works! as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies
if it works it works! as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies