if it works it works! as long as the test set is reasonably large and diverse its better than nothing. You could characterize how robust it is by throwing dozens of different types of work at it and see how much the confidence varies