A good point. I initially had a model for a judge, but it seemed to give very lenient scores. I'm open to learning about how best to benchmark the phenomenon, though.