The judgements are available on huggingface so training is possible. But given that we evaluate on new queries daily there would need to be some generalization happening for this to show up in the benchmark.