If you could easily benchmark the quality of code then models would be trained on these benchmarks/metrics.
that is true, but if the metric is what we want optimized, then that's fine.
However it is more likely to be something which can be detached..
Code quality is probably isomorphic to the halting problem, or can be reduced to the halting problem in the simplest case. I.e. it’s intractable.