That seems to be the key and its risky when the measure becomes the target itself. When you already know the benchmarks, and that drives the definition of success outcome then there is every incentive to just chase them
Yes, but that's with mostly the same data as before, just with more of the preexisting dataset and more RLHF. We haven't seen the ouroborous eat itself yet.
> as per benchmarks
That seems to be the key and its risky when the measure becomes the target itself. When you already know the benchmarks, and that drives the definition of success outcome then there is every incentive to just chase them