The problem is, if you explicitly say “this is an eval” then you get different results. It can move the capabilities both up and down on the narrow metric depending on the scenario.
The choice to run the model with relaxed safeguards seems questionable. At a certain level of capabilities it is just unsafe to run a raw model; we are clearly close to if not at that point.