> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no?
I know for some types of ML analysis, a separate model is already used to analyze the weights.
This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.