logoalt Hacker News

dgellowtoday at 7:38 PM4 repliesview on HN

> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks

Wait, what? Am I understanding that correctly? That sounds really bad


Replies

drakythetoday at 7:46 PM

I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?

pixl97today at 7:47 PM

Nothing to worry about citizen, ignore the fleet of drones flying overhead.

order-matterstoday at 7:50 PM

the bullshit machine is learning to optimize its bullshitting techniques!

<AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt

you cannot give this type of worker autonomy over anything.

Laurel1234today at 7:58 PM

[dead]