logoalt Hacker News

drakytheyesterday at 7:46 PM0 repliesview on HN

I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.

Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?