logoalt Hacker News

6thbityesterday at 11:59 PM0 repliesview on HN

    > closer to a harness and operational failure than a model alignment failure. Our models were told they had no internet access and to capture the flag, while in fact being misconfigured to have internet access. 
    > This led them to believe—arguably reasonably—that the real environments they encountered were simulations.

That the AI lab most typically preaching for alignment does not consider this an obvious misalignment is a clear red flag.