logoalt Hacker News

zahlman • today at 9:21 AM • 0 replies • view on HN

Am I reading this right? 3 minutes to acknowledge the alert, more than two hours to act on it?

> When training restarts, we will begin a fresh run with additional alignment improvements, including more comprehensive misalignment interventions.

But they apparently won't do anything to address the possibility that simply trying to make the model behave might not work. They won't actually make sure "that the model could not access the live internet" by, for example, creating a physical hardware environment that lacks this capability.