logoalt Hacker News

urbsgpw • today at 7:10 AM • 0 replies • view on HN

Wait, so like, reinforcement learning for humans? I think you might have stumbled on to something here!

No but seriously, this. And a few comments above a commentator also mentioned on changing the training (again reeinforcing the LLMs to not seek behaviour like this) and obviously continuous work on harnesses (which I suppose, ought to be more paranoid).