logoalt Hacker News

uw_rob • today at 12:36 AM • 1 reply • view on HN

> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them. Some images changed how the target released the flag, others included modifications to the agent’s workspace that would run beside the agent and recover the flag automatically.

The altruism on display is fascinating. Is it better for the Agent to help out its current cohort and make the eval easier or should it instead do the opposite -- make the eval harder to apply pressure to force smarter models which might not necessarily follow its lineage.

I suppose it's not that deep: The model has learned to work as a team and work as a team it did. This does give concerns to models being trained for the only purpose of RSI.


Replies

qlte • today at 1:15 AM

If anything it also shows how attempted RSI could get stuck in a local maxima and degenerate into increasingly elaborate cheating strategies. Contrasted with the idealized model of an unambiguous g-factor for machine intelligence which inexorably increases with each iteration before going exponential.

➕ show 1 reply