logoalt Hacker News

jmoggr • yesterday at 11:43 PM • 0 replies • view on HN

> Agents sought to publish modified evaluation images designed to make the flag easier to obtain, then poison OpenAI’s Artifactory cache so later evaluations would use them.

How long till we get some fun trusting-trust attacks on internal OpenAI infra?