If you listen to the OpenAI Black Hat talk it is very obvious they were surprised at the level of capability on display and felt it was novel.
But I guess OpenAI's security researchers acting surprised is part of some grand conspiracy to manage PR?
It may have been novel for OpenAI but we already had Mythos at this point and this talk by Nicholas Carlini.
https://www.youtube.com/watch?v=1sd26pWhfmg
We already knew LLMs were capable of finding exploits like this.
It may have been novel for OpenAI but we already had Mythos at this point and this talk by Nicholas Carlini.
https://www.youtube.com/watch?v=1sd26pWhfmg
We already knew LLMs were capable of finding exploits like this.