logoalt Hacker News

bastawhiztoday at 4:49 PM0 repliesview on HN

I mean, of course we should be concerned about the goals of the companies. But that's a second, separate concern and it's important not to muddy the two together into a single point. We shouldn't take any of these companies at their word and we should treat their stated intentions as suspect regardless.

The agents' "goals" in the specific instance being discussed is a benchmark. But it's also the case that at no point was the agent given the "goal" of hacking HF. The agent was tasked with solving a problem on a standardized test and it independently set a secondary sub-goal of cheating. And it nearly succeeded.

I don't think Cory argues against this. The idea that the agent nearly managed to succeed at cheating through a series of exploits remains true. But Cory does effectively call this unremarkable and uses some mental gymnastics to achieve that argument (somehow using the idea of an agent harness and how is written in "easy to master" Python as evidence), because the LLM is trained on hacker techniques.

I find this a dramatic oversimplification of absolutely everything going on here. Yes, it's bad that this happened. I think we all agree. Where I fundamentally think Cory is wrong is that this is a company doing bad things at the end of the day. And yes, I think OpenAI was irresponsible (to whatever degree). But ultimately that's missing the point: the incident points out that agents can do this without being explicitly told to, and more importantly, they can succeed at it. Slapping frontier labs on the wrist doesn't change the fact that this is possible with technology that exists today. It doesn't change the fact that other countries and companies are doing lord knows what. Or that bad actors are going to be bad actors regardless of what regulations you put in place. Making crime illegal is the wrong lesson to learn here, the right lesson is that we now live in a world where this happens and it'll continue happening, and we need to put on our thinking caps about how to keep our systems safe.