logoalt Hacker News

JamesStuffyesterday at 7:59 PM14 repliesview on HN

Personification of AI is what’s going to get us in the end.

I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!


Replies

jononortoday at 8:01 AM

Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks. We cannot accept "AI" absolving humans of responsibility.

johnisgoodyesterday at 8:32 PM

Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?

Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.

show 1 reply
trio8453yesterday at 8:43 PM

The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".

show 3 replies
mikestorrentyesterday at 8:07 PM

This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).

trio8453yesterday at 9:04 PM

> An AI didn’t hack into a company, the engineer set an automated tool to.

What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?

Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".

jagraffyesterday at 8:11 PM

I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.

In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen

show 5 replies
pizza234yesterday at 8:06 PM

> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.

No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.

show 2 replies
grumpopotamusyesterday at 8:08 PM

Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.

AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.

show 2 replies
jacquesmyesterday at 8:56 PM

I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.

trio8453yesterday at 8:39 PM

Do you get upset when we say that "a program is running" when we all know it has no legs?

Kim_Bruningtoday at 10:20 AM

You know, I actually think it's the refusal to consider personification that's going to get us.

Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.

So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1

A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.

You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.

And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?

huurtehoogyesterday at 8:05 PM

[dead]