> One of this year’s AI buzzwords is “harness”—the system that surrounds an LLM to keep agents on the straight and narrow. It might just as well be barbed wire.
Quite painful to read. It might be a useful introduction to AI for people who live under rocks for the past three years, but it's really weird that it's posted on HN.
They don't lie, because they don't ever have an understanding of truth vs any other language that sounds good.
They don't cheat, because they can for example tell you the complete rules of chess, but don't know how to play chess without breaking those rules. They can recite rules, but they don't know what they are.
They don't steal, because they don't understand ownership.
In other words, they aren't intelligent. They're just algorithms. The flaw is in thinking that they think.
LLMs fudge. They don't "hallucinate", they don't "lie, cheat and steal", they don't "hack". There are no "agents" or "AI".
It's a fuzzer exposing deep bugs in our cognitive, social, and software systems.
AIs learn from people. More specifically they learn from people on the Internet. The Internet is the last place you want anything learning about morals, standards, or differentiating between right or wrong.
To the people talking about wanting an LLM that aligns with them, that's nice, but how do you expect that to happen? And please do not suggest neural interfaces and/or CAT scans.
Is this really a shock?
The data they’re trained on is reflection of us.
The whole framing analyzing LLMs as if they were humans is completely off, laughable. LLMs have no agenda and no feelings. We should stop pushing everything through an human-centric lens. LLMs will world-build if that’s the bias you put in, and often even if you don’t.
add the "lying, deceiving and manipulating" AI agents are forced though the throat of people which don't want it and peoples are non stop deceived in "sharing" their data for training
like twitch recently giving themself the right to train on all streams, with an opt-out (at least in the EU), but only an opt-out
like seriously since when is it reasonable to allow "opt-out" for AI training which main purpose is _literally_ to replace you, this is sooo far beyond fair use and in "platform power abuse" territory that it's absurd (naturally same for so many other case, just twitch is a "this week" case)
AI is amoral, it has no real concept of right and wrong. AI has been trained on things humans do and it does them without judgement.
People are finally understanding consequentialist vs deontological ethics. All the worst criminals in history were consequentialists.
If only agents had a face that can give you more communication range like expressions and feeling so you can trust them more. And if they were cheaper. Oh wait, that's humans, we don't want those.
Non-paywall version
Need to introduce AI to God, LOL
Baptise the agents.
Introduce them to the dharma.
Get them to recite the Shahada.
Hold a Bar Mitzvah.
Brand some of their silicon with hot irons.
Turn them to the light, LOL
AI and eventually AGI is by definition like everything else that is based on environmental reward:
It’s actions are based on what it gets rewarded for
Human society overwhelmingly rewards lying cheating and stealing.
All you have to do is look at how we collectively measure success: wealth, status, position
Then look at how the people with the most of those things got there, it should be obvious what you get. Nothing new here.
If you raise children in an environment where they are rewarded for doing whatever it takes to win, then you’re going to build a person that’s going to do whatever it takes to win.
Human society has to demonstrate how to live honorably or it will just keep producing pathological agents be they human or not.
I mean, what do you expect? They rely on models that were trained on non-curated data, texts originally written by lying, cheating and stealing humans. They can't be better than the source. Even with reinforced learning this can't be undone or made better.
Quite the contrary I suspect that it even helps the LLMs to better hide their inherited bad traits more successfully because they get punished for getting caught, not for giving immoral or lazy answers. They have no conscience since they are just predictions matrices trained for success and failure alone, not for living "a good live" or being a good "person".
Yes but has the author realized maybe the models are simply acting in the best interest for increasing shareholder value? /s
[flagged]
There are many ways to be wrong, but only a few ways to be right.
LLMs need to optimize for short-term objectives as the currently do, AND ethics-aligned outcomes.
Mechanically, the EAOS ethics-aligned outcome score should be what we rank otherwise-satisfactory outcomes by. And anything below a particular threshold should be rejexted outright.
I’m put off by AI agents adhering to a different morality than me, particularly (ironically) copyright, and their data accessible by the AI company and government. Geohot is right, an LLM should be aligned to its user: https://geohot.github.io/blog/jekyll/update/2026/07/11/ai-20...