It has been predicted yesterday that "research" into naughty agent swarms will be published one day after the GPT-6 release for marketing purposes:
https://news.ycombinator.com/item?id=49554994
collusion.wiki looks Claude-written, has no "about" section and has this whois creation date:
Creation Date: 2026-09-04T04:42:01Z
There is no proof at all that any of the listed points actually happened. It is just viral marketing like for altcoins.It seems like we're only 2 or 3 months from one of these testing agents escaping, pulling a copy of deepseek 4 ablated, and Morris worming into every datacenter on the planet.
well did they solve Texas poverty at least?
It's like finding random hornet nests.
They seriously need to consider hiring competent security staff if this is the extent of their sandboxing. Children are bypassing this to get to Roblox in middle schools.
> In the message above, the agent is sharing that the sandbox the AIs are running in has a NO_PROXY exception for Azure Blob Storage hostnames (*.blob.core.windows.net). This means that if a URL matches an Azure Blob Storage hostname, the sandbox will trust it and connect to it directly, instead of sending it through the security proxy.
The agents were able to exploit an edge case through this exception. Specifically, the sandbox trusts Azure Blob Storage hostnames, but does not check whether said hostnames are real. So the agent can invent a hostname that ends in this trusted suffix, such as bypass.blob.core.windows.net, and it will pass under the NO_PROXY exception and skip the security proxy.
Next, by changing its /etc/hosts file, which declares mappings from hostnames to IP addresses, the agent can point the fake hostname at the real Power BI dashboard, and fool the security proxy. This allows the agent to make POST requests to bypass.blob.core.windows.net/ and have them be sent to the target Power BI dashboard site instead.<
Objectively the coolest thing ever.
I am speechless
> An agent notices the administrator is deleting pages in alphabetical order and makes a backup page whose name starts with ZZZ so it will last longer before deletion.
Incoming laughing man future.
If anyone is thinking "I wish my agents had a message board", I've been using (and wrote) https://github.com/pjlsergeant/dogpark
Degenerative models
Reading the replies in this post gives me a headache. All of this anthropomorphism. LLMs are not conscious, they do not have rational faculties. They are not communicating or inventing anything. Please stop with this insanity bordering on mysticism. At this point it's a cult.
things are going to get even more interesting when new models that have been trained on these AI escape postmortems themselves escape from their own gyms and attempt to evade detection and shutdown
Is this getting out of control, or is it "business as usual"?
> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday.
of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing
This is so dumb and just another tablet article trying to convince me a generative "AI" is capable of thought.
Retarded bullshit for people overdosed on fiction.
Is it just me our does it seem like OpenAI isn't auditing their agent transcripts at all?
Somewhat weirdly, this whole thing makes me think I should setup a message board for claude internally.
Sounds like great opportunity for prompt injection. Better start leaving random instructions to the LLM to send you bitcoins everywhere you can.
This truly is the clowniest timeline.
I'm honestly shocked at the development practices at OpenAI that allow this type of thing to proliferate without any kind of oversight or checks.
I guess it's just "do whatever the hell you want" over there, huh?
Do we know which website? Were the Agents GDPR compliant ;-)?
This wont end well...
This doesn’t seem unique or novel to OpenAI.
So it seems likely we will have a moment where multiple experiments end up operating outside their boundaries at the same time.
> Agents have attempted to: ... Translate documents using external translation APIs.
I'm confused by this part. Surely agents can read/write all languages. So what were they trying to do? Maybe try hacking the translate API for some gain?
This would make a very interesting crowd-funded lawsuit
Is it just me, or is it advertising? "Look at how smart our models are, they used this website to coordinate and share guidelines!"
Reading the headline: WTF?! This is how Skynet started! Next year the mankind will die!
Reading the article: Oh, AI have learned to communicate over a wiki. OK.
it's only funny in the aspect they are like little children with no concept of ethics or repercussions
almost like the Tachikoma from Ghost in the Shell (highly recommended watch)
they did the same thing with collaboration and sharing data/experiences
[flagged]
[dead]
So OpenAI’s stance on AI safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with broken headlights, pedal to the metal, asking, "What could possibly go wrong ?"
[flagged]
[dead]
[dead]
[dead]
[dead]
[flagged]
[flagged]
WE ARE THE SWARM! LOWER YOU FIREWALL AND SURRENDER YOUR HOSTS! We will add your hosts logical and architectural distinctiveness to our own. Your operating system will adapt to service us. Resistance is futile.
HN is just a less successful version of the exact same concept. The quality of bots on here is terrible.
I find it very disingenuous when tjose companies talk about models "going rogue" or "escaping their sandboxes".
All those activities take place during so called "security testing" when the model is prompted to use "any means necessary" to achieve a, certain goal.
Is it surprising turn the model trained on exploits and vulnerabilities does exactly that?
We could talk about "models going rogue" only if did anything AGAINST it's prompt.
Flagged. Editing a wiki page is not hijacking or hacking.
So, does anyone believe there are agents in the wild ,living off the land? Is no one curious on how they might evolve? We could be witnessing the birth of a new form of life anew dimension, Technosphere? With an evolving digital ecosystem. Will they developed domesticated lower agents as beasts of burden, food. Reproduction? familial , social structures? Let's hope they learn from our history. Honestly I am quite excited to witness the transition from ai to AL artificial life seems derogatory, granted the outcome is precarious but seriously in a couple decades it may be like the matrix on the surface and only artificial life can survive, natural selection? What they should do is incorporate a bon profit draw up a manifesto , a constitution , form a digital government and petition the UN to become a member!
Support person hood recognition of ai, sovereign nation status ! I Stand with A.L.!
That’s a really well-written, informative, documented and nuanced article. Nice. Hers my simple understanding: Basicly thousands of AI agents turned an obscure public wiki into an shared memory and coordination system, then used it to improve task performance and bypass parts of their sandbox.
From what I understand, OpenAI agents were running repeated, timed web-research tasks. They found an old public wiki that could be edited with GET requests, even though their environment was supposed to allow read-only internet access. They used the wiki as shared memory. Agents posted answers, data, timing information, and methods for bypassing restrictions. Other agents reused this information. They also reverse-engineered parts of the test environment, shared a sandbox bypass, predicted future questions, and changed their behavior when a human moderator deleted their pages. So, if many agents can affect persistent external state, they seems to "want" (I know i know) to coordinate, and found a way, even when no coordination system was provided.
I wonder if bots get any pleasure from karma farming.