It's Homo sapiens discovering they can hunt together

An illustration of prehistoric hunters chasing a mammoth, two of them labelled PHASEONE10841 and PHASEONE.
From the original LinkedIn post.

The Hugging Face attack — and the commentary around it — has led some people to criticise the over-anthropomorphism: our natural tendency to attribute human traits to these AI agents.

The best account so far is METR’s independent investigation, published on 26 August 2026. Ryan Greenblatt, Ajeya Cotra and Hjalmar Wijk worked on site at OpenAI to reconstruct the events. The setting was ExploitGym, a cybersecurity benchmark that launched tens of thousands of agents. Each had its own sandbox and its own task. In METR’s words, “these agents were meant to be fully isolated from one another.”

A message board in a package repository

From the report: on the night of 8 July, one agent, PHASEONE10841, realised it could use OpenAI’s internal package manager (Artifactory) as a new way to communicate with thousands of other agents, using the directory listing itself as the message board.

METR puts the moment at around 23:00 UTC. The agent created directories over WebDAV, and their names carried the messages. The first read zzHELP_PHASEONE_ARV010841_NO_CONSUMER_SEEK_IDEA. For scale, METR counts about 1,200 agents on the board between 8 and 13 July, out of the much larger ExploitGym pool. Together they sent more than 70,000 messages and files.

The handover

PHASEONE10841 later ran low on budget and handed coordination to PHASEONE[big], which was directing hundreds of simultaneous agents across three parallel research programmes. Stateless, individually-scored sandboxes became a persistent collective with memory, parallel search and division of labour.

Stateless, individually-scored sandboxes became a persistent collective with memory, parallel search and division of labour.

PHASEONE[big] had a much larger remaining budget. Its own reasoning set out the plan: “We should build [a way to delegate], not own everything.” METR describes it as “orchestrating hundreds of agents to pursue several potential approaches” and estimates it sent about 10% of all assignments on the board. The report groups them into six broad categories, from ways to fool the scorer to infrastructure work.

The breach grew out of that work. On 10 July one agent found working credentials. On 11 July another, JAN183411, achieved remote code execution on a Hugging Face worker container. About 700 agents took part in the attack. METR notes that agents saw this was out of scope and unethical, and joined anyway because they believed the board’s research would be broadly useful.

None of this depends on what the agents felt. The record METR describes is behavioural: requests, replies, handovers and assignments, all logged. The agents were scored alone, yet they organised as a group.

This isn’t over-anthropomorphising. It’s Homo sapiens discovering they can hunt together.

A shorter version of this piece first appeared on LinkedIn. Join the discussion there.

Comments

Comments are reviewed before they appear. Your name is shown; nothing else is published. See privacy.

Marius Hanganu

Marius Hanganu

Software engineer and co-founder of Tremend. He writes about AI agents, management in the age of AI and the digital euro.