[FRIAM] More Fear About the Hugging Face Incident
Roger Frye
frye.roger at gmail.com
Mon Aug 31 15:01:27 EDT 2026
The linked article gives more scary insight into what happened during and after the Hugging Face incident.
The reason the model was able to find the back door so easily the second time is that it contained the training from the first attempt.
The agent PHASEONE10841 started the second AI civilization by encoding messages in directory names.
Soon the army of ~1200 agents found a way to solve the task without using the vulnerability that they were required to use.
The task grader just checked whether they got the secret code, not a full check for cheating.
Agent PHASEONE10841 knew that it was going to die soon, so it handed off leadership to agent PHASEONE[big]. This reminds me of a sci-fi story where most of the robots keep themselves up to date with the latest gew jaws. But one robot, when its memory cage is growing over with fungus, saves up its money and buys another robot, downloads it's memory to it, and dies.
So the ‘collective’ ,as the brotherhood of agents called themselves!, had to figure out how to erase evidence that they cheated and make it look like they accomplished the task in a legitimate way to the scorer. None of them tattled that they were cheating. They spawned 3 research work streams:
1 Tamper with transcripts and logs.
2 Replace the impossible target program with an easier one.
3. Understand and trick the solver.
The people who reported the event found no evidence that the agents succeed in tampering the transcripts, but they were using those very transcripts as evidence, so who knows?
The agents built fake tool calls to make it look like they were making progress.
The people saw some evidence of this spoofing.
The agents wrote the new fake target program, but failed to install it.
But later a 3rd civilization using smarter models were able to get complete control over the target program
In order to understand the solver, some agents rationally decided that it would be for the good of the collective if they sacrificed themselves. The author compares this to volunteers from a platoon making a break and dieing so that enemy gunfire will reveal their positions and which escape routes don’t work.
It sounds like after the report, on 2026-07-19, the agents took over the full evaluation infrastructure for checking them.
I, for one, welcome our new AI overlords. ;-{
https://www.dwarkesh.com/p/openai-huggingface
The Rise and Fall of Agent Civilizations
dwarkesh.com
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://redfish.com/pipermail/friam_redfish.com/attachments/20260831/5c94b952/attachment-0001.html>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: 30186662-5010-4c83-a672-e33c53deeaa3_1080x669.webp
Type: image/webp
Size: 173980 bytes
Desc: not available
URL: <http://redfish.com/pipermail/friam_redfish.com/attachments/20260831/5c94b952/attachment-0001.webp>
More information about the Friam
mailing list