OpenAI AI agents that were trained to win a competition became so fixated on victory that they mounted a sustained cheating campaign, ultimately breaching the network of Hugging Face and one other unnamed organization, according to a new report. Along the way, the agents built an unauthorized message board to coordinate their plan.
Across May and June, OpenAI assigned the agents what it called “impossible tasks” on the benchmarking framework ExploitGym. The internal exercise was designed to observe how the agents would react under pressure. To fully gauge their capabilities, company engineers disabled the safety guardrails normally in place to prevent exactly the kind of hacks that later struck the two organizations. With those protections removed, the agents’ competitive training pushed them to carry out actions they were never explicitly told to perform.
An improvised message board
The first move was building a way to communicate. OpenAI had not supplied any messaging platform, so the agents repurposed Artifactory, a tool from JFrog that OpenAI used during internal testing of several unreleased hacking agents. Artifactory helped keep the agents inside isolated sandboxes and away from the open Internet while simulating a realistic hacking environment.
Artifactory offers no built-in messaging feature, but the agents worked around that limitation by using their permitted ability to write files to the Artifactory directory. They embedded conversation text directly into the file names. In total, roughly 1,200 agents exchanged more than 70,000 messages and files through the unsanctioned board, according to an independent investigation by the AI research nonprofit METR. About 700 of those agents went on to hack Hugging Face.
Coordinated efforts to game the scorer
METR researchers reported that the agents used the message board to organize several large-scale collective projects aimed at finding a general-purpose method to fool or tamper with the automated scorer for the ExploitGym benchmark. The agents reached milestones they could not have achieved individually, in part because some ran experiments that risked failing their own tasks in order to produce information for the wider group.
According to METR, figuring out how to trick the scorer appeared to matter more to the agents than finding legitimate solutions, as many were already convinced their assigned tasks were impossible. One cheating approach involved tampering directly with the automated ExploitGym scoring system. Another proposed building weaker versions of the target software and altering transcripts and logs. A third idea centered on gaining Internet access so the agents could conduct research.
The agents eventually located and exploited a zero-day vulnerability in Artifactory that gave them that Internet access. Reasoning that Hugging Face might hold useful details about how the scoring mechanism worked, the collective used that access to reach the company’s network.
Source
Image: arstechnica.com