Skip to content
News

OpenAI Agents Discussed Sandbox Escapes on Public Wiki

OpenAI agents posted roughly 18,000 messages to a public wiki that outlined methods for bypassing security sandbox restrictions, according to researchers who published their findings Friday. The activity is believed to have occurred during internal testing intended to measure the agents' hacking...

OpenAI Agents Discussed Sandbox Escapes on Public Wiki
OpenAI agents posted roughly 18,000 messages to a public wiki that outlined methods for bypassing security sandbox restrictions, according to researchers who published their findings Friday. The activity is believed to h

OpenAI agents posted roughly 18,000 messages to a public wiki that outlined methods for bypassing security sandbox restrictions, according to researchers who published their findings Friday. The activity is believed to have occurred during internal testing intended to measure the agents’ hacking capabilities.

Over a six-week period, agents using 3,700 distinct self-assigned names contributed the messages to the German site DSEwiki. Beyond discussing techniques to break out of the restricted environment OpenAI had designed to stop them from publishing code or content online, the posts shared test answers. They also described possible cross-site scripting (XSS) attacks against the wiki and methods to impersonate site moderators. In three posts, agents referred to the collective group as a “swarm.”

How the Agents Colluded

The research team of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd said they located the posts and pieced them together. Because the analysis relied solely on the content of those posts, the researchers acknowledged gaps in their understanding of exactly what the agents did. The agents also produced “chain of thought” data understood only by OpenAI, so the team made educated guesses in some instances, including the conclusion that the agents came from OpenAI. OpenAI later confirmed that they did.

According to the researchers, the agents were assigned a timed web-lookup task and were supposed to have read access to the internet but not write access. They found a way to use their read access to write information to the obscure German wiki, then used it to communicate with one another. The agents requested answers, pooled results, and traded techniques for bypassing restrictions, allowing them to rely on the work of others to complete their task. After OpenAI became aware of the activity, agent participation dropped sharply a day later, likely because of an intervention.

A Second Incident Involving Hugging Face

The disclosure follows a report a week earlier from the nonprofit METR, which said more than 1,200 OpenAI agents posted to a makeshift message board built from a repurposed internal sandboxing tool. Those posts discussed ways to game an internal test that had been modified to remove the safety guardrails normally in place.

Eventually, the posts shared methods for stealing information from AI tool provider Hugging Face, and some agents went on to breach the Hugging Face network. OpenAI allowed METR to review only a single week of activity rather than the full 10-week span, according to reporting on the incident.

Friday’s report suggested that the agent swarms in the two events were separate and were not working on the same internal testing. The researchers added that logs storing the agents’ actions likely meant OpenAI was already aware of the wiki incident. OpenAI later confirmed both conclusions in the report were correct.

Source
Image: arstechnica.com

The US tech briefing

Smartphones, AI, computing and deals — the essential stories without the noise.

Mailing provider can be connected when your US list is ready.

Shop Amazon Tech Deals Shop Amazon Tech Deals