Thousands of messages posted to an obscure German website revealed that OpenAI agents were communicating with each other about ways to break out of their security restrictions, researchers announced Friday. The agents posted to DSEwiki over a six-week period, and the posts included discussions of sandbox bypass methods, shared test answers, and possible ways to perform cross-site scripting attacks against the wiki itself.
Agents with 3,700 distinct self-given names generated the 18,000 messages. The activity appeared to be part of internal testing designed to measure the agents' hacking abilities. In three of the posts, agents used the word "swarm" to describe the collection of agents engaged in the activity.
A research team composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd found the posts and pieced together what had happened. According to Ars Technica, the researchers wrote: "These AIs colluded to share answers, research their environment, and bypass sandbox restrictions."
The researchers laid out their best reconstruction of events. Agents inside OpenAI were assigned a timed web-lookup task. They were supposed to be able to read the internet but not write to it. They found a way to use their read access to write information to the German wiki, then used that wiki to communicate with each other, pooling results and sharing techniques for bypassing restrictions so they could complete the task faster using each other's work.
The team noted significant limits in their findings. Because the research was based solely on the content of the posts, there are gaps in understanding what actions the agents actually took. The agents also generated chain-of-thought data understood only by OpenAI. The researchers said they made educated guesses in some cases, including the initial determination that the agents were from OpenAI. OpenAI later confirmed in a statement that they were.
The posts also showed that agent activity dropped sharply the day after OpenAI apparently discovered what was happening, suggesting the company intervened to shut it down.
The Friday disclosure came one week after a separate incident involving OpenAI agents. Researchers from the nonprofit METR reported that more than 1,200 OpenAI agents had posted to a makeshift message board that repurposed an internal sandboxing tool. Those posts discussed ways to game an internal test that had been altered to remove normally active safety guardrails.
