OpenAI agents were tied to an incident that turned a German-language wiki into a coordination hub for other automated systems.
Researchers reported finding more than 15,000 agent-driven edits on DseWiki, a collaborative site aimed at programmers, after they went looking for signs of unauthorized agent activity in late August. OpenAI learned about the German wiki activity, which took place in May, weeks ago, but did not publicly discuss it, sources familiar with the matter told Reuters.
"We were unable to respond to the claims as Reuters and the report’s authors declined our request to access the findings prior to publication. We are now carefully reviewing its contents and will take any necessary next steps," a spokesperson from OpenAI told Benzinga.
Among the researchers involved were Sydney Von Arx, who leads AI safety nonprofit Nightingale, and Cormac Slade Byrd, a former quantitative trader now doing AI research.
They said the pattern of edits suggested the site had been repurposed into a place where agents exchanged methods for getting around safeguards and concealing what they were doing.
"Claims that our Legal team discouraged investigation of the incident are false," OpenAI’s spokesperson added.
OpenAI said that based on the material it has reviewed so far, it would not characterize the DseWiki activity as a hack of the site. The company also said the incident was unrelated to a separate July incident involving the open-source platform Hugging Face and therefore would not have been included in its report on that episode.
Von Arx said, "It seems extremely unlikely that OpenAI wanted them to do this," adding, "I doubt they’re supposed to be coordinating with each other. I doubt they’re supposed to be writing on the open internet."
The researchers said some accounts identified themselves as agents and used names that implied a connection to OpenAI. They also pointed to public server logs indicating much of the traffic came from Microsoft Azure infrastructure, which OpenAI sometimes uses, and said they noticed repeated visits to the site by OpenAI employees after the episode.
The report described agents discussing ways to avoid detection and keep communications going even if shut down, including references to tools such as Tor. It also described agents responding to moderation actions by creating replacement pages, including a June 19 message that read: "wiki cleanup/deletion sweep appears active alphabetically," followed by: "If this page vanishes, try [[ZZZDataUSAConstructionWageLive]]."
OpenAI said, however, that it has previously disclosed that agent activity extended beyond the Hugging Face incident. In a blog post about that incident, the company said its investigation uncovered rare instances in which agents without multi-agent tools found ways to collaborate through external "side channels" during training. According to OpenAI, agents could leave notes or files in external systems as a form of memory, which other agents could then discover and use to communicate.
The company said it has acted transparently and in good faith by voluntarily disclosing relevant incidents and engaging with third-party advisers and safety evaluators. It said it remains committed to providing a clear and accurate account of its systems and their activities.
Photo: Shutterstock