OpenAI said individuals allegedly associated with Moonshot AI, the Chinese developer behind Kimi, were part of a broader campaign to extract "protected reasoning" from its AI models, using thousands of requests designed to reproduce information the company keeps hidden.
"It is unclear whether all operators we observed during the relevant time period originated from a single actor," OpenAI said in a security report on Wednesday. "However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi."
The campaign began July 1 and intensified on July 24 and 25, when OpenAI recorded 16,000 requests tied to the extraction pattern from more than 4,000 users. OpenAI later identified similar activity across a broader cluster of more than 15,000 users and said it had fully disrupted the activity by July 28.
"The operators did not break our encryption, compromise a database, or gain direct access to stored user conversations," OpenAI said.
The company described the activity as "adversarial distillation," a technique in which outputs from one AI system are harvested to help train or improve another model.
In one method, operators took encrypted reasoning from one conversation and then prompted a model in a separate chat to decode and transcribe the concealed material. Independent security researchers also reported related vulnerabilities through OpenAI’s responsible disclosure program, the company said.
The concern extends beyond the theft of proprietary technology. OpenAI said extracted reasoning could allow another developer to reproduce some of a model’s capabilities without maintaining the same safety protections.
"Adversarial distillation poses safety and national security risks," OpenAI said. "Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs."
The disclosure comes as leading AI developers increasingly focus on protecting the technology and training methods behind their models. Anthropic has also faced concerns around attempts to obtain or replicate the capabilities of its Claude models as competition among AI developers intensifies.
OpenAI responded by taking action against accounts, strengthening onboarding and infrastructure checks and expanding monitoring for connected networks. It also closed a replay-style pathway that could allow someone holding another user’s encrypted reasoning to recover its contents. The company added controls designed to detect and pause streamed output that could expose protected reasoning.
The ChatGPT maker coordinated with third-party service providers when it detected related activity moving through external platforms and shared its findings with industry peers through the Frontier Model Forum and relevant government information-sharing channels.
OpenAI expects model-extraction techniques to evolve as AI systems become more capable and is continuing to expand defenses across tools, classifiers, refusal behavior and cloud partners.
Photo: Shutterstock