Disrupting a coordinated model-distillation campaign
OpenAI disrupted a coordinated campaign to extract protected reasoning from its models, attributing core activity to individuals linked to Moonshot AI and sharing findings to strengthen industry defenses.
OpenAI identified and disrupted a coordinated campaign in early July aimed at extracting protected reasoning from its models through adversarial distillation. The activity involved unauthorized use of model outputs to train or reproduce another model, violating terms of service without breaching encryption or databases. Operators manipulated model interactions to reproduce hidden reasoning in visible forms, a method not unique to OpenAI’s systems. The company shared details with industry partners via the Frontier Model Forum to bolster collective defenses against such attacks.
The campaign began on July 1, with low initial volume before spiking to 16,000 requests on July 24 and 25 from over 4,000 users using an extraction pattern. Further investigation revealed related activity across more than 15,000 users, which was fully disrupted by July 28. Independent security researchers also reported related vulnerabilities, which OpenAI confirmed and used to accelerate mitigations. The evolving nature of the activity highlighted adversarial distillation as an ongoing security challenge requiring adaptive defenses.
OpenAI attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi, though it remains unclear if all operators originated from a single actor. Adversarial distillation poses safety and national security risks by enabling the transfer of advanced capabilities without preserving original safeguards. The technique can accelerate model capability transfer at scale, particularly in dual-use domains, making it a shared industry-wide concern.
OpenAI mitigated the campaign through account enforcement, technical controls, and partner coordination, including banning fraudulent accounts and strengthening signup and infrastructure safeguards. The company closed pathways allowing replay of encrypted reasoning and added checks to detect exposed reasoning in streamed outputs. When activity occurred via third-party services, OpenAI worked with providers to disrupt involved accounts and shared findings through the Frontier Model Forum and government channels to enhance collective defenses.