HeadlinesBriefing HeadlinesBriefing.com

OpenAI Disrupts Coordinated Model-Distillation Campaign

OpenAI Blog •
×

We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. The operators manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service.

We observed operators attempt to extract protected reasoning in novel ways, including by copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content. The activity began on July 1, with high-volume spikes on July 24 and 25 consisting of 16,000 requests from over 4,000 users. Further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which we fully disrupted by July 28.

We attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. Adversarial distillation poses safety and national security risks. Extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs.

We mitigated this recent distillation campaign through a combination of account enforcement, technical controls, and partner coordination. We also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families.

Source: OpenAI Blog · Summarized by HeadlinesBriefing