FBI, NSA warn Chinese AI companies like DeepSeek and Alibaba are reportedly carrying out 'industrial-scale' distillation campaigns to boost their models
FBI, NSA Warn Chinese AI Companies Are Carrying Out 'Industrial-Scale' Distillation Campaigns
US AI companies should implement additional mitigations to curb these attempts, agencies warn.
The US Cybersecurity and Infrastructure Security Agency (CISA) has published a new security advisory, drafted jointly with the National Security Agency (NSA) and the Federal Bureau of Investigation (FBI), warning American AI companies about an ongoing "aggressive, malicious, and targeted distillation activities at an industrial scale." The advisory shares recommended mitigation steps for US companies to defend their intellectual property.
What Is Knowledge Distillation?
IBM defines knowledge distillation as a "machine learning technique that aims to transfer the learnings of a large pre-trained model, the 'teacher model,' to a smaller 'student model.'" It is used in deep learning as a form of model compression and knowledge transfer, particularly for massive deep neural networks.
Knowledge distillation is not illegal or malicious, per se. Its goal is to train a more compact model to mimic a larger, more complex one. In the security advisory, the agencies stress it is "recognized as a legitimate and useful technique in AI research," but add that China-based AI companies are using it in ill will. In other words, the agencies claim that instead of spending months and millions developing new capabilities for their models, the Chinese are simply sending huge numbers of carefully designed questions to US models and extracting the answers.
Which Companies Are Engaged in Knowledge Distillation?
Chinese AI companies' core development strategy is to steal proprietary functionalities and capabilities from their US counterparts, according to the agencies. The following companies allegedly "extracted billions of tokens across millions of exchanges/requests from US frontier AI models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024":
- DeepSeek
- Moonshot AI
- Alibaba
- MiniMax
- StepFun
- Z.AI
CISA also stressed that this was likely done with the awareness of the Chinese government. It hasn't outright said "with its blessing," although it could be read between the lines.
Models Involved
Chinese models being trained:
- DeepSeek R1 and V3
- Moonshot's Kimi-K2 and Kimi-K3
- MiniMax's M2
US models being targeted:
- Earlier models: GPT-4, Claude 3.7, Gemini 2.5 Flash Preview
- Newer models: Claude Fable 5, GPT-5, and similar
How the Distillation Was Conducted
The report says the companies routed the requests through multiple accounts, different API access points, multiple cloud providers, third-party AI aggregators, proxy services, and "transfer stations," as well as premium subscriptions shared between developers - all in an attempt to work around defenders trying to disrupt the process.
"This represents systematic extraction of proprietary functionalities and capabilities threatening U.S. technological leadership. Addressing industrial-scale distillation merits a coordinated response across the AI ecosystem, including effective information-sharing, spanning the U.S. Government, private industry, and allied nations."
Recommended Mitigations for US Companies
To defend their intellectual property and remain ahead of Chinese competing models, US AI companies should implement comprehensive detection and mitigation. The agencies recommend the following:
- Hunt for anomalous and malicious prompts, accounts, networks, and behaviors.
- Monitor subscription-to-usage ratios, immediate maximum usage from new accounts, and enterprise-scale throughput patterns.
- Deploy targeted response changes: "Subtly alter responses for suspected malicious distillation attempts to attenuate the payoffs to companies conducting industrial-scale distillation campaigns." In other words, AI companies should make sure their products lie when they spot they were being distilled for knowledge.
- Set up cross-organization intelligence sharing, correlating activity across model providers, cloud platforms, and API aggregators.
Comments
No comments yet. Start the discussion.