In a significant escalation of the ongoing geopolitical technological rivalry, the United States intelligence and security apparatus has formally accused several leading Chinese artificial intelligence companies of orchestrating "industrial-scale" campaigns to illicitly harvest data from American frontier AI models.
In a sweeping joint cybersecurity advisory, the National Security Agency (NSA), the Cybersecurity and Infrastructure Security Agency (CISA), and the Federal Bureau of Investigation (FBI) have explicitly identified a cohort of Chinese firms—including DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI—as the primary actors in a strategy known as "model distillation." This practice involves systematically querying powerful American AI systems to extract their proprietary reasoning patterns, capabilities, and outputs, which are then used to train and refine the Chinese companies’ own competing models.
The advisory marks a major turning point in how the US government views the proliferation of AI, shifting from a policy of economic competition to one of national security defense, characterizing the unauthorized extraction of data as a direct threat to American technological supremacy.
The Mechanics of Distillation: How the "Heist" Occurs
At the heart of this controversy lies the technical process of distillation. In standard, legitimate AI development, distillation is a known technique used to create smaller, more efficient models that retain the intelligence of much larger "frontier" systems. However, the US government alleges that these Chinese firms have weaponized the process through unauthorized, high-volume automation.
According to the agencies, these companies have been extracting "billions of tokens across millions of exchanges and requests" since 2024. By relentlessly querying models like OpenAI’s GPT, Google’s Gemini, Anthropic’s Claude, and xAI’s Grok, these firms are effectively "reverse-engineering" the cognitive logic and nuanced behaviors of the world’s most advanced AI systems.
The intelligence report specifically notes that DeepSeek utilized data harvested from Claude, Gemini, GPT, and Grok to train its own R1 model—the open-source reasoning powerhouse that gained global attention in early 2025. Similarly, Moonshot AI is accused of leveraging "significant" data extractions from Anthropic’s Fable model to bolster its own Kimi K3, currently touted as one of the most sophisticated AI systems developed within China.
Chronology of Escalation: From Concerns to Formal Advisory
The current standoff is the culmination of nearly two years of rising tensions between Silicon Valley developers and Chinese labs.
- Early 2024: Reports begin to circulate within the AI research community that specific accounts, often masked behind VPNs and obfuscated proxies, are systematically stress-testing frontier models with repetitive, highly structured prompts designed to elicit detailed reasoning steps rather than simple answers.
- Late 2024: OpenAI and Microsoft report an uptick in "malicious distillation" attempts. OpenAI begins an aggressive campaign to ban accounts suspected of harvesting its model outputs to train competing systems.
- Early 2025: DeepSeek’s R1 model is released to global acclaim, quickly becoming one of the most popular AI tools on the Apple App Store. Almost simultaneously, US developers and security researchers begin flagging the "uncanny" performance similarities between R1 and top-tier American models.
- Mid-2025: Anthropic publicly accuses DeepSeek, Moonshot, and MiniMax of large-scale abuse, claiming the companies have systematically violated terms of service to improve their own internal capabilities.
- Late 2025 to Early 2026: The friction reaches a breaking point as domestic concerns over the pace of Chinese AI development lead to a formal investigation by the FBI and CISA.
- Present Day: The release of the joint NSA/CISA/FBI advisory cements the government’s position that these activities are not merely intellectual property disputes but state-tolerated industrial espionage.
Supporting Data: The Scope of the Extraction
The joint advisory provides a chilling look at the scale of the alleged operation. By mapping millions of requests, US agencies have identified a "patterns of life" for these malicious bots. The data shows that the targeted models—specifically those with advanced reasoning capabilities like Claude’s Fable (which derives from the state-of-the-art Mythos cybersecurity model)—are being subjected to "probing" queries designed to reveal the underlying weights and training biases of the model.
Furthermore, the document notes that Moonshot’s Kimi K2 model showed measurable performance jumps that correlated directly with increased query volume directed at GPT-4o. This evidence suggests that the Chinese firms are not just building models from scratch, but are effectively "stealing" the thousands of man-years and billions of dollars spent on RLHF (Reinforcement Learning from Human Feedback) by American companies.

Official Responses and Industry Impact
The tech giants involved have responded with varying levels of alarm. Anthropic, which has been the most vocal about the potential risks of its models being distilled, has called for a "global framework" to prevent the weaponization of AI research. OpenAI, while focusing on the security of its API access, has intensified its efforts to implement "watermarking" and "behavioral analysis" to detect and block non-human, systematic queries.
The reaction from the Chinese firms has been dismissive. In public statements and social media posts, representatives for companies like DeepSeek have maintained that their innovations are the result of "highly efficient engineering" and the talent of their research teams, rather than the copying of Western intellectual property. They argue that the US is attempting to stifle legitimate technological progress under the guise of "national security."
However, the US government is not waiting for a diplomatic resolution. The advisory includes a detailed list of "mitigation strategies" for American firms, including:
- Rate Limiting: Implementing stricter thresholds on how many queries a single user or IP address can make in a given timeframe.
- Behavioral Fingerprinting: Utilizing AI-based monitoring to detect "non-human" query patterns that mimic distillation efforts.
- Adversarial Training: Purposely introducing "noise" or subtle traps in model outputs that can identify when an output is being used to train another system.
Implications: A New Cold War for AI
The implications of this advisory are profound. First, it signals that the US government is moving toward a more protectionist stance on AI, potentially leading to restricted access to powerful models for users in certain jurisdictions. This could mean the end of the "open" era of AI, where frontier models are freely accessible via APIs to anyone with a credit card.
Second, the accusation underscores the extreme importance of "frontier capabilities." If, as the government suggests, the secret to the world’s most advanced AI is the data gathered from the previous generation, then whoever controls the most robust, proprietary data pipelines will inevitably win the "AI arms race."
Third, the inclusion of companies like Elon Musk’s xAI in the broader discussion—following Musk’s admission during a lawsuit that xAI had used OpenAI outputs for training—complicates the narrative. It highlights that the "distillation" problem is not strictly a foreign threat, but an industry-wide practice that has pushed the boundaries of intellectual property law.
As the US and China continue to compete for dominance in the foundational technologies of the 21st century, this advisory serves as a stark warning: the era of unchecked open-model access is closing. For American companies, the challenge is now to protect their "digital gold" while maintaining the pace of innovation. For Chinese firms, the path forward is increasingly precarious, as the US signals it will no longer tolerate the "industrial-scale" harvesting of its technological crown jewels.
Ultimately, this struggle over AI training data is about more than just software. It is about the future of global power. If American models are the architects of the new digital age, the US government is determined to ensure that those blueprints do not become the property of its primary geopolitical rivals. The battle for the "brain" of the next generation of computing has only just begun.







