Kimi K3 Trails US Frontier Models in Cyber Exploits: The Distillation Gap
A joint evaluation by the British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) has revealed a significant capability gap in offensive cyber operations between Chinese and U.S. AI models. While Moonshot AI’s Kimi K3 shows promise, it remains substantially behind leading American frontier models when tasked with complex software exploitation and network attacks.
ExploitBench: The Gap in Vulnerability Research
To measure exploit development skills, researchers utilized ExploitBench, a benchmark from Carnegie Mellon University involving 41 vulnerabilities found in Chrome’s V8 engine post-2023. The results highlighted a stark performance divide. Leading U.S. models achieved an average score of 76.2%, whereas Kimi K3 scored significantly lower at 32.2%. While Kimi K3 outperformed China’s GLM-5.2 (24.4%), it failed to reach the highest level of exploitation: Arbitrary Code Execution (ACE).
In contrast, the top-tier U.S. models successfully achieved ACE in 20 of the 41 tested tasks, granting full control over target systems—a level of sophistication Kimi K3 could not replicate.
Simulating Network Attacks via "The Last Ones"
The evaluation also tested models using "The Last Ones" (TLO), a simulation of a complex 32-step corporate network attack involving 20 hosts across four subnets. This test is designed to measure an agent's ability to navigate multi-stage intrusion paths.
Kimi K3 reached an average of 17 steps out of 32, demonstrating a middle-ground capability. While it completed the full attack path in one out of ten attempts, it lacked the consistency of U.S. models, which averaged 28.5 steps. The findings suggest that while Kimi K3 is capable of autonomously attacking weakly defended enterprise systems, it lacks the reliability of Western frontier models to execute sophisticated, long-chain operations.
The Distillation Theory: Why the Gap Exists
A critical question remains: why do Chinese models trail in cyber tasks even when they perform well on general benchmarks? The report suggests this may be due to "distillation"—the practice of using high-quality outputs from frontier models (like Anthropic's Claude) as training data.
If Moonshot AI utilized distillation to train Kimi K3, they likely missed critical cyber-offensive data. Leading U.S. models employ strict safety classifiers that block advanced offensive queries. Consequently, a dataset built from these models' public outputs would be naturally depleted of the very knowledge required for high-level exploitation. This explains how Kimi K3 can match Western models in general programming and reasoning while failing significantly in specialized offensive cyber tasks.
Growing Capabilities and Persistent Risks
Despite the current gap, CAISI's time-series analysis shows that both U.S. and Chinese models are on an upward Elo-based trajectory. The performance gap for open models has narrowed from six to ten months at the start of 2025 to just four to seven months recently. The UK AISI warns that this narrowing gap does not equal safety; the rising capabilities of open-weight models represent a "persistent and irreversible risk of misuse" in the global cybersecurity landscape.
Key Takeaways
- Cyber Performance Gap: Kimi K3 scored 32.2% on ExploitBench, significantly trailing the 76.2% average of leading U.S. models, and failed to achieve Arbitrary Code Execution (ACE).
- Network Attack Capability: In TLO simulations, Kimi K3 averaged 17 steps in a 32-step attack path, proving it can target vulnerable systems but lacks the reliability of top-tier U.S. models.
- The Role of Distillation: The deficit in cyber skills may stem from distillation practices, as training on censored public outputs from models like Claude lacks the offensive data necessary for advanced exploitation.
