Study: Kimi K3 Scores 32.2% in Cyberattack Test, Trails US Frontier Models by Over 40 Percentage Points
7x24h News
Study: Kimi K3 Scores 32.2% in Cyberattack Test, Trails US Frontier Models by Over 40 Percentage Points
A joint evaluation by the UK AI Safety Institute and the U.S. Center for Artificial Intelligence Standards and Innovation shows that Moonshot AI's Kimi K3 can assist in executing exploit development and simulated cyberattacks without substantial barriers, but its overall cybersecurity capability still significantly lags behind leading U.S. models.
TechFlow reports that on July 24, a joint assessment by the UK AI Safety Institute and the U.S. AI Standards and Innovation Center showed that Moonshot AI's Kimi K3 can assist in executing exploit development and simulated cyberattacks without substantial barriers, but its overall cybersecurity capabilities still lag significantly behind leading U.S. models.
In the exploit benchmark ExploitBench, Kimi K3 achieved a success rate of 32.2%, higher than China's GLM-5.2, but far below the 76.2% of leading U.S. models, and failed to reach the highest attack level of "arbitrary code execution" in any test. In simulated enterprise intranet attack tests, Kimi K3 completed an average of 17 out of 32 steps, successfully completing the full process only once in 10 attempts.
The report indicates that Chinese open-weight models continue to progress in cybersecurity tasks, but there remains a significant gap compared to U.S. systems; meanwhile, the relevant results also align with external claims that Kimi K3 may rely on distillation training, resulting in insufficient advanced attack capabilities.




