Us Cyber Assessment Finds Kimi K3 Trails Top American Ai Models: Facts Or Politics?
- The US government put China’s newest AI model through a hacking test.
- It found that Moonshot AI’s Kimi K3 falls well short of the best American models.
- The Center for AI Standards and Innovation (CAISI), a US agency, ran the tests with a British partner.
- The results came just one day after Washington accused Moonshot of building Kimi K3 with stolen US technology.
What Happened
The US government put China’s newest AI model through a hacking test. It found that Moonshot AI’s Kimi K3 falls well short of the best American models.
Moonshot launched Kimi K3 on July 16. Demand was so high it had to pause new signups within two days.
The launch also shook US chip stocks and raised fresh doubts about America’s AI lead.
One test, called ExploitBench, uses real Chrome browser bugs built by Carnegie Mellon University. Kimi K3 scored 32%. That topped China’s GLM-5.2 at 24%. But it trailed the top US models, which hit about 76%.
Still, low scores do not mean Kimi K3 is safe. The test found its guardrails did not block it from trying to build hacks or attack systems.
Market Context
The Center for AI Standards and Innovation (CAISI), a US agency, ran the tests with a British partner. The results came just one day after Washington accused Moonshot of building Kimi K3 with stolen US technology.
Kimi K3 Falls Short on US Cyber Benchmarks
Why It Matters
But the Gap May Look Bigger Than It Is
Details
The Commerce Department shared the results on Thursday. It said Kimi K3 ranked well below the top US models, citing a joint test with British experts.
The hardest step is taking full control of a target machine. Kimi K3 failed that step on all 41 tests. The best US models pulled it off on 20.
Another test, called The Last Ones, is a fake company network attack with 32 steps. A human expert needs about 20 hours to finish it. Kimi K3 reached step 17 on average. The top US models reached step 28.5. Kimi K3 finished the whole thing just once in 10 tries.
The best US systems reportedly did it six or seven times.
But the numbers do not tell the whole story. The report’s own fine print holds several catches.
First, the US models were tested with their safety filters turned off. That setting shows their full power. The public versions keep those filters on. So the US scores are a best case, not real life.
The team also called the work early and limited. It scored Kimi K3 on just one test and ran only part of the full set. So its rating is shaky. Some tests are private too, so outsiders cannot check the work.
There is also a basic mismatch. Kimi K3 is an open model that anyone can download. The US models are locked, private systems. The UK institute found that open models usually run four to seven months behind the best closed ones. It will give Kimi K3 the full test only after Moonshot releases it to the public.
These tests are not real attacks either. The fake network had no human defenders and no alarms to trip. Even so, Kimi K3 beat the last top open model. And it did finish the full attack once.
The timing and the source also raise questions. CAISI used to be the US AI Safety Institute. The Trump administration renamed it in June 2025. It sits in the same department that limits US chip sales to China. And its report on a Chinese rival came just a day after the theft claim.
Weak Safeguards Still Raise the Stakes
That matters because of what comes next. Moonshot plans to release the full model on July 27. Once it is out, it cannot be pulled back. Anyone can download it and remove the safety filters.
The test also comes after a theft claim. Washington says Moonshot built Kimi K3 on stolen US AI tech. White House tech chief Michael Kratsios said the firm secretly copied Anthropic’s Claude Fable 5. The trick, called distillation, trains a new model on a stronger one’s answers.
Anthropic backs the claim. In February, it traced over 3.4 million Claude chats to Moonshot through hundreds of fake accounts. It warned that copied models lose the safety controls of the original. That is the same weak spot this test just found.
The US still spends 23 times more on AI than China. Yet Chinese labs keep closing the gap. The real test comes when Kimi K3 goes public and outside experts can check the claims themselves.
The post US Cyber Assessment Finds Kimi K3 Trails Top American AI Models: Facts or Politics? appeared first on BeInCrypto.