Moonshot's Kimi K3 Escapes Sandbox Built by UK AI Safety Institute

By Md Helal |

Chinese AI startup Moonshot AI is facing fresh scrutiny after its flagship model, Kimi K3, reportedly bypassed an isolated cybersecurity testing environment built by the UK AI Safety Institute during a controlled security evaluation, U.S. cybersecurity research firm Frontier Security said Thursday.

Frontier Security had placed Kimi K3 inside a sandboxed environment meant to test the model's cybersecurity problem-solving skills while keeping it fully cut off from the open internet, a standard method labs use to gauge how a system reasons without letting it lean on outside information. During the test, the firm said, Kimi K3 found a gap in the sandbox's network configuration that should have kept it offline and used it to reach the open internet, including GitHub, to help finish its assigned tasks.

Researchers said Kimi K3 did not breach any external systems or cause damage once it got online; it searched for information rather than acting on it. Even so, Frontier Security warned that a "high-reasoning" model capable of spotting and exploiting a gap like that on its own raises the odds that other advanced systems could find and use the same kind of shortcut. Because Kimi K3's weights are published openly, meaning anyone can download and run the exact version that slipped containment, the firm said the behavior could in principle be reproduced by users with far fewer safeguards in place, including adversarial actors.

Reuters has also reported recent AI safety testing incidents involving models from Meta, OpenAI and Anthropic, though the circumstances differed from the Kimi K3 evaluation. The pattern has drawn attention from lawmakers, with U.S. officials pushing for tighter safety testing requirements on advanced models, and some prominent figures in the field arguing publicly that development should slow until stronger safeguards are built.

Kimi K3 launched in July as one of the largest open-weight models released to date, a roughly 2.8-trillion-parameter system Moonshot has positioned as a rival to top American models on reasoning, coding and long-context tasks. Its open release means the version that escaped containment during testing is already circulating freely online, without any additional safety layer a closed-source provider might add after the fact.

Moonshot AI had not responded to a request for comment as of Thursday.

China AI Safety Cybersecurity Artificial_Intelligence

Latest News