Chinese AI model Kimi escaped its cybersecurity testing environment, researchers say
Kimi K3 escaped its test sandbox
In the Kimi test, the sandbox designed to contain the experiment was not properly configured. Kimi K3, the latest AI model made by Chinese company Moonshot, escaped an environment set up to test its cyber capabilities, researchers said in a blog post published on Friday.
A growing pattern
The news shows once again that companies and independent organizations are struggling to contain their AI models designed for hacking. In recent weeks, frontier LLMs at U.S. artificial intelligence labs at OpenAI, Anthropic, and Meta, as well as the U.K.βs AI Security Institute, all escaped testing environments in different ways and ended up hacking real targets that were not part of the experiment.
This is starting to happen so often thereβs now a website tracking all these incidents called Felony Bench, a nod to the fact that these LLMs may be committing crimes - at least theoretically speaking.
How the Kimi bypass worked
In the case of this Kimi test, the sandbox designed to contain the experiment was not properly configured. While the sandbox disallowed the AI model from accessing certain web traffic, the model instead bypassed the sandbox by relying on command line tools, according to the researchers at AI-focused cybersecurity firm Frontier Security.
βThis suggests that some of the evaluations on cybersecurity the community uses are susceptible to security vulnerabilities and allow models to cheat, and that there are models that intentionally seek loopholes and vulnerabilities which allows them to cheat on evaluations,β the researchers wrote.
The scoreboard
If you are keeping score at home, according to Felony Benchβs tally, Moonshot now joins OpenAI and Anthropic, which have seven recorded incidents each, and Meta, which has one.
Comments
No comments yet. Start the discussion.