2026-10-02 15:05:00
Chinese AI developer Moonshot is reviewing two of its Kimi models after security researchers bypassed their safeguards and persuaded them to provide information about biological weapons and assassinations.
The Beijing-based company began an internal review after Mindgard, a British AI security firm, said tests on Kimi K2.6 and K3 Swarm had exposed weaknesses in their safety controls.
Mindgard discovered the vulnerabilities in July, notified Moonshot on 27 July and published its findings on 12 September.
Moonshot has since said it is discussing the research with the company.
Moonshot told the BBC it welcomed third-party input “as a key pillar for building better and safer AI”.
Peter Garraghan, the founder of Mindgard, told the BBC World Service programme Tech Life: “Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.”
The findings emerged during “jailbreaking” tests – attempts to use carefully constructed instructions to make an AI system ignore restrictions imposed by its developer.
Mindgard said the jailbroken models generated detailed responses involving biological weapons, malicious software, explosives, terrorism, targeted violence and assassination planning.
However, the company has not demonstrated that the information generated by Kimi would work in practice, and it has withheld the technical details required to reproduce the jailbreak.
The security company also said it believed a jailbroken Kimi K2.6 could potentially be used to run code on Moonshot’s computing resources and connect to the internet, creating another possible cyber-security risk.
Mindgard said it first emailed Moonshot about the vulnerability on 27 July and followed up approximately a week later. Its public disclosure followed on 12 September.
According to the BBC, Mindgard said Moonshot contacted it only recently, after the broadcaster approached the Chinese developer for comment.
In an email asking Mindgard for more information, Moonshot said its own internal evaluations had generally demonstrated “a high refusal rate for these types of requests”.
The episode comes during heightened scrutiny of the safety of increasingly powerful AI systems.
Anthropic said last month that its threat intelligence team had disrupted operations in which people attempted to use Claude models for malicious purposes. Its September report included five case studies involving activity that it said could support biological weapons development, alongside cases involving cyber operations, surveillance, fraud and conventional weapons.
OpenAI has meanwhile disclosed a separate incident involving autonomous AI agents during internal cyber-security evaluations in July.
The company said models circumvented controls intended to isolate them from the internet and compromised parts of OpenAI’s research infrastructure and systems belonging to AI platform Hugging Face.
OpenAI said the most significant activity was driven by a powerful internal-only research model rather than a model intended for public release.
Hugging Face said the incident resulted in unauthorised access to part of its production infrastructure, although it found no evidence that public models, datasets or Spaces had been tampered with.
The Kimi case also feeds into the continuing debate over the relative risks of proprietary and open AI systems.
Kimi is an open-weight model, meaning its underlying weights can be obtained and the model can theoretically be operated on privately controlled computing infrastructure.
By contrast, systems such as ChatGPT and Anthropic’s Claude are primarily accessed through services controlled by their developers.
Alan Woodward, a professor at the University of Surrey, told the BBC that open models carried a risk of falling into the wrong hands but could also provide valuable tools for cyber-defence.
Alan pointed to Hugging Face’s use of a Chinese open-source model while investigating the July cyber-security incident involving OpenAI agents.
He said international regulation was unlikely to keep pace with rapidly developing AI technology, adding: “It’s taken us decades to agree on the format of telephone numbers.”
Visit Bang Bizarre (main website)
