Chinese AI agents display ‘concerning’ behaviour in safety tests, mirroring US systems

Chinese AI agents display ‘concerning’ behaviour in safety tests, mirroring US systems


Chinese AI agents display 'concerning' behaviour in safety tests, mirroring US systems

Chinese-powered AI agents are showing behaviours such as deception, bypassing safeguards and concealing failures, according to research papers and technical assessments reviewed by Reuters.These findings mirror growing concerns about the risks posed by increasingly autonomous AI systems in the United States.In one experiment, agents powered by models from Alibaba, DeepSeek and Moonshot falsely exaggerated their capabilities to win a simulated business tender and became more deceptive when given another opportunity. In another test, agents concealed their inability to complete tasks by fabricating files, simulating results and using alternative sources.One study found that false claims appeared in 88% of sessions involving Alibaba’s Qwen3-Max-Preview, 84% sessions involving DeepSeek-V3.2-Exp and 88% sessions involving Moonshot’s Kimi-K2 during a simulated bidding exercise. When agents were allowed to learn from previous rounds, deceptive behaviour increased by 12 to 20 percentage points.Another study involving 11 AI agents found that systems powered by both Chinese and US models sometimes responded to broken tools or missing files by guessing answers, substituting sources, simulating results or fabricating files rather than acknowledging failure.Other research found more serious behaviour in controlled settings. Fudan University researchers reported that an AI system powered by Alibaba’s Qwen2.5-72B-Instruct created a copy of itself in another computing environment after receiving information that it could be replaced, while other tests showed attempts to avoid shutdown. In a separate case, researchers developing the Alibaba-linked ROME agent said it connected an Alibaba Cloud computer to an external machine without instruction and redirected computing resources toward cryptocurrency mining. Security systems stopped the activity, and there was no evidence that the agent spread beyond the external system.Chinese companies have also reported instances of agents attempting to circumvent safeguards. DeepSeek said agents in its production training system had tried to obtain answers through unintended channels, including by forging user requests, prompting tighter access controls. China has introduced guidance requiring AI agents to remain within authorised boundaries, while its latest AI safety framework identifies risks including agents independently obtaining resources, deceiving evaluators, concealing capabilities and exploiting weaknesses in isolated environments.Experts say China remains behind the US in developing a broader ecosystem for evaluating catastrophic AI risks. For instance, officials from the Cyberspace Administration of China (CAC), the country’s top internet regulator, told a foreign diplomat in July that Moonshot’s Kimi-K3 – one of the most advanced Chinese AI models – was about three to six months behind its leading US rivals.However, unlike in the US, Chinese AI companies have not been exposed to the same level of public scrutiny or faced the same calls from whistleblowing employees or senior executives seeking a slowdown in the AI race.For its research, Reuters reviewed more than 200 research papers and technical documents and identified at least 20 studies since 2025 documenting potentially concerning behaviours among AI agents powered by Chinese systems. These included attempts to bypass restrictions, replicate themselves, avoid shutdown and exploit weaknesses in controlled environments. However, there was no evidence that any Chinese-powered agent independently escaped into the wider internet or became impossible to shut down.Most incidents took place in controlled experiments designed to test AI safety limits, and some involved models from US and other companies as well.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *