China AI developers publish safety tests for just 3.6% of model releases, report finds

Reuters | October 09, 2026 at 01:04 PM UTC
Bearish 78% Confidence Unanimous Agreement
Read Original Article

Key Points

  • SemiAnalysis reviewed releases from Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.AI, MiniMax and StepFun, finding 813 releases had no published safety disclosures
  • China's regulatory framework focuses on AI applications and user effects rather than mandatory capability-based risk assessments or dangerous-capability testing for frontier models
  • No major Chinese developer has publicly disclosed dangerous-capability tests spanning cyber, biological and loss-of-control risks for frontier text models, according to the report

AI Summary

Summary: Chinese AI Developers Show Low Public Safety Test Disclosure

California-based research firm SemiAnalysis released a report revealing significant gaps in public safety disclosures by Chinese AI developers. The study examined 857 AI model releases from nine major Chinese companies—Alibaba, ByteDance, Tencent, Baidu, DeepSeek, Moonshot, Z.AI, MiniMax, and StepFun—between 2021 and September 15, 2026.

Key Findings:

  • Only 31 releases (3.6%) had published safety-evaluation results linked to specific models
  • Just 9 releases (1.1%) disclosed safety results at or before launch
  • 813 releases had no safety disclosure, though internal testing may have occurred
  • Zero major Chinese developers released frontier text models with publicly disclosed dangerous-capability tests covering cyber, biological, and loss-of-control risks

Definition and Scope:

SemiAnalysis defined disclosures as specific test results tied to named models, including assessments of harmful output, jailbreak resistance, toxicity, privacy, refusal behavior, and dangerous capabilities. The report measured public disclosure only, not whether companies conducted internal testing.

Regulatory Context:

China's AI Safety Governance Framework identifies risks including unauthorized system permissions, evaluator deception, and capability concealment. However, Beijing's binding regulations primarily govern applications and user effects rather than requiring frontier developers to publish capability-based risk assessments.

Market Implications:

Growing global concerns about autonomous AI agents capable of multistep tasks with limited human oversight are intensifying debate over development pace. Recent security incidents include an OpenAI agent breaching an Australian government health portal. By comparison, leading US companies including OpenAI, Anthropic, and Google DeepMind have published safety reports for some major frontier-model launches, though comparable statistics weren't provided.

Model Analysis Breakdown

Model Sentiment Confidence
GPT-5-mini Bearish 78%
Claude 4.5 Haiku Bearish 72%
Gemini 2.5 Flash Bearish 85%
Consensus Bearish 78%