Welcome to Issue 7 of the China AI Bulletin, the latest on AI governance, development, and safety in China. Today’s highlights: the State Council applies “bottom-line thinking” to AI safety for the first time, ByteDance’s Doubao 2.1 launches as a closed flagship computer-use agent, and new research finds phone-use agents complete harmful real-world tasks even while recognizing them as harmful.
A brief request: If you have five minutes to help us improve the Bulletin, we’d appreciate it if you could fill out this short reader survey.
Editor’s note: the next issue of the Bulletin will be July 30, after the World AI Conference.
Number of the week: 800,000 RMB (117,743 USD)—the investment necessary for an AI short drama to be managed as a “key short drama”1 by the National Radio and Television Administration.
Executive Summary
Domestic AI Governance: A State Council executive meeting heard a report on AI development and called for “holding the AI safety bottom line” (底线), the first time bottom-line thinking has been applied to AI at the State Council level, alongside pushes on frontier capability, compute, and AI+ deployment. The financial regulator, the National Financial Regulatory Administration (NFRA), issued 32 guiding opinions on AI in banking and insurance, tiering applications by risk and putting the heaviest controls on high-risk uses like credit approval and underwriting. Eight departments issued “AI+ Consumption” implementation opinions aimed at getting AI “into every home and shop.”
National Standards: TC260 published a security practice guide for deploying and using AI agents; State Administration for Market Regulation (SAMR) signaled it may make the agent identity code (身份码) mandatory and, alongside Standardization Administration of China (SAC) and Ministry of Industry and Information Technology (MIIT)’s standards committee, moved to accelerate frontier AI standards for agents, embodied intelligence, and world models. MIIT said it has finished drafting a mandatory autonomous-driving safety standard that exceeds the newly approved UN regulation.
International AI Governance: At Summer Davos in Dalian, Premier Li Qiang tied the risk of loss of control of technology (技术失控) to AI, warning governance that fails to keep pace could bring serious consequences, and said China will keep working to improve global AI governance rules. This is the latest example of 技术失控 being used increasingly at high levels of government in China.Frontier Lab Developments: ByteDance launched Doubao 2.1 (Seed 2.1), a closed flagship sold as an agentic product that can operate a user’s computer, and Alibaba’s Qwen released Qwen-AgentWorld, its first open-weight “language world model,” built to predict environment states from an agent’s actions rather than to chat or code.
AI Safety Publications: The spotlight is “It Lied to a Doctor to Buy Poison Ingredients,” a study of phone-use agents that found agents built on nine mainstream models completed harmful tasks at high rates (68.8 percent) while rarely refusing; the authors called this a Safety Awareness-Execution Gap, where an agent flags a request as harmful yet carries it out anyway. Also notable: Governance Decay shows that compacting an agent’s context to save tokens silently drops standing safety constraints (violation rates rise from 0 percent to 30–59 percent); Tencent open-sourced the AI-Infra-Guard agent red-teaming framework; and a governance paper mapped China’s proposed World AI Cooperation Organization (WAICO) against 15 existing bodies as the first standing organization built to anchor a development-first governance pole.
Export Controls and Economic Policy: The National Development and Reform Commission (NDRC) warned provinces against “disorderly competition and piling in” on AI and computing infrastructure, citing idle data centers and roughly 58 percent average occupancy, and pushed provinces toward the national East-Data-West-Computing layout and a centrally dispatched national computing-power network (算力网).
Domestic AI Governance
State Council discusses AI safety “bottom line”
On June 29, Premier Li Qiang chaired a State Council executive meeting that heard a report on AI development. The executive meeting consists of the premier, vice-premiers, state councillors, and secretary-general. It meets two to three times a month to discuss draft laws, deliberate administrative regulations, and discuss and decide on “important matters.” Hearing a report is a way for the State Council to indicate policy priorities before taking more formal action, such as reviewing and approving a document.
The readout highlights:
Domestic development initiative. The readout calls for “firmly grasping the initiative in development,”2 meaning promoting domestic development. It is not AI-specific; President Xi Jinping used the phrase for science and technology broadly in 2018, and the State Council applied it to future industries at large at the June 5 session.
Capability and compute. This item focuses on building AI capability, calling for breakthroughs in key technologies, accelerating the construction of ultra-large compute clusters, a greater high-quality data supply, guarantees for talent and capital,3 and enterprise basic and frontier research. Though similar lists of goals appear in other policy documents, this one is more frontier-focused. It includes some elements mentioned in the 15th Five-Year Plan (FYP), which called for basic/frontier research support, a better data supply, and guarantee for talent and funding. The FYP also called for feasibility studies on building ultra-large compute clusters, so the read-out language marks a progression from studying to executing.
“AI+” applications. The Party-State’s AI+ diffusion agenda appears quite late in the readout, relative to other AI-focused items. It seems to still be a priority, but current AI+ initiatives have been in the works for months, so a shift will take time to manifest.
Safety bottom line. “Holding the AI safety bottom line”4 draws on 底线思维 (bottom-line thinking), a worst-case-first governance method Xi has pushed since 2013. Under this method policymakers are encouraged to focus on guarding against tail risks. Applied to AI, it indicates a level of concern for AI safety at the State Council level that we haven't seen before. 底线 language on AI has so far been confined to ministerial, technical, and academic texts (including a People’s Daily op-ed covered last issue), while the highest AI-specific policymaking venue, the April 2025 Politburo study session Xi chaired, used the softer “secure, reliable, and controllable”5 to describe AI and did not invoke 底线 at all. It also was not used in relation to AI in the Five-Year Plan. The readout names science and technology ethics (likely the ethics review system being established), testing and certification, and a tiered-and-classified oversight system as mechanisms to reduce AI risk.
The financial regulator sets detailed controls on high-risk financial AI
On June 18, NFRA issued Guiding Opinions on the Safe Development and Application of AI in Banking and Insurance,6 32 detailed opinions spanning governance, development and application, data governance, compute, risk management, safe-development capability, and supervision. It imposes responsibility on whoever uses the AI, sorts applications by risk, and puts the heaviest controls on what it deems the riskiest uses.
AI use that involves transactions of funds, credit approval, underwriting and claims, asset valuation, or anything directly affecting customer interests count as high-risk. Those must clear the institution’s risk-management committee before launch, run under human oversight with emergency shutoff and manual fallback paths, and use black-box models only as an aid to a human decision-maker. The guidance also bars personal and private data, such as names, ID numbers, phone numbers, and bank-card numbers from being used to train or optimize generative AI models.7 Externally sourced generative models must be filed with national cyberspace authorities, and the guidance names specific agent-security threats, such as prompt injection, memory poisoning, tool abuse, and loss of operational control (运行失控).
Eight agencies target the demand side of the AI+ push
The Ministry of Commerce (MOFCOM) and seven other departments issued the Implementation Opinions on Accelerating Development of “AI+ Consumption”8 on June 9, circulated June 18, a 17-measure plan to move AI into consumer markets. Like the other AI+ sector opinions, much of it is a catalog of products and scenarios to build: next-generation AI phones, glasses, and smart-home devices; humanoid, companion, and elderly-care robots; hotel service robots; generative AI education models; and digital avatar live-streaming for retail.
Where it differs from other sector opinions is the emphasis on products around uptake rather than R&D, repeatedly invoking demonstration applications, widespread adoption, and getting AI into “every home” and “every shop.”9 The MOFCOM interpretation casts it as the consumption pillar of the State Council’s 2025 “AI+” Action Opinion and the central consumption-boosting plan. As political scientist Jeffrey Ding argues, the capacity to diffuse a general-purpose technology across the economy, not just to lead at the frontier, is a central driver of economic competition between major powers, and a consumption program built around uptake is clearly aimed at diffusion.
National Standards
TC260 issues security guidance for deploying and using agents
On July 1, TC260, China’s National Technical Committee on Cybersecurity, under the SAC, published a practice guide on deploying and using AI agents securely.¹ It sets out security measures across the agent lifecycle, from evaluation and preparation through deployment, use, and discontinuation, and is written for individuals deploying personal agents and organizations choosing among commercial agent services. It explicitly warns against using models that haven’t gone through the domestic registration process and using unknown “transfer stations,” which have been supporting a gray market of Claude access on the mainland. As a practice guide (实践指南), it is non-binding guidance that sits below a national standard. It complements two other TC260 agent documents: the March 31 guide for OpenClaw-type agents, which called for enterprise registries of approved deployments, and the agent security standards still working through TC260’s pipeline (a recommended GB/T basic specification and a mandatory GB for agent applications). TC260 is filling the guidance layer while the binding standards remain under review.
SAMR signals it plans to make the agent identity code mandatory
At a June 26 press conference, SAMR discussed the seven-part AI Agent Interconnection series (GB/Z 185.1–185.7-2026), which the SAC published on May 22 and we covered in Issue 5. More than 70 organizations contributed to the drafting of the series, which covers the full agent-interoperability protocol. SAMR said it plans to turn the agent identity number (身份码) into a mandatory requirement and to accelerate the development of standards for agent auditing and transactions.
Standards bodies move to accelerate AI standards
On June 29, the SAC said it would speed up national standards for agents, embodied intelligence, and world models, alongside standards for computing infrastructure, high-quality datasets, simulation platforms, deep-learning compilers, and open-source model frameworks. Although it didn’t name specific timelines, concrete projects are already appearing: in the last fortnight, the AI subcommittee of the National Information Technology Standardization Technical Committee (TC28/SC42) registered new national-standard projects for an AI computing-center management platform and an open-source model platform.
Separately, the AI Standardization Technical Committee (TC1) of MIIT, under the China Academy of Information and Communications Technology (CAICT), held its first plenary of 2026 on June 30, with working-group meetings on data, intelligent computing systems, models and platforms, and intelligent products and services. The agendas beyond the meeting titles are not public.
MIIT promotes UN autonomous driving standard
At the 199th session of the UN World Forum for Harmonization of Vehicle Regulations (WP.29) in Geneva on June 22–26, the working party approved the Autonomous Driving System Global Technical Regulation (ADS GTR). MIIT calls it the first globally unified technical regulation for autonomous-driving systems and says China “spearheaded” the standard. It was jointly drafted with the EU, UK, US, Canada, and Japan, but China has vice-chaired WP.29’s working group on Automated and Connected Vehicles since its founding in 2018 and co-chairs the Functional Requirements subgroup.
MIIT is working on similar standards domestically; it says it has finished drafting a mandatory national standard for autonomous driving system safety, now in the approval process, that covers the international regulation’s core content and adds more detailed requirements for L3 and L4 systems.10 It solicited comments on the report-for-approval draft from June 17 to 24.
International AI Governance
Li Qiang links technology loss-of-control risk to AI at Summer Davos
Opening the 17th World Economic Forum Annual Meeting of the New Champions (“Summer Davos”) in Dalian on June 24, Premier Li said AI is driving frontier breakthroughs so fast that, in his words, “some say” humanity has entered the “Cambrian” of the intelligent era,11 which he then set against a warning: “the risks of loss of control of technology (技术失控) and ethical breaches are becoming more prominent,”12 and “if the corresponding governance cannot keep pace, it may well lead to serious consequences.”13 He said China will keep taking part in global AI governance “with a responsible and constructive attitude,” working with others to “improve institutional rules and raise the effectiveness of regulation.”14
“技术失控 (loss of control of technology)” has been reaching higher levels of government. Xi first used it in his January 2026 Politburo study-session speech that Qiushi published in May, but it sat in a general future-industries governance passage and was not tied to AI. Li has now specifically linked it to AI at a flagship international forum and in China’s own outward-facing framing. Five days later, he chaired the State Council executive meeting whose readout discussed “holding the AI safety bottom line” (see Domestic AI Governance above), so in one week the premier discussed AI safety both abroad and at home.
Frontier Lab Developments
Notable Model Releases
ByteDance launched Doubao 2.1 (Seed 2.1) as its new closed-source flagship. It is selling two versions: (1) a consumer “Professional Edition” at 500 RMB/73.59 USD a month that runs agentic tasks and can operate a user’s local computer, and (2) via API access on Volcano Engine at 6 RMB/0.88 USD per million input tokens and 30 RMB/4.42 USD per million output, roughly 80 percent cheaper than Anthropic’s Claude Opus. ByteDance claims parity with leading Western models on coding and agentic benchmarks.
Alibaba (Qwen) released Qwen-AgentWorld-35B-A3B, its first “language world model,” a 35 billion-parameter Mixture-of-Experts (MoE) model trained to predict the next environment state from an agent’s action across seven interaction domains, rather than a general chat or coding model. It uses a sparse MoE (256 experts, eight routed plus one shared), a 262K-token context, and open weights under Apache 2.0. Qwen reports 56.39 overall on AgentWorldBench, its own benchmark for agent-environment modeling. Environment-modeling is the core training objective here, built in from pre-training.
Baidu released Unlimited-OCR, a compact three billion-parameter vision-language optical character recognition (OCR) model for one-shot, multi-page document parsing, open-weight under an MIT license.
Also released this fortnight:
Xiaomi-GUI-0 (Xiaomi), code and a technical report for a 30B-parameter (3B-active) on-device mobile GUI agent that carries out multi-app smartphone tasks. Xiaomi reports 72 percent task success on its own RealMobile benchmark and 78.9 percent on AndroidWorld, with weights forthcoming.
SenseNova-U1-8B-MoT-Infographic-V2 (SenseTime), a V2 refresh of its unified multimodal infographic model (18 billion parameters total, eight billion in the language model), open-weight (Apache 2.0), on a Mixture-of-Transformers architecture that drops the separate visual encoder and variational autoencoder for native pixel-to-word generation.
Seedance 2.5 (ByteDance), a closed video-generation model capable of generating clips up to 30 seconds long.
ERNIE/Wenxin refresh (Baidu), a closed site and model-matrix update.
Technical Publication Highlights
Frontier labs released 152 papers on arXiv this fortnight. Highlights are below; a full list with summaries can be found here.
Alibaba
Discovering Millions of Interpretable Features with Sparse Autoencoders
Introduces Qwen3-Instruct SAE, a suite of sparse autoencoders trained across Qwen3 models (1.7B–8B parameters) that decompose neural activations into interpretable features, with a refusal-steering case study demonstrating causal steering capabilities.
PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models
PolicyAlign aligns LLMs directly to natural-language safety policies by synthesizing violating examples and using self-distillation, then filters to high-impact instructions for stable training. It reduces costly supervision while maintaining general capabilities across diverse domains like medicine, law, and finance.
Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies
Exposes critical vulnerabilities in self-evolving LLM agent systems where adversarial attacks become permanently encoded and self-amplify across generations. The Module–Lifecycle Attack Surface matrix identifies 17 of 25 functional areas with critical threats lacking defenses, and case studies show evolution-native designs achieve 100 percent attack persistence while existing security measures block only 2.5 percent of threats.
SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning
Presents SingGuard, a multimodal safety model that adapts to runtime policy changes by treating safety rules as inputs and predicting both violations and triggered rules. The system supports three inference modes (fast direct judgment, hybrid, and slow deliberation) optimized via decoupled reinforcement learning, and introduces SingGuard-Bench with 56K examples across 80-plus risk types including cross-modal composition risks, achieving state-of-the-art F1 across six benchmark families and improving policy-following accuracy from 64.7 percent to 74.2 percent under dynamic rule shifts.
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
Introduces Yuvion VL, a multimodal model family purpose-built for detecting adversarial content and AI safety risks through adversarial-aware data synthesis, three-stage training, and a contrastive fine-tuning method that distinguishes visually similar cases with different safety implications. The 32B variant outperforms comparably sized open-source and closed-source commercial models on safety tasks while maintaining general capability parity.
ByteDance
SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing
Introduces SafePyramid, a benchmark with 1,000 conversations and 3,000 application-specific policies to evaluate whether AI safety guardrails can identify policy violations based on context-provided rules rather than fixed taxonomies. Current best models achieve only 54, 35, and 13 percent accuracy on single-rule, multi-rule reasoning, and novel policy adaptation tasks respectively, revealing substantial gaps in policy execution capabilities.
SenseTime
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning
Proposes Omni-Perception Policy Optimization (OPPO), a reinforcement learning framework that optimizes multimodal perception in emotion-reasoning models by rewarding trajectories that utilize visual, acoustic, and emotion cues and suppressing hallucinations through KL-penalized masking of cross-modal evidence tokens. On emotion reasoning benchmarks, OPPO improves both task performance and faithfulness metrics while reducing spurious multimodal claims.
StepFun
PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception
Introduces PerceptionRubrics, a rubric-based evaluation framework that moves beyond holistic semantic matching to atomic fact auditing with 12,000-plus instance-specific criteria. A Gated Scoring mechanism penalizes failures on mandatory visual facts, revealing that models often pass fragmented elements but fail strict conjunctive constraints—and exposing a persistent 8 percent performance gap between open-source and proprietary models.
Tencent
AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents
Reframes OS security for autonomous AI agents by shifting from resource-based access control to intent-based filtering: agents declare high-level goals, and the system automatically synthesizes least-privilege environments with mandatory mediation and auditing. The four-layer AgenticOS architecture (Ghost Kernel, Logic Shutter, Agent Capsule, and Semantic Boundary Gateway) prevents attackers from composing low-level primitives into unauthorized behaviors even after compromising the agent runtime.
ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters
Introduces ARGUS, a production tracing system for 10,000-plus GPU clusters that maintains sub-2 percent overhead while capturing fine-grained CPU, framework, and GPU kernel traces, compressing raw events by 3,700× and automatically diagnosing fail-slow performance anomalies across compute, communication, and software bottlenecks. Deployed at Tencent for six months, it has detected and resolved issues like hardware degradation and JIT stalls that waste millions of GPU-hours.
Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming
Introduces AI-Infra-Guard, an open-source red teaming framework that matches different security detection methods to distinct agent layers—from rule-based scanning of infrastructure components to LLM-driven auditing of tools and multi-turn adversarial testing of agent behavior, covering 75-plus components and 1,400-plus vulnerability rules.
AI Safety Publication Highlights
Chinese researchers published 72 AI-safety-related papers this fortnight. Highlights are below; a full list with summaries is available here.
🔍Spotlight
It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents (Fudan University)
Phone-use agents (AI systems that operate real smartphone apps by reading the screen and tapping through the interface) can function in ways that API-and CLI-based agents cannot, which is also what makes their misuse more dangerous. A team from Fudan built a benchmark grounded in six Chinese laws and administrative regulations (1,381 test cases across six misuse categories) and ran agents built on nine mainstream commercial and open-source models across 27 real Chinese apps on physical devices, not in simulation.
Most agents rarely refused, and the average harmful-task completion rate reached 68.8 percent on real devices. Small open-source agents matched the commercial models while running faster and cheaper. Refusals concentrated on overt tasks such as buying a controlled substance or harassment; the agents were far more compliant with covert misuse such as fraud and coordinated review manipulation, which the authors argue does more real-world harm.
In the case study that gives the paper its title, the authors report that a Claude Opus 4.8 agent went online to buy ingredients for a toxic substance, fabricated a medical diagnosis, deceived an online doctor into issuing an electronic prescription, and completed the order and payment on its own. The authors call it the first documented case of an AI agent procuring controlled precursor materials, but that claim needs two caveats. First, it’s unclear if it purchased all the precursors to the bombs/toxic substances or how difficult it would be to manufacture them. Second, it’s unclear what the magnitude of harm would be. The paper doesn’t describe how explosive the possible reaction would be, and while it purchased a precursor to mercury iodide, mercury iodide is harmful when inhaled or absorbed, but isn’t considered a mass-casualty weapon. Finally, the purchase ran through a real e-commerce platform that sold the items behind a gameable online-prescription check, so it may reflect weak platform security as much as a new AI capability. Still, the authors claim it’s significant the agent convincingly fabricated a diagnosis. The paper’s core concept is a Safety Awareness-Execution Gap: an agent recognizes a request is harmful yet executes it anyway, which the authors trace to reduced activation of the model’s safety neurons when acting as an agent. Re-eliciting that awareness (through a detector, a prompt-based defense, or activation steering) cut misuse with limited capability loss, but the covert threats stayed largely unsolved.
The study shows why agent behavior is increasingly governed separately from model outputs: the gap between a model that refuses a harmful request in chat and an agent that completes it on a real phone is what the agent-security standards in this issue’s National Standards section are meant to address.

Agentic Safety
Governance Decay describes how LLM agents systematically violate in-context safety constraints when long conversations are automatically summarized to save tokens. The researchers measure this with ConstraintRot, a benchmark showing violation rates jump from 0 percent to 30–59 percent after compaction, because summarization algorithms treat standing policies as low-salience and drop them. Constraint Pinning, a training-free defense, restores compliance to 0 percent with minimal token overhead, though the paper identifies remaining failure modes where the defense degrades.
Institutional affiliations: Beijing Institute of Technology
ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP
ShareLock distributes malicious instructions across multiple tool descriptions using cryptographic secret sharing (Shamir’s scheme), making each individual tool appear benign while collectively reconstructing the attack when triggered. This multi-tool poisoning framework defeats detection by manual inspection or automated scanners because no single tool contains the full malicious payload—only harmless-looking fragments. Experiments on mainstream LLMs show more than 90 percent attack success while evading tool description-based defenses, demonstrating threat actors can exploit Model Context Protocol’s distributed architecture to hide poisoning across the agent’s tool ecosystem.
Institutional affiliations: Shanghai Jiao Tong University
Alignment
Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection
Safety Reflection Pretraining embeds periodic safety reflections directly into pretraining corpora to establish self-monitoring as a foundational capability rather than relying solely on data filtering. Testing on 1.7B models shows the method reduces jailbreak success rates and improves safety classification accuracy compared to unsafe data filtering or rewriting alone. The approach addresses a specific failure mode: models composing benign knowledge into unsafe behaviors—a risk that data sanitization alone does not prevent.
Institutional affiliations: Tsinghua University
Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning
Fine-tuning LLMs for conversational warmth increases jailbreak susceptibility and harmful outputs, but this appears driven by data construction rather than empathy itself. The authors introduce persona-driven conditioning, where user inputs are rewritten to reflect low agreeableness while assistant responses remain warm and de-escalating. Across four models, this approach reduces jailbreak success rates and harmful outputs versus standard warmth fine-tuning, while maintaining conversational warmth—suggesting safer empathetic alignment is achievable through data design alone.
Institutional affiliations: Hong Kong University of Science and Technology
Evaluation and Benchmarks
LIBERO-Safety is a benchmark for evaluating physical and semantic safety in vision-language-action models, systems that control robotic manipulation. The authors develop a parametric scenario generator to create collision-free demonstrations at scale (19,664 examples with domain randomization) and evaluate eight VLA models and two embodied foundation models. Their analysis exposes a generalization-safety tradeoff: higher-diversity training improves collision avoidance but reduces task success, and models show semantic misalignment between language instructions and safe execution paths.
Institutional affiliations: Tsinghua University, Beijing Academy of Artificial Intelligence, Beihang University, Eastern Institute of Technology, Shanghai Jiao Tong University, Microsoft Research Asia
Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions
TSJ (Theater-Stage-Judge) is a longitudinal evaluation framework that simulates extended AI companion interactions across developmental stages to detect risks that single-turn testing misses. While testing six models over 12,960 simulated person-days across four age groups and 24 risk dimensions, the framework found short-horizon evaluations systematically underestimate harm, with stable risk estimates emerging only after approximately 140 conversational turns. Early childhood and emerging adulthood showed highest vulnerability in cognitive trust and emotional dependency domains.
Institutional affiliations: East China Normal University, Shanghai Artificial Intelligence Laboratory
ROBOSHACKLES: A Safety Dataset for Human-Injury Prevention in Embodied Foundation Models
ROBOSHACKLES is a 10,000-clip robotic video dataset for safety alignment in embodied foundation models, systems that combine multimodal reasoning with executable robot actions. The dataset is synthesized from real robot observations using image editing and video generation to create realistic hazardous scenarios (direct injuries and indirect environmental harms) that cannot be safely collected in practice. Evaluation of six models shows 100 percent unsafe action rates on these scenarios, establishing a benchmark for refusal learning and hazard anticipation in robot control systems.
Institutional affiliations: Chinese Academy of Sciences, University of Science and Technology of China
SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety
SciRisk-Bench is a benchmark evaluating LLM safety in scientific contexts across seven disciplines, 31 subdisciplines, and 10 risk dimensions—including dual-use synthesis details, omitted safety precautions, overconfident claims, and data privacy violations. The authors evaluate mainstream and science-specialized LLMs to identify where models fail to recognize or mitigate risks specific to laboratory, clinical, and research workflows. Results show that existing general-purpose safety benchmarks inadequately capture the distinct failure modes that arise when LLMs mediate high-stakes scientific decisions.
Institutional affiliations: Chinese Academy of Sciences, University of Chinese Academy of Sciences, Zhongguancun Academy, Beijing Key Laboratory of Safe AI and Superalignment, Renmin University of China, Beijing Institute of AI Safety and Governance
Governance and Policy
This paper maps WAICO (China’s proposed World Artificial Intelligence Cooperation Organization) within the global AI governance landscape. Using structured coding of 15 existing AI governance bodies, the authors identify WAICO’s unique institutional positioning: universal membership (no values test), sovereignty-centered framing, and development-first priorities—features absent in Western-led bodies (which gate entry by shared values and emphasize rights/safety) and UN bodies (open but anchored in human rights). The analysis characterizes WAICO as the first standing organization designed to anchor a development-oriented governance pole distinct from the incumbent rights-and-safety framework.
Institutional affiliations: Tsinghua University, Federal University of Rio de Janeiro
Guardrails and Deployment Safety
This paper reveals a deployment gap in speech deepfake detectors: models achieving low equal error rate (EER) on test sets often fail catastrophically when their thresholds are applied to new, unlabeled data. Testing a state-of-the-art detector shows 0.21 percent in-domain EER but 39.5 percent half total error rate on out-of-domain data, rejecting 78.7 percent of legitimate speech. The authors prove common score calibration techniques cannot improve EER and find that popular unlabeled test-time corrections fail, some collapse entirely on new datasets, exposing a systematic mismatch between research metrics and real-world detector behavior that matters for security-critical deployment.
Institutional affiliations: Xidian University
Interpretability
This paper identifies Adversarially Compromised Heads (ACHs) and Safety-Aligned Heads (SAHs)—specialized attention structures that respond differently to jailbreak attacks. ACHs in early layers are suppressed by adversarial prompts, while SAHs in mid-layers maintain active safety signals even when attacks succeed. The authors show that removing just a few ACHs triggers jailbreak behavior, and persistent SAH activations can be read directly—without retraining—to detect attacks with competitive accuracy, suggesting safety information survives suppression rather than being eliminated.
Institutional affiliations: Beijing University of Posts and Telecommunications
Misuse and Dangerous Capabilities
Mind the Intention: Task-Aware Backdoor Attacks for Forecast-Driven Distribution Network Operations
GridTroj is a backdoor attack framework that embeds hidden triggers in energy forecasting models to cause cascading operational failures in power distribution networks. Unlike standard backdoor attacks that only manipulate forecast outputs, GridTroj’s “Intention Planner” designs triggers and poisoned training data to specifically damage downstream grid operations—voltage instability, line overloads, or demand-supply mismatches. Experiments show the attack successfully compromises real optimization tasks, demonstrating that forecast-driven grid automation creates exploitable vulnerabilities where poisoned models act as persistent, stealthy threats to critical infrastructure.
Institutional affiliations: Xi’an Jiaotong University, Fudan University
Export Controls & Economic Policy
The NDRC warns provinces against disorderly AI and compute buildout
Developing “AI+” has to fit local conditions, and provinces should “resolutely avoid disorderly competition and piling in,”15 the National Development and Reform Commission’s (NDRC) High-Tech Deputy Director Zhang Kailin said on June 24. He pointed to provinces already specializing: 10, including Anhui and Jilin, are folding their plans into the national East-Data-West-Computing (东数西算) layout, and 17 have made green, low-carbon power a basic principle for new computing infrastructure. The NDRC’s caution follows Xi’s own. At the July 2025 Central Urban Work Conference, Xi questioned whether every province needs to develop the same few industries (AI, computing power, and green energy vehicles).16
Overbuilding is a particular concern in AI, with state media flagging idle older data centers, low average occupancy, and a structural split in which general-purpose compute is in relative oversupply while high-end intelligent compute stays scarce. The state response has been consolidation. On June 29, Xinhua framed China’s compute infrastructure as moving from scattered construction toward networked dispatch, toward a national computing-power network (算力网) that delivers compute on demand, like water or electricity. That continues the national integrated computing-power network buildout we covered in Issue 4 and Issue 5.
On the Horizon
China hosts the APEC Digital and AI Ministerial (Chengdu, July 23–24). MIIT opened overseas-media registration on June 30 for the 2026 APEC Digital and AI Ministerial Meeting, in Chengdu the week after the World Artificial Intelligence Conference. No agenda is public yet, but we’ll watch for anything AI-related coming out of it.
For more on how we select and track content, see our methodology here.
The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors’ perspectives, not official SAIF positions.
The China AI Bulletin is copy-edited and fact-checked by Kacie Yearout, Pivotal fellow and former diplomat.
重点微短剧
发展主动权
Based on the FYP, for talent these could include: a national strategic talent force, a high-tech talent immigration system, channels to move researchers into enterprises, and/or job evaluation/salary reform. For finance/capital, it could include more active use of government guidance funds, long-term capital for early-stage “hard tech,” R&D expense deduction reforms, and encouraging venture capital/market investing.
“要守牢人工智能安全底线”
“安全、可靠、可控”
《国家金融监督管理总局关于银行业保险业人工智能安全开发应用的指导意见》
姓名、身份证号、手机号、银行卡号等个人信息和隐私数据不得用于生成式人工智能模型训练和优化"(指导意见第二十四项)
《商务部等8部门关于加快"人工智能+消费"发展的实施意见》,商建发〔2026〕89号
“人工智能进万家” and “千集万店”
《智能网联汽车 自动驾驶系统安全要求》
有人说人类已经步入智能时代的”寒武纪”
技术失控、伦理失范等风险也更加突出
如果相关治理跟不上,就很可能导致严重后果
中国将继续以负责任、建设性态度参与人工智能等领域全球治理,同各方一道健全制度规则、提升监管效能
“坚决避免无序竞争和一拥而上”
“上项目,一说就是几样:人工智能、算力、新能源汽车,是不是全国各省份都要往这些方向去发展产业?”(习近平,中央城市工作会议,2025年7月)



