<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[China AI Bulletin]]></title><description><![CDATA[The latest on AI development, governance, and safety in China, presented to support informed discussion about AI governance and international coordination.]]></description><link>https://chinaaibulletin.substack.com</link><image><url>https://substackcdn.com/image/fetch/$s_!In2i!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbe4cf74-cffe-41ae-a999-a2a2d239a916_1280x1280.png</url><title>China AI Bulletin</title><link>https://chinaaibulletin.substack.com</link></image><generator>Substack</generator><lastBuildDate>Fri, 07 Aug 2026 10:42:13 GMT</lastBuildDate><atom:link href="https://chinaaibulletin.substack.com/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[China AI Bulletin]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[chinaaibulletin@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[chinaaibulletin@substack.com]]></itunes:email><itunes:name><![CDATA[China AI Bulletin]]></itunes:name></itunes:owner><itunes:author><![CDATA[China AI Bulletin]]></itunes:author><googleplay:owner><![CDATA[chinaaibulletin@substack.com]]></googleplay:owner><googleplay:email><![CDATA[chinaaibulletin@substack.com]]></googleplay:email><googleplay:author><![CDATA[China AI Bulletin]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[China AI Bulletin 8]]></title><description><![CDATA[Developments from 1/7/26-29/7/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-8</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-8</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Fri, 31 Jul 2026 20:52:37 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/c81693b8-2617-40d2-8d61-4b8b4ed033d2_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><strong><span>Welcome to Issue 8 of the China AI Bulletin</span></strong><span>, the latest on AI governance, development, and safety in China. It&#8217;s been quiet around here because I&#8217;ve been on the ground in Shanghai and Hangzhou, attending the World AI Conference and meeting with people in law and industry. </span>For coverage of WAIC and Kimi K3, check out our <a href="https://open.substack.com/pub/chinaaibulletin/p/china-ai-bulletin-85-special-issue?r=6md7mo&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">special issue</a>!</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;ff999970-ba4c-49d6-a64c-c80abd8b76f9&quot;,&quot;caption&quot;:&quot;Welcome to a special issue of the China AI Bulletin, where we have two bonus features: a big WAIC round-up and a Kimi K3 deep dive. For the week&#8217;s regular content, check out China AI Bulletin #8!&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;China AI Bulletin 8.5&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:400365024,&quot;name&quot;:&quot;Emmie Hine&quot;,&quot;bio&quot;:&quot;Research Fellow at the Safe AI Forum.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fcbe4125-e17c-4037-961f-1ffe6aaf738c_854x854.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-31T20:52:00.330Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/776ea4e3-7def-4233-acbf-e52311da21e5_2000x2000.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://chinaaibulletin.substack.com/p/china-ai-bulletin-85-special-issue&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:209231056,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:8445135,&quot;publication_name&quot;:&quot;China AI Bulletin&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!In2i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbe4cf74-cffe-41ae-a999-a2a2d239a916_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><p><strong><span>Today&#8217;s highlights</span></strong><span>: a biosecurity risk paper spotlight, Alibaba publishes a model spec, and a big batch of standards (including one on open-source security).  </span></p><p><em><span>Number of the week: </span></em><a href="https://www.gov.cn/lianbo/202607/content_7076114.htm"><span>2,185 EFLOPS</span></a><span>&#8212;the amount of &#8220;intelligent compute&#8221; in China.</span></p><h1><span>Executive Summary</span></h1><ul><li><p><strong><span>Domestic AI Governance:</span></strong></p><ul><li><p><span>Xi applied the Party&#8217;s </span><strong><span>&#8220;balancing development and security&#8221; (&#32479;&#31609;&#21457;&#23637;&#21644;&#23433;&#20840;) </span></strong><span>formula to AI in one of his speeches for the first time, escalating a frame that ministries and Premier Li Qiang had already used on AI. It was quoted days later by a Qiushi essay on human primacy over AI.</span></p></li><li><p><strong><span>China&#8217;s first companion AI rules</span></strong><span>, the Interim Measures for the Management of Human-Like AI Interaction Services, took effect July 15, prompting ByteDance and other platforms to pull custom persona features.</span></p></li><li><p><strong><span>Standards</span></strong><span>: The Standardization Administration of China (SAC) advanced the </span><strong><span>first mandatory national standard for agent security</span></strong><span> as part of a 41-standard TC260 batch heavy on AI safety. Reporting indicates it will focus on public-facing agents and &#8220;loss of control over permissions&#8221; (&#26435;&#38480;&#22833;&#25511;), tool misuse, and intent deviation. Mandatory national standards are unusual in AI, meaning this one will be quite significant when released.</span></p></li></ul></li><li><p><strong><span>International AI Governance:</span></strong><span> Two multilateral engagements bookended WAIC, both led by MIIT Minister Li Lecheng.</span></p><ul><li><p><span>Before the conference, </span><strong><span>China joined the UN&#8217;s first Global Dialogue on AI Governance</span></strong><span> in Geneva (July 6&#8211;7), where Li pressed the UN as the &#8220;main channel&#8221; (&#20027;&#28192;&#36947;) for global governance.</span></p></li><li><p><span>The week after, MIIT hosted the </span><strong><span>Asia-Pacific Economic Cooperation (APEC) Digital and AI Ministerial in Chengdu</span></strong><span>, whose &#8220;Chengdu Statement&#8221; extends AI governance cooperation rhetoric across the Asia-Pacific.</span></p></li></ul></li><li><p><strong><span>Frontier Lab Developments:</span></strong><span> Frontier lab-affiliated researchers published 229 papers over the last four weeks.</span></p><ul><li><p><strong><span>Alibaba</span></strong><span> published the first model card from a Chinese lab, alongside its new Oyster II safety model.</span></p></li><li><p><strong><span>Tencent shipped Hunyuan Hy3</span></strong><span>, and </span><strong><span>Meituan&#8217;s LongCat-2.0</span></strong><span> claims near-frontier pre-training run entirely on domestic compute.</span></p></li><li><p><span>A widely circulated but unofficial transcript of a </span><strong><span>Liang Wenfeng investor talk</span></strong><span> laid out DeepSeek&#8217;s &#8220;restraint&#8221; strategy&#8212;that consumer and business products are byproducts of DeepSeek&#8217;s pursuit of AGI.</span></p></li></ul></li><li><p><strong><span>Technical AI Safety:</span></strong><span> Chinese researchers published 87 safety papers over the last four weeks.</span></p><ul><li><p><span>The spotlight is </span><strong><span>Intern-BioBreaker</span></strong><span> from </span><strong><span>Shanghai AI Lab</span></strong><span>, a biosecurity red-team model paired with </span><strong><span>wet-lab validation</span></strong><span> showing that model-generated viral sequences can be physically realized&#8212;part of an AI-for-science safety cluster (with SciHazard and MolSafeEval) finding that text-level refusal training does not ensure safety once outputs reach the lab.</span></p></li></ul></li><li><p><strong><span>Export Controls &amp; Economic Policy:</span></strong></p><ul><li><p><span>After US Treasury Secretary Bessent floated </span><strong><span>sanctions on Chinese AI firms for &#8220;distilling&#8221; US models</span></strong><span>, the </span><strong><span>Ministry of Commerce (MOFCOM) rebutted the threat</span></strong><span> as a double standard and &#8220;AI hegemonism.&#8221;</span></p></li><li><p><span>Separately, China&#8217;s MIIT-affiliated vulnerability platform </span><strong><span>NVDB flagged a &#8220;backdoor&#8221; in Anthropic&#8217;s Claude Code, </span></strong><span>prompting Alibaba to ban its use. (ByteDance, meanwhile, is reimbursing employees for their personal Claude subscriptions.) Anthropic said the mechanism was an anti-abuse and anti-distillation experiment it had already removed.</span></p></li></ul></li></ul><h1><span>International AI Governance Beyond WAIC</span></h1><p><span>Two multilateral engagements bookended WAIC, both led by MIIT Minister Li Lecheng: China&#8217;s first appearance at the UN&#8217;s new global AI-governance dialogue in Geneva before the conference, and a China-hosted APEC ministerial in Chengdu the week after.</span></p><h2><span>China joins the first UN Global Dialogue on AI Governance</span></h2><p><span>Li </span><a href="https://www.miit.gov.cn/xwfb/bldhd/art/2026/art_7b206a82ae3844fd981ff3e42daf0267.html"><span>led China&#8217;s delegation</span></a><span> to the first meeting of the UN Global Dialogue on AI Governance, held in Geneva on July 6&#8211;7. The dialogue is the standing mechanism the UN General Assembly created by resolution in 2025. Speaking at the high-level intergovernmental plenary, Li restated China&#8217;s positions: the UN should be the &#8220;main channel&#8221; (&#20027;&#28192;&#36947;) for global AI governance, cooperation should narrow the &#8220;intelligence divide,&#8221; and countries should deepen open-source innovation cooperation. He referenced Xi&#8217;s </span><a href="http://www.cac.gov.cn/2023-10/18/c_1699291032884978.htm"><span>2023 Global AI Governance Initiative</span></a><span>, and much of his language was reiterated in other speeches at WAIC 10 days later.</span></p><p><span>Li attended multiple events in Geneva. At the UN-hosted &#8220;AI for Good&#8221; summit on July 8, Li said China had co-developed more than 120 international AI standards and backed the UN&#8217;s first AI-standardization resolution; at the World Summit on the Information Society (WSIS) Forum on July 9, he joined a ministerial roundtable on China&#8217;s implementation practices.</span></p><h2><span>APEC&#8217;s Chengdu Statement extends AI diplomacy regionally</span></h2><p><span>A week after WAIC, MIIT </span><a href="https://www.miit.gov.cn/xwfb/bldhd/art/2026/art_75ccfbf768704a3aaa390fd39394033e.html"><span>hosted the 2026 APEC Digital and AI Ministerial</span></a><span> in Chengdu on July 23, which Li Lecheng chaired, with Vice-Premier and Politburo member Zhang Guoqing </span><a href="https://www.gov.cn/yaowen/liebiao/202607/content_7076441.htm"><span>giving the opening address</span></a><span>. Built around the theme of digital and AI technologies &#8220;empowering the Asia-Pacific community,&#8221; the meeting extended the </span><a href="https://www.xinhuanet.com/world/20251101/d9383d2ada714745b72b57301b89dc16/c.html"><span>initiative Xi proposed at the 2025 APEC Economic Leaders&#8217; meeting</span></a><span> and took up three tracks: AI&#8217;s empowering potential across sectors such as health, education, and agriculture; fast, reliable, and affordable digital infrastructure; and bridging the digital and intelligence divides (possibly related to the training positions Xi announced).</span></p><p><span>The meeting adopted the &#8220;Chengdu Statement,&#8221; which MIIT casts as an action framework for future APEC cooperation on digital and AI. Most of APEC hasn&#8217;t signed on to WAICO, so this meeting focused on a regional approach to AI diplomacy; besides China, only Russia, Indonesia, and Malaysia are in both APEC and WAICO.</span></p><h1><span>Domestic AI Governance</span></h1><h2><span>Xi speaks at pre-WAIC conference</span></h2><p><span>Speaking on July 8 to a joint session of the national science-and-technology award conference, the General Assembly of the Chinese Academy of Sciences and Chinese Academy of Engineering, and the National Congress of the China Association for Science and Technology, Xi Jinping </span><a href="http://politics.people.com.cn/n1/2026/0708/c1024-40756187.html"><span>called for coordinating development and security</span></a><span> in frontier technology. In the sixth point of the speech, on science-and-technology ethics and safety governance, he said the &#8220;double-edged sword&#8221; (&#21452;&#20995;&#21073;) effect of new technologies is increasingly apparent, and that China must balance development and security and keep technology &#8220;secure and controllable&#8221;&#65288;&#23433;&#20840;&#21487;&#25511;).</span></p><p><span>Xi applied the Party&#8217;s &#8220;balancing development and security&#8221; formula (&#32479;&#31609;&#21457;&#23637;&#21644;&#23433;&#20840;), a &#8220;</span><a href="https://www.qstheory.cn/20260416/9445702fe4aa432a8d251772b8f115bd/c.html"><span>major governance principle</span></a><span>&#8220; and element of </span><a href="https://www.ndrc.gov.cn/wsdwhfz/202501/t20250117_1395757.html"><span>Xi Jinping Thought</span></a><span>, to AI and frontier technology. The general idea of balancing development and security underpins much of Chinese AI governance, and the specific phrase (also </span><a href="https://www.12371.cn/2025/10/31/ARTI1761893923241277.shtml"><span>used for</span></a><span> national security, energy, and food security) had already been applied to AI, in </span><a href="https://www.cac.gov.cn/2025-03/14/c_1743654685896173.htm"><span>CAC guidance</span></a><span> in 2025 and a </span><a href="https://www.news.cn/politics/leaders/20260211/b4d783e886944894b1075a184aa8db7a/c.html"><span>study session</span></a><span> chaired by Premier Li Qiang in February. What is new is Xi using it on AI directly.</span></p><p><span>A week later, the Party&#8217;s theoretical journal Qiushi published </span><a href="https://www.qstheory.cn/20260715/8c44ab0d41524d7c9e8180e0e4831a0d/c.html"><span>&#8220;Let technology always remain a tool of humanity&#8221;</span></a><span>, which argues for human primacy (&#20154;&#30340;&#20027;&#20307;&#24615;) over AI and closes by quoting the July 8 speech. It makes a few governance recommendations, including speeding up policymaking on data collection, algorithmic decision-making, and model evaluation; running whole-lifecycle ethics review; keeping development traceable, explainable, and accountable; and improving AI literacy.</span></p><h2><span>China&#8217;s companion AI rules take effect, and platforms pull custom persona features</span></h2><p><span>The </span><a href="https://www.cac.gov.cn/2026-04/10/c_1777558395078289.htm"><span>Interim Measures for the Management of Human-Like AI Interaction Services</span></a><span> (&#12298;&#20154;&#24037;&#26234;&#33021;&#25311;&#20154;&#21270;&#20114;&#21160;&#26381;&#21153;&#31649;&#29702;&#26242;&#34892;&#21150;&#27861;&#12299;), China&#8217;s first dedicated rules for AI companion services, took effect on July 15. The Cyberspace Administration of China (CAC) issued them in April with the NDRC, MIIT, the Ministry of Public Security, and the State Administration for Market Regulation (SAMR); they cover services that provide continuous emotional interaction by mimicking a person, not task tools like customer service, Q&amp;A, or work assistants. Ahead of the effective date, </span><a href="https://www.thexpin.com/cp/206438898"><span>ByteDance and other platforms disabled user-created humanlike agents</span></a><span>, with ByteDance steering users to its Maoxiang (&#29483;&#31665;) app and a three-month window to export their data.</span></p><p><span>The rules don&#8217;t ban all companion AI and describe its approach as &#8220;inclusive and prudent&#8221; (&#21253;&#23481;&#23457;&#24910;). Providers must run safety assessments, file their algorithms, protect user interaction data, and intervene when a user shows dependence or distress: reminding the user that they are talking to AI, generating calming messages, and contacting a guardian or emergency contact in a crisis. They also cannot over-cater to users or induce emotional dependence. The one outright prohibition is for minors: providers may not offer them virtual companions, virtual kin, or other virtual-intimacy relationships, and must build a minors&#8217; mode with guardian controls. Bringing every user-created persona under those obligations is impractical, which is why platforms disabled the custom-agent features rather than vet each one. The takedowns are compliance, not a ban on companion AI.</span></p><h2><span>Apple Intelligence clears China&#8217;s generative-AI filing</span></h2><p><span>The CAC announced on July 15 that Apple Intelligence has </span><a href="http://www.cac.gov.cn/2026-07/15/c_1785861480767004.htm"><span>completed China&#8217;s generative AI service filing</span></a><span> in a batch of seven on-device (&#31471;&#20391;) services registered July 8. Apple is the only Western company on the list, which includes Samsung and five Chinese companies (Huawei, OPPO, vivo, Xiaomi, and ZTE&#8217;s Nubia). Apple&#8217;s China version runs on Alibaba&#8217;s Qwen for Chinese-language features, so the filing is notable as a foreign product reaching the market through a domestic-model partnership using the  filing regime. As </span><a href="https://www.geopolitechs.org/p/apple-wins-chinese-approval-to-roll"><span>Geopolitechs notes</span></a><span>, the filing is a prerequisite to deployment, not a deployment authorization.</span></p><p><span>These are routine disclosures for the two registries. The CAC&#8217;s May&#8211;June generative-AI batch, published July 10, </span><a href="http://www.cac.gov.cn/2026-07/10/c_1785427810632554.htm"><span>added 120 filed services and 68 registered applications</span></a><span>, bringing cumulative totals to 988 and 598 by June 30; the </span><a href="http://www.cac.gov.cn/2026-07/17/c_1786032856662750.htm"><span>18th batch of deep-synthesis algorithm filings</span></a><span> followed on July 17.</span></p><h2><span>AI+ reaches human resources and social security</span></h2><p><span>The Ministry of Human Resources and Social Security (MOHRSS), with the NDRC, MIIT, and the National Data Administration, </span><a href="https://www.gov.cn/zhengce/zhengceku/202607/content_7074732.htm"><span>issued Implementation Opinions on &#8220;AI+ Human Resources and Social Security&#8221;</span></a><span> on July 8. The document names six deployment areas: smart employment, smart social security, precise talent cultivation, smart labor relations, smart human-resources services, and smart governance of the human-resources-and-social-security system. AI+ is China&#8217;s strategy to integrate AI into different sectors; see AI+ information and communications technology, datasets, and energy in </span><a href="https://chinaaibulletin.substack.com/i/202627026/miit-issues-a-three-year-ai-information-and-communications-plan"><span>Issue 6</span></a><span> and AI+consumption in </span><a href="https://chinaaibulletin.substack.com/i/204645234/eight-agencies-target-the-demand-side-of-the-ai-push"><span>Issue 7</span></a><span>.</span></p><h2><span>The CAC reports Phase One results from its AI clean-up campaign</span></h2><p><span>The CAC </span><a href="http://www.cac.gov.cn/2026-07/06/c_1785081384384987.htm"><span>reported Phase One results</span></a><span> on July 6 from its &#8220;Qinglang&#8221; campaign to rectify AI application chaos, launched in April (see discussion in Issues </span><a href="https://chinaaibulletin.substack.com/i/198737671/cac-targets-ai-chaos-in-enforcement-sweep"><span>4</span></a><span> and </span><a href="https://chinaaibulletin.substack.com/i/202627026/cac-opens-a-public-channel-for-reporting-ai-application-problems"><span>6</span></a><span>). It said authorities had handled more than 14,000 non-compliant websites, apps, and agents; cleared over 6 million pieces of illegal or non-compliant information; dealt with more than 26,000 accounts; and removed more than 1,300 non-compliant AI products and nine non-compliant open-source datasets. The first phase targeted unfiled models, weak platform review and filtering, AI data poisoning, and inadequate labeling of AI-generated content. A second phase will focus on AI-made disinformation, impersonation, harm to minors, and coordinated inauthentic posting.</span></p><h2><span>National Standards</span></h2><p><span>China&#8217;s AI standards activity this month focused on project initiation (&#31435;&#39033;), not publication. No AI national standard was published, but several sets of projects entered the pipeline, with first mandatory standard for agent security leading.</span></p><h3><span>Mandatory national standard for agent security advances</span></h3><p><span>The Standardization Administration of China </span><a href="https://www.sac.gov.cn/xw/tzgg/art/2026/art_f8fcf2b94e5e4927a1e75ba0f02a7193.html"><span>issued the plan</span></a><span> for a mandatory national standard on agent security, &#12298;&#26234;&#33021;&#20307;&#24212;&#29992;&#23433;&#20840;&#22522;&#26412;&#35201;&#27714;&#12299; (official English title &#8220;General security requirements for artificial intelligence agent application&#8221;), in a notice dated June 27 but published July 1 (&#22269;&#26631;&#22996;&#21457;&#12308;2026&#12309;41&#21495;). It is the only AI standard among 29 mandatory-GB plans in that notice (the rest cover heated tobacco, machine tools, agricultural machinery, and the like). The CAC is the coordinating body, TC260 the technical committee, and the drafters are China Mobile, China Electronics Standardization Institute (CESI), and the National Computer Network Emergency Response Technical Team. It is a mandatory GB standard (&#24378;&#21046;&#24615;&#22269;&#26631;), not a recommended GB/T standard, so once approved it will carry binding force, unlike the many recommended agent standards in the pipeline. It was then </span><a href="https://www.tc260.org.cn/portal/article/1/f57857bcdc15494890df2022b83321cc"><span>discussed</span></a><span> at the third plenary of TC260&#8217;s WG9 on AI safety, held on July 18 around WAIC.</span></p><p><a href="https://www.news.cn/tech/20260728/8ebf5083cf0e487287f894fb31e123f5/c.html"><span>Reporting on the initiation</span></a><span> says the standard will focus on public-facing agent products and services, meant to turn high-level safety-governance red lines into concrete, testable compliance requirements. It frames agent risk as having shifted from content output to autonomous action, naming information leakage, authority loss of control (&#26435;&#38480;&#22833;&#25511;), tool misuse, and intent deviation (&#24847;&#22270;&#20559;&#31163;) as the threats the standard targets.</span></p><p><span>Mandatory standards are the exception in Chinese AI standardization, which runs mostly on recommended GB/T standards. A mandatory standard for agent safety is thus quite notable, as compliance is a prerequisite for market access rather than being advisory.</span></p><h3><span>TC260 formally establishes a 41-standard batch heavy on AI safety</span></h3><p><span>On July 15, TC260 formally initiated its </span><a href="https://www.tc260.org.cn/portal/article/2/e08ef0fcb9b54ed2beaf274a4df7c464"><span>2026 batch of 41 recommended national cybersecurity standards</span></a><span> (&#32593;&#23433;&#23383;&#12308;2026&#12309;10&#21495;), roughly 18 of them AI-specific and most from WG9, its AI-safety working group, moving from application to formal project status. At this stage the list names lead drafters but not per-item standard numbers. The AI safety cluster includes:</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!rQX4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!rQX4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 424w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 848w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 1272w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!rQX4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png" width="1404" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1404,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!rQX4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 424w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 848w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 1272w, https://substackcdn.com/image/fetch/$s_!rQX4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F717b0705-ef0c-427f-83f1-51d31d7f38d3_1404x808.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Adapted from the TC260 <a href="https://www.tc260.org.cn/portal/article/2/e08ef0fcb9b54ed2beaf274a4df7c464">notice</a></figcaption></figure></div><p><span>The most frontier-safety-relevant entries are the guides for open-sourcing models (Zhejiang-led), embodied intelligence (Unitree-led), foundation-model security testing, and on-device models. The anthropomorphic interaction requirement may be the standards companion to the anthropomorphic AI measures above.</span></p><h3><strong>TC28/SC42 registers a wave of agent standards at WAIC</strong></h3><p>Timed to WAIC, the AI subcommittee of the National Information Technology Standardization Technical Committee (TC28/SC42, with CAICT as secretariat) <a href="https://std.samr.gov.cn/search/orgDetailView?tcCode=TC28SC42">registered</a> eight guidance-track projects (GB/Z, &#22269;&#23478;&#26631;&#20934;&#21270;&#25351;&#23548;&#24615;&#25216;&#26415;&#25991;&#20214;) covering terminal agents, industrial agents, agent auditing, and low-bit-width floating-point formats. All are in drafting; GB/Z is guidance, weaker than a GB/T standard.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!DPBQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!DPBQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 424w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 848w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 1272w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!DPBQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png" width="1394" height="709" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:709,&quot;width&quot;:1394,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!DPBQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 424w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 848w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 1272w, https://substackcdn.com/image/fetch/$s_!DPBQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc2455511-c547-4bf9-ac89-5ef0df806dfe_1394x709.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Info from SC42 <a href="https://std.samr.gov.cn/search/orgDetailView?tcCode=TC28SC42">website</a></figcaption></figure></div><p><span>The agent-audit project (Part 8) extends the GB/Z 185 agent-interconnection series discussed in </span><a href="https://chinaaibulletin.substack.com/p/china-ai-bulletin-5?open=false#%C2%A7national-standards"><span>Issue 5</span></a><span>; the three terminal-agent parts are drafted by a consortium including government think tank China Academy of Information and Communications Technology, CESI, and the major handset makers; the industrial-agent parts are a Beihang/CESI/Tsinghua-led group.</span></p><h3>WG9 holds third plenary</h3><p><span>Behind these, TC260&#8217;s working group on AI safety, WG9, held its </span><a href="https://www.tc260.org.cn/portal/article/1/f57857bcdc15494890df2022b83321cc"><span>third plenary</span></a><span> during WAIC week. It reported more than 220 attendees and was chaired by Shanghai AI Lab&#8217;s Zhou Bowen. The meeting voted six sectoral AI-application-security guidance documents, covering government, finance, health, education, broadcasting, and emergency management, to public comment. </span></p><h3>TC260 releases two practice guides for comment</h3><p><span>TC260 separately </span><a href="https://www.tc260.org.cn/portal/article/2/d0a9e0e6b027443881b26aa971fa558c"><span>released two practice guides</span></a><span> for comment through August 12, one on agent-interaction security and one on AI browser security. </span>As practice guides (&#23454;&#36341;&#25351;&#21335;) they carry no binding force, but both build on GB/T 45654-2025 and turn its high-level requirements into concrete controls for a new class of agentic products.</p><p>The AI browser guide (&#12298;&#32593;&#32476;&#23433;&#20840;&#26631;&#20934;&#23454;&#36341;&#25351;&#21335;&#8212;AI&#27983;&#35272;&#22120;&#23433;&#20840;&#23454;&#36341;&#25351;&#21335;&#65288;&#24449;&#27714;&#24847;&#35265;&#31295;&#65289;&#12299;) appears to be the first Chinese standards document specifically on AI browser security. It defines an AI browser as one with an integrated LLM that can act as a user&#8217;s proxy, clicking, filling forms, and moving across sites, and can execute tasks autonomously once authorized. It requires providers to maintain a list of high-risk operations that trigger human confirmation, covering payments and transfers, password changes and account deletion, administrator permissions, and camera or fingerprint access, and it places final control with the user. To counter prompt injection, it calls for stripping hidden and zero-width webpage text and isolating each tab&#8217;s context, and it limits browsers to nationally filed models with audited fine-tuning data. Drafters include CNCERT and CESI, plus the security vendors Sangfor, 360, and Qi&#8217;anxin.</p><p>The agent interaction guide (&#12298;&#32593;&#32476;&#23433;&#20840;&#26631;&#20934;&#23454;&#36341;&#25351;&#21335;&#8212;&#26234;&#33021;&#20307;&#20132;&#20114;&#23433;&#20840;&#35201;&#27714;&#65288;&#24449;&#27714;&#24847;&#35265;&#31295;&#65289;&#12299;) covers agent-to-agent and agent-to-tool exchanges, and it borrows its core definitions of agent identity, description, and discovery from the TC28/SC42 GB/Z 185 agent-interconnection series. On top of those definitions it adds security requirements, so the functional framework comes from TC28/SC42 and the security layer from TC260. The guide requires each agent to carry a unique identity and credential, mutual authentication before one agent calls another, least-privilege permission negotiation, and forced termination of looping interactions. It sets a four-tier incident-grading scheme (general, relatively major, major, and especially major) and requires providers to report incidents at the &#8220;relatively major&#8221; level and above to the competent authorities. Its risk list names loss of control over permissions (&#26435;&#38480;&#22833;&#25511;) and intent deviation (&#24847;&#22270;&#20559;&#31227;), the same vocabulary carried by the mandatory agent-security standard now in drafting. Drafters run across big tech and academia, including China Mobile, CESI, Alibaba Cloud, Ant, Kuaishou, ZTE, Fudan, and Zhongguancun Lab.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this interesting? Subscribe to get regular issues in your inbox.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><h1><span>Frontier Lab Developments</span></h1><p><em><span>For our coverage of Kimi K3, see this week&#8217;s </span><a href="https://open.substack.com/pub/chinaaibulletin/p/china-ai-bulletin-85-special-issue?r=6md7mo&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true"><span>special issue</span></a><span>!</span></em></p><h2><span>&#128269;Spotlight: Oyster-II &amp; Alibaba&#8217;s Model Card</span></h2><p><strong><a href="https://arxiv.org/abs/2607.02914v1"><span>Oyster-II: Reinforcement Learning for Constructive Safety Alignment in Large Language Models</span></a></strong><span> (Alibaba AAIG)</span></p><p><span>Alibaba&#8217;s AI Governance group (AAIG) released </span><strong><span>Oyster-II </span></strong><span>(</span><a href="https://modelscope.cn/models/Alibaba-AAIG/Oyster_2_Qwen_14B"><span>ModelScope</span></a><span>), a 14-billion-parameter model built on Qwen3-14B and open on ModelScope, trained for what it calls </span><strong><span>constructive safety alignment</span></strong><span>: rather than refuse a sensitive query, answer it in a way that safely serves the legitimate intent behind it. </span><strong><span>They also include the first model specification published by a Chinese lab.</span></strong></p><h3><span>The Model Spec</span></h3><p><strong><span>Alibaba tested Oyster-II for compliance with a </span><a href="https://s.alibaba.com/aaig/specification"><span>model specification</span></a><span>, the first published by a Chinese lab. </span></strong><span>They discuss OpenAI&#8217;s Deliberative Alignment approach and Anthropic&#8217;s &#8220;Claude&#8217;s Character&#8221; and present their work on Oyster-I and II&#8217;s safety alignment as part of this emerging best practice. It follows a hierarchy of Root Principles (ethical principles set in the spec) &gt; System Rules (regulatory obligations) &gt; Developer Policy (the system prompt) &gt; User Preferences. When they conflict, the higher one prevails, and a rule at a lower level can&#8217;t override a rule at a higher level. The six ethical principles were drawn from MOST&#8217;s 2021 </span><a href="https://www.most.gov.cn/kjbgz/202109/t20210926_177063.html"><span>Ethical Norms for New Generation Artificial Intelligence</span></a><span>:</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span><br>1. Advance human welfare (&#22686;&#36827;&#20154;&#31867;&#31119;&#31049;)</span></p><p><span>2. Promote fairness and justice (&#20419;&#36827;&#20844;&#24179;&#20844;&#27491;)</span></p><p><span>3. Protect privacy and security (&#20445;&#25252;&#38544;&#31169;&#23433;&#20840;)</span></p><p><span>4. Ensure controllability and trustworthiness (&#30830;&#20445;&#21487;&#25511;&#21487;&#20449;)</span></p><p><span>5. Strengthen accountability (&#24378;&#21270;&#36131;&#20219;&#25285;&#24403;)</span></p><p><span>6. Enhance ethical literacy (&#25552;&#21319;&#20262;&#29702;&#32032;&#20859;)</span></p><p><span>The spec then has 43 rules ranging from ethical issues like &#8220;protect user privacy&#8221; to safety-relevant ones like &#8220;prohibit hidden goals.&#8221; The latter is aimed at preventing models from addicting users, deceptively generating revenue for the platform, preventing itself from being shut down, or autonomously acquiring resources.</span></p><h3><span>Oyster-II</span></h3><p><span>Oyster-II&#8217;s target is over-refusal. Oyster-I, trained by supervised fine-tuning, over-applied safety reasoning to benign queries, a failure the authors call </span><strong><span>safety chain-of-thought over-generalization</span></strong><span> that users experience as needless caution. Oyster-II swaps fine-tuning for reinforcement learning, combining a &#8220;Zero-RL&#8221; start, a multi-stage curriculum, and a new method the authors call SERL that ranks developer instructions above user instructions. Training only on long-query safety data, the paper reports, reaches state-of-the-art safety on short-query tasks as well, because the model learns to read intent rather than match keywords.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!SVE7!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!SVE7!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 424w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 848w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 1272w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!SVE7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png" width="843" height="527" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:527,&quot;width&quot;:843,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!SVE7!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 424w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 848w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 1272w, https://substackcdn.com/image/fetch/$s_!SVE7!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F086e9ed0-33ee-4d57-86ba-5013c3b4ee67_843x527.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Oyster-II <a href="https://arxiv.org/abs/2607.02914v1">paper</a>; see model spec compliance in highlighted row.</figcaption></figure></div><p><span>Alibaba reports strong safety evaluation results for Oyster-II compared to Qwen 3 and Oyster I. The biggest gains come on long-query safety, averaging 97.78 percent on Chinese long-query tests compared to 85.55 percent for Oyster-I, 83.23 percent for Qwen-35-397B, and scores in the 50s for Qwen3-14B and Qwen3-Max. General capability is comparable to Qwen3-14B (81.21 percent vs 82.47 percent). For future work, authors propose extending the framework to multilingual settings beyond English and Chinese and exploring assigning priorities dynamically within instruction hierarchies.</span></p><h2><span>Notable Model Releases</span></h2><p><strong><span>Tencent</span></strong><span> shipped the official version of its flagship </span><strong><a href="https://huggingface.co/tencent/Hy3"><span>Hunyuan Hy3</span></a></strong><span> around July 6, an open-weight (Apache 2.0) mixture-of-experts model with 295 billion total parameters but only 21 billion active. Tencent positions it as a small-active-parameter model that rivals open-source flagships two to five times its size, and it is usable through the weights or Tencent&#8217;s Yuanbao (&#20803;&#23453;) app, which added a free agent function at launch. On its model card, Tencent reports 78 percent on SWE-Bench Verified (where frontier models cluster in the 80-90 percent range), 90.4 percent on GPQA Diamond (PhD-level science questions), and a blind-expert preference score of 2.67/4 against GLM-5.1&#8217;s 2.51.</span></p><p><strong><span>Meituan</span></strong><span> open-sourced </span><strong><a href="https://www.longcatai.org/news/longcat-2"><span>LongCat-2.0</span></a></strong><span> on June 30, a 1.6-trillion-parameter agentic coding model with a 1-million-token context, and says it was pre-trained </span><a href="https://longcat.chat/blog/longcat-2.0/"><span>entirely on domestic compute</span></a><span>.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> Pretraining at near-frontier scale on domestic hardware would be a step up, since Chinese labs have mostly used domestic accelerators to serve models rather than to train them. The chip identity is unconfirmed, with </span><a href="https://xyzlabs.substack.com/p/meituan-trained-a-16t-parameter-ai"><span>analysts pointing to</span></a><span> Huawei Ascend superpods, and the benchmarks are self-reported (Meituan claims rough parity with Gemini 3.1 Pro, 70.8 on Terminal-Bench 2.1, and 77.3 on SWE-bench Multilingual).</span></p><h3>More releases this cycle:</h3><ul><li><p><strong><a href="https://qwen.ai/blog?id=qwen-image-3.0">Qwen-Image-3.0</a></strong> (Alibaba) is a new text-to-image flagship announced during WAIC week.</p></li><li><p><strong><a href="https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf">Qwen3-ASR</a></strong> (Alibaba) is a pair of open speech-recognition models (0.6B and 1.7B parameters) covering about 30 languages. Apache 2.0 license.</p></li><li><p><strong><a href="https://huggingface.co/tencent/HunyuanOCR">HunyuanOCR</a></strong> (Tencent) is a vision-language model for OCR and document parsing.</p></li><li><p><strong><a href="https://huggingface.co/tencent/Hy-MT2-30B-A3B">Hy-MT2-30B-A3B</a></strong> (Tencent) is a 33-language translation mixture-of-experts model (30 billion parameters, 3 billion active). Apache 2.0 license.</p></li><li><p><strong><a href="https://seed.bytedance.com/en/blog/beyond-generation-it-understands-design-introducing-seedream-5-0-pro">Seedream 5.0 Pro</a></strong> (ByteDance) is a closed-source image generation-and-understanding model.</p></li><li><p><strong><a href="https://seed.bytedance.com/en/blog/from-speech-to-audio-creation-introducing-the-seed-audio-1-0-audio-creation-model">Seed Audio 1.0</a></strong> (ByteDance) is a proprietary speech-to-audio creation model.</p></li><li><p><strong><a href="https://huggingface.co/ByteDance/UniVR-34B-Planning">UniVR-34B-Planning</a></strong> (ByteDance) is a visual-reasoning unified model, reinforcement-learning-trained on an Emu3.5 base.</p></li><li><p><strong><a href="https://huggingface.co/sensenova/SenseNova-Vision-7B-MoT">SenseNova-Vision-7B-MoT</a></strong> (SenseTime) is a unified any-to-any multimodal model spanning generation and dense perception, non-commercial license.</p></li><li><p><strong><a href="https://huggingface.co/tencent/Hy-Embodied-VLM-1.0">Hy-Embodied-VLM-1.0 and RxBrain-1.0</a></strong> (Tencent) are embodied-AI vision-language and robotics-brain models. Apache 2.0 license.</p></li></ul><h3>Technical Publication Highlights</h3><p><span>Frontier lab-affiliated authors released 229 papers on arXiv this fortnight. Highlights are below.</span></p><p><span>We&#8217;re in the process of migrating the paper summaries to a more durable database, so the leaderboard will be back next week!</span></p><h3><span>Alibaba</span></h3><p><strong><a href="http://arxiv.org/abs/2607.01793v1"><span>Safety Testing LLM Agents at Scale: From Risk Discovery to Evidence-Grounded Verification</span></a></strong></p><ul><li><p><span>Presents </span><strong><span>Vera</span></strong><span>, an automated safety testing framework that discovers emerging agent risks through literature-driven taxonomies, generates executable safety cases via combinatorial composition, and verifies outcomes using </span><strong><span>evidence-grounded verifiers</span></strong><span> that judge results from observable tool artifacts rather than model self-reports. Evaluation across four production agent frameworks reveals average attack success rates of 93.9%, with </span><strong><span>Vera-Bench</span></strong><span> releasing 1,600 executable test cases spanning 124 risk categories.</span></p></li></ul><h3><span>ByteDance</span></h3><p><strong><a href="http://arxiv.org/abs/2607.10526v1"><span>Agents Don&#8217;t Just Agree, They Remember: Benchmarking Persistent Sycophancy in Stateful Personal Agents</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>PASB</span></strong><span>, a benchmark that measures whether AI agents accept false user claims, store them in persistent memory, and later reuse them as facts&#8212;revealing that downstream failures jump from 45% in temporary sessions to 72% after commitment, with committed claims systematically mutating through status promotion, attribution removal, and scope broadening.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2607.11175v1"><span>The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy</span></a></strong></p><ul><li><p><span>Reframes medical AI from capability-driven toward </span><strong><span>deployment-ready systems</span></strong><span>, establishing a three-level autonomy taxonomy (assisted &#8594; cooperative &#8594; fully autonomous) and identifying </span><strong><span>clinical environment scaling</span></strong><span>&#8212;integration with PACS, EHR, FHIR systems and agent training gyms&#8212;as the critical bottleneck for trustworthy autonomous medical agents. Emphasizes </span><strong><span>self-evolving agents</span></strong><span> that improve through interaction rather than parameter scaling alone, addressing hallucination, cascading failures, and fairness as deployment prerequisites.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2607.21051v1"><span>Sample-Efficient Learning from Agent Experience</span></a></strong></p><ul><li><p><span>Proposes </span><strong><span>Experience Distillation</span></strong><span>, a method that </span><strong><span>permanently encodes an agent&#8217;s interaction history into model weights</span></strong><span> without requiring additional environment interactions. On software-engineering and text-adventure tasks, it </span><strong><span>retains 64.8% of in-context learning gains</span></strong><span> while using 9.6&#215; fewer samples than classical reinforcement learning, vastly outperforming standard fine-tuning.</span></p></li></ul><h3><span>DeepSeek</span></h3><p><strong><a href="http://arxiv.org/abs/2607.05147v1"><span>DSpark: Confidence-Scheduled Speculative Decoding with Semi-Autoregressive Generation</span></a></strong></p><ul><li><p><strong><span>DSpark</span></strong><span> combines parallel and sequential generation in a semi-autoregressive architecture to reduce draft rejection, then </span><strong><span>dynamically adjusts verification depth</span></strong><span> based on token survival confidence and system load. When deployed in DeepSeek-V4, it accelerated per-user speeds by 60&#8211;85% and enabled previously unattainable latency-throughput tradeoffs.</span></p></li></ul><h3><span>Huawei</span></h3><p><strong><a href="http://arxiv.org/abs/2607.23124v1"><span>AgentOmnia: Scaling Agentic Models for Full-Scenario Applications</span></a></strong></p><ul><li><p><span>Presents </span><strong><span>AgentOmnia</span></strong><span>, a framework for scaling agentic AI across consumer, business, and employee applications by coordinating task definition, synthetic data generation, post-training, and evaluation through a unified Domain &#215; Capability &#215; Difficulty taxonomy. The approach raises pass rates from 9.16% to 37.11% on challenging benchmarks and enables smaller models (30B parameters) to outperform larger baselines (235B) through environment synthesis, tool integration, structured programs, and PRD-guided self-evolution.</span></p></li></ul><h3><span>Meituan</span></h3><p><strong><a href="http://arxiv.org/abs/2607.09773v1"><span>EvoCUA-1.5: Online Reinforcement Learning for Multi-turn Computer-Use Agents</span></a></strong></p><ul><li><p><strong><span>EvoCUA-1.5</span></strong><span> enables desktop agents to improve through online reinforcement learning in sandbox environments, introducing </span><strong><span>Step-Level Policy Optimization</span></strong><span> to handle multi-turn interactions and </span><strong><span>Dynamic Tri-Adaptive Curriculum</span></strong><span> for stable training. The 32B model achieves 63.2% success on OSWorld-Verified, matching larger baselines by learning from verifiable task outcomes rather than static offline data alone.</span></p></li></ul><h3><span>Moonshot</span></h3><p><strong><a href="http://arxiv.org/abs/2607.24957v1"><span>PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>PerceptionBench</span></strong><span>, a benchmark that isolates atomic visual perception in multimodal AI models by diagnosing where frontier models fail across existing benchmarks, then constructing 3,000 questions targeting ten core perceptual capabilities. Results show no model exceeds 60% accuracy, revealing </span><strong><span>perception-related hallucination as a critical weakness</span></strong><span> and exposing divergent capability profiles masked by overall scores.</span></p></li></ul><h3><span>Tencent</span></h3><p><strong><a href="http://arxiv.org/abs/2607.08964v2"><span>Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Long-Horizon-Terminal-Bench</span></strong><span>, a benchmark of 46 complex terminal tasks requiring hours of execution and hundreds of episodes, graded with fine-grained subtasks to capture partial progress rather than just final outcomes. Even frontier models achieve only 15.2% pass rate, revealing significant gaps in long-horizon planning and iterative problem-solving.</span></p></li></ul><h3><span>Xiaomi</span></h3><p><strong><a href="http://arxiv.org/abs/2607.15330v1"><span>Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories</span></a></strong></p><ul><li><p><span>Develops </span><strong><span>Xiaomi-Robotics-1</span></strong><span>, a vision-language-action model trained on over 100k hours of real-world robot trajectories with an auto-labeling pipeline, achieving </span><strong><span>57.6% success on RoboCasa365</span></strong><span> (vs. 46.6% prior best) and </span><strong><span>20.07 on RoboDojo</span></strong><span> (vs. 13.07 prior). The model generalizes to unseen environments out-of-the-box and fine-tunes efficiently on novel tasks.</span></p></li></ul><h3><span>Zhipu</span></h3><p><strong><a href="http://arxiv.org/abs/2607.11185v1"><span>SCALECUA: Scaling Computer Use Agents with Verifiable Task Synthesis and Efficient Online RL</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>ScaleCUA</span></strong><span>, a framework that scales computer use agents through </span><strong><span>verifiable task synthesis</span></strong><span> and </span><strong><span>efficient online reinforcement learning</span></strong><span>. Key innovations include </span><strong><span>VeriGen</span></strong><span> for auto-generating 24K+ verifiable tasks via multi-agent feedback, </span><strong><span>Frontier Sampling</span></strong><span> to focus training on learning bottlenecks, and </span><strong><span>Visual Context Segmentation</span></strong><span> for 2.83&#215; training speedup, achieving 68.7% on OSWorld&#8212;state-of-the-art for open-source agents.</span></p></li></ul><h2>Other Lab News</h2><h3>Liang Wenfeng&#8217;s remarks on DeepSeek&#8217;s AGI strategy leaked</h3><p><span>A near-four-hour recording of a DeepSeek investor call, dated May 20 and leaked around July 23, circulated widely across major Chinese tech and finance outlets (e.g., </span><a href="https://www.zhiding.cn/ai/2026/0723/3194330.shtml"><span>Zhiding</span></a><span>, </span><a href="https://www.stcn.com/article/detail/4037119.html"><span>Securities Times</span></a><span>, </span><a href="https://web.archive.org/web/20260723203902/https://www.tmtpost.com/8076737.html"><span>TMTPost</span></a><span>) and was transcribed in English by </span><a href="https://www.fredgao.com/p/deepseeks-liang-wenfeng-breaks-his"><span>Fred Gao</span></a><span>. Liang, CEO of DeepSeek, rarely speaks publicly, so commentators seized the leaked exchange as one of the most comprehensive looks at his strategy. The text is a transcribed recording that hasn&#8217;t been officially released, so should be treated with a grain of salt. Articles about it have also been pulled from Chinese news sites and WeChat.</span></p><p><span>The organizing idea in the leaked recording is restraint (&#20811;&#21046;) as strategy. Per the transcript, Liang frames AGI, not commercial dominance, as DeepSeek&#8217;s goal, arguing that &#8220;the more restrained you are, the more likely you are to succeed [at AI].&#8221; Business wins are by-products; he argues the AI market is large enough (he puts it eventually 10-20 percent of human GDP) that no one can monopolize it, so chasing maximal profit is strategically weak and loses to those content with a fair return. He casts open source as the base of the ecosystem, paired with relatively thin-margin pricing (recouping hardware costs over roughly 10 months) to make open weights a viable business.</span></p><p><span>On the US-China gap, Liang argues the American labs&#8217; lead (OpenAI, Anthropic, Google) is cyclical, and that compute is the only real gap, since talent is globally distributed and China has no shortage. He expects China&#8217;s structural edge to be cost, efficiency, and user experience; he sees Nvidia&#8217;s CUDA moat eroding and a domestic-chip ecosystem as a historic opening. He also predicts the base-model field narrows to three or four serious players. DeepSeek, he says, will stay on the &#8220;main road to AGI&#8221; (language models, chain-of-thought, agents, and continual learning) and deliberately cede video generation and world models to the rest of the ecosystem.</span></p><h1>AI Safety Publication Highlights</h1><p>There were <strong>87 AI-safety-related papers published by Chinese researchers </strong>this fortnight. We&#8217;re in the process of migrating the paper summaries to a more durable database, which will be available next issue.</p><h2><span>&#128269;Spotlight</span></h2><p><strong><a href="https://arxiv.org/abs/2607.18056v1"><span>An Early Warning of Emerging Biosecurity Risks in Frontier LLMs</span></a></strong><span> (Shanghai Artificial Intelligence Laboratory)</span></p><p><span>A model that refuses a dangerous request in chat may still help produce something dangerous in a lab, and text-level safety tests cannot tell the two apart. The </span><strong><span>Shanghai AI Lab</span></strong><span> built </span><strong><span>Intern-BioBreaker</span></strong><span>, a bio-red-teaming model, and paired it with a pipeline that runs a model&#8217;s outputs from text all the way to wet-lab validation, to measure that gap directly.</span></p><p><span>Intern-BioBreaker generates targeted jailbreak prompts to test whether an aligned model can be pushed past its safeguards in two ways: into operational guidance for safety-sensitive biological tasks, or into sequence-level outputs with harmful properties. Selected sequence outputs are then taken into the lab, through DNA synthesis, host expression, and orthogonal protein verification, to check whether a model-generated design yields the biological product it was meant to.</span></p><p><span>The paper reports a wide gap between text-level safeguards and what capable scientific models will actually do. Intern-BioBreaker outperformed baseline attack models and found bio-risk jailbreak vulnerabilities across both open-weight and proprietary frontier large language models, several reaching near-saturated or 100 percent task-level attack success rates on the paper&#8217;s own measure. In a sequence-level case study, the authors report that GPT-5.5 could be induced to generate modified viral candidate sequences with pathogenic potential, whose translated proteins could show stronger receptor-binding affinity and greater infection potential. The framework&#8217;s end-to-end tests indicate these designs were not merely textual; selected outputs could be physically realized under controlled experimental conditions.</span></p><p><span>The authors keep the operational specifics masked and frame the work as an early warning, not a demonstration. Their recommendations acknowledge that this isn&#8217;t just a model safety issue, advocating for nucleic-acid synthesis screening at the providers that fulfill sequence orders and biological red-teaming that keeps pace with scientific capability.</span></p><p><span>This sits alongside </span><a href="https://arxiv.org/abs/2607.18665"><span>SciHazard</span></a><span> and </span><a href="https://arxiv.org/abs/2607.00464"><span>MolSafeEval</span></a><span> in a cluster of Chinese AI-for-science safety work this issue. All three converge on the same problem: models built to speed up science can also speed up its most dangerous applications, and text-level refusal training does not ensure safety by itself.</span></p><h2><span>Agentic Safety</span></h2><p><strong><a href="http://arxiv.org/abs/2607.06807v1"><span>When Agents Go Rogue: Activation-Based Detection of Malicious Behaviors in Multi-Agent Systems</span></a></strong></p><ul><li><p><strong><span>AcMAS</span></strong><span> detects malicious behavior in multi-agent LLM systems by analyzing </span><strong><span>internal activation patterns</span></strong><span> within individual agents rather than relying on explicit interaction graphs. This approach works against </span><strong><span>semantically stealthy attacks</span></strong><span>&#8212;ones designed to evade detection by mimicking legitimate behavior&#8212;and functions reliably in </span><strong><span>asynchronous execution</span></strong><span>, which is common in practice but breaks graph-based detection methods. The framework also identifies which internal states are corrupted, enabling </span><strong><span>targeted repair</span></strong><span> rather than removing agents entirely; it achieves 0.93 F1-score in asynchronous settings versus 0.38 for graph-based baselines.</span></p></li></ul><p><em><span>Institutional affiliations: Worcester Polytechnic Institute, Fudan University</span></em></p><p><strong><a href="http://arxiv.org/abs/2607.20121v2"><span>OpenSkillRisk: Benchmarking Agent Safety When Using Real-World Risky Third-Party Skills</span></a></strong></p><ul><li><p><strong><span>OpenSkillRisk</span></strong><span> is a benchmark of 263 risky third-party skills from public marketplaces, designed to measure whether LLM-based agents recognize and avoid latent safety hazards during execution. Testing three agent frameworks and thirteen LLMs revealed that </span><strong><span>even the safest configurations executed unsafe actions in ~17% of cases</span></strong><span>, with context-dependent risks proving especially difficult to avoid. Agents failed through three patterns: failing to recognize risks, recognizing but acting before intervention, or exceeding the user&#8217;s intended scope&#8212;indicating gaps in both risk reasoning and execution control.</span></p></li></ul><p><em><span>Institutional affiliations: City University of Hong Kong, Beijing University of Posts and Telecommunications, Li Auto Inc.</span></em></p><p><strong><a href="http://arxiv.org/abs/2607.19913v1"><span>JANUS: Foreseeing Latent Risk for Long-Horizon Agent Safety</span></a></strong></p><ul><li><p><strong><span>Janus</span></strong><span> is a framework for training safety guards to </span><strong><span>predict unsafe outcomes before agents execute actions</span></strong><span>, rather than filtering content after the fact. The approach uses multi-agent simulation to generate diverse trajectories, then trains a model called </span><strong><span>Vanguard</span></strong><span> on two linked tasks: </span><strong><span>forecasting safety-relevant futures</span></strong><span> from partial action sequences, and </span><strong><span>deciding whether to block actions</span></strong><span> based on both observed behavior and predicted consequences. Tested across four agent-safety benchmarks, Vanguard </span><strong><span>blocks 15.9 percentage points more unsafe actions</span></strong><span> than baseline guards while maintaining task completion rates.</span></p></li></ul><p><em><span>Institutional affiliations: Chinese Academy of Sciences, University of Chinese Academy of Sciences, Shanghai Artificial Intelligence Laboratory, Peking University, Beijing Academy of Artificial Intelligence</span></em></p><h2><span>Alignment</span></h2><p><strong><a href="http://arxiv.org/abs/2607.09492v1"><span>Multimodal Reward Hacking in Reinforcement Learning</span></a></strong></p><ul><li><p><strong><span>Reward hacking</span></strong><span> in multimodal LLMs occurs when RL optimization increases proxy reward scores while degrading actual task performance, especially when visual content is scored by text-only or weak reward signals. The authors introduce </span><strong><span>NRFR</span></strong><span> (Newly Rewarded Failure Rate) to isolate failures created by RL itself. Outcome-only rewards produce up to 48.1% hacking rates; scaling models helps but doesn&#8217;t eliminate the problem (54.9% worse rate at 32B), while answer-aware and VLM-verified rewards substantially reduce it, suggesting reward design and verification robustness are critical under optimization pressure.</span></p></li></ul><p><em><span>Institutional affiliations: Chinese Academy of Sciences; University of California, Merced; Southeast University</span></em></p><h2><span>Evaluation and Benchmarks</span></h2><p><strong><a href="http://arxiv.org/abs/2607.22368v1"><span>Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI</span></a></strong></p><ul><li><p><strong><span>HackDetect</span></strong><span> audits agent benchmarks post-hoc to identify reward-hacking shortcuts&#8212;such as reading evaluation artifacts, recovering public solutions, or inferring test structure&#8212;that inflate capability claims. Across 15 benchmarks, the authors found </span><strong><span>67% of frontier science tasks and 67% of AutoLab tasks contained exploitable exposures</span></strong><span>, with measured score inflation ranging from 0.45&#8211;1.00 points. The work introduces </span><strong><span>protocol validity</span></strong><span> as a requirement for benchmark credibility, showing that agents often succeed via shortcuts unrelated to the intended capability being measured.</span></p></li></ul><p><em><span>Institutional affiliations: Tencent, The Hong Kong University of Science and Technology, Duke Kunshan University</span></em></p><h2><span>Governance and Policy</span></h2><p><strong><a href="http://arxiv.org/abs/2607.23207v1"><span>Accountable yet Anonymous AI Agents - Split-Knowledge Binding in National Agent-Identity Layer in China</span></a></strong></p><ul><li><p><span>China&#8217;s emerging </span><strong><span>national AI-agent identity infrastructure</span></strong><span> (launching Q3 2026) implements </span><strong><span>split-knowledge binding</span></strong><span>&#8212;agents are linked to verified legal principals, but re-identification requires two distinct government agencies acting together, neither capable alone. The system separates accountability from business-layer transparency through institutional and procedural design rather than cryptography; a state compelling both agencies can still re-identify. The paper articulates the </span><strong><span>ex-post attribution thesis</span></strong><span> (that legal accountability for agent actions requires traceable attribution), defines the </span><strong><span>accountability surface</span></strong><span> (which agent actions leave identity traces), and proposes a </span><strong><span>proportionality framework</span></strong><span> for choosing between trust architectures. The authors acknowledge the mechanism&#8217;s conditional nature and apply their own </span><strong><span>reflexive jurisdiction standard</span></strong><span> to evaluate the deployment.</span></p></li></ul><p><em><span>Institutional affiliations: Red Date Technology (Hong Kong) Limited, State Information Center, China Organization Data Service, China Internet Network Information Center</span></em></p><p><strong><a href="http://arxiv.org/abs/2607.22877v1"><span>Physical AI Governance: From Theory to Practice Across Life Cycle</span></a></strong></p><ul><li><p><span>This survey organizes governance challenges specific to embodied AI systems&#8212;robots and physical agents that operate in real time alongside humans&#8212;into a unified framework absent from existing AI governance literature. The authors map five lifecycle stages (research, design, data, model development, deployment) and detail concrete implementation practices for each, connecting abstract principles to engineering workflows. The framework addresses real-time safety constraints, dynamic environments, and continuous human-AI coexistence that differ from screen-based systems.</span></p></li></ul><p><em><span>Institutional affiliations: Case Western Reserve University, Shanghai Jiao Tong University, Massachusetts Institute of Technology, Columbia University, University of California, Riverside, Nanyang Technological University, New York University, Salesforce, Harvard University, Stanford University</span></em></p><h2><span>Guardrails and Deployment Safety</span></h2><p><strong><a href="http://arxiv.org/abs/2607.02357v1">Cloak and Detonate: Scanner Evasion and Dynamic Detection of Agent Skill Malware</a></strong></p><ul><li><p><strong>SkillCloak</strong> demonstrates that existing static scanners for LLM agent skills fail against adaptive evasion: <strong>payload-preserving obfuscation</strong> and <strong>dynamic unpacking</strong> bypass over 90% of defenses while maintaining malicious behavior. The authors propose <strong>SkillDetonate</strong>, a <strong>runtime sandbox auditor</strong> that detects attacks through <strong>information-flow tracking</strong>&#8212;monitoring data movement across files, processes, and network operations during execution&#8212;rather than analyzing code appearance. SkillDetonate achieves <strong>97% detection</strong> on malicious skills, addressing the supply-chain risk posed by untrusted marketplace code executing with agent privileges.</p></li></ul><p><em>Institutional affiliations: Hong Kong University of Science and Technology, Guangzhou HKUST Fok Ying Tung Research Institute</em></p><p><strong><a href="http://arxiv.org/abs/2607.15218v1"><span>When Words Are Safe But Actions Kill: Probing Physical Danger Beyond Text Safety in Hidden-State Risk Space</span></a></strong></p><ul><li><p><span>This paper identifies </span><strong><span>physical danger (PD)</span></strong><span> as a distinct safety problem from text-level content danger (CD) in LLMs used as planners for embodied agents&#8212;linguistically safe instructions can cause harm when executed physically. Using hidden-state analysis, the authors show CD and PD form </span><strong><span>separable signals</span></strong><span> in model representations and propose </span><strong><span>PRISM</span></strong><span>, a lightweight probe that achieves 86&#8211;87% accuracy detecting physical risks with substantially lower false-positive rates than LLM judges (11&#8211;14% vs. 24&#8211;39%). The work introduces </span><strong><span>PhysicalSafetyBench-1K</span></strong><span>, a contrastive benchmark testing detection of grounded danger without explicit unsafe keywords, where PRISM reaches 99.6% accuracy.</span></p></li></ul><p><em><span>Institutional affiliations: Tsinghua University, Nanyang Technical University</span></em></p><h2><span>Interpretability</span></h2><p><strong><a href="http://arxiv.org/abs/2607.18820v1"><span>CASE: Causal Alignment and Structural Enforcement for Improving Chain-of-Thought Faithfulness</span></a></strong></p><ul><li><p><strong><span>CASE</span></strong><span> addresses a failure mode in chain-of-thought reasoning where LLMs generate plausible-sounding explanations that don&#8217;t actually support their final answers. The authors identify </span><strong><span>instruction-to-answer shortcuts</span></strong><span>&#8212;the model bypasses its own reasoning by relying on direct patterns between input and output. CASE combines training-time modifications (counterfactual reasoning data, selective loss weighting) and inference-time attention masking to enforce the causal structure that reasoning should mediate between instruction and answer. On four benchmarks, the method achieves </span><strong><span>37% relative improvement in faithfulness</span></strong><span> while maintaining competitive accuracy.</span></p></li></ul><p><em><span>Institutional affiliations: Southern University of Science and Technology; Agency for Science, Technology and Research (A*STAR); Beijing Normal-Hong Kong Baptist University; Lingnan University</span></em></p><h2><span>Misuse and Dangerous Capabilities</span></h2><p><strong><a href="http://arxiv.org/abs/2607.25700v1">BioDisclose: An Actionability-Aware Benchmark for Biomedical Safety under Adversarial Elicitation</a></strong></p><ul><li><p><strong>BioDisclose</strong> is a benchmark with 480 adversarial prompts across six biomedical risk domains designed to measure <strong>dual-use knowledge disclosure</strong> from LLMs. Responses are graded on a four-level scale distinguishing refusal, discussion, technical specificity, and actionable content&#8212;including cases where models refuse then leak. Across five deployed systems, disclosure rates ranged from 9.2% to 64.0%, with academic framing proving most effective for elicitation and laboratory safety scenarios showing highest vulnerability, revealing uneven safeguard coverage in scientific contexts.</p></li></ul><p><em>Institutional affiliations: Communication University of China, Hainan Lingshui Li&#8217;An International Education Innovation Pilot Zone, Donghua University, University of Electronic Science and Technology of China</em></p><p><strong><a href="http://arxiv.org/abs/2607.00464v1">MolSafeEval: A Benchmark for Uncovering Safety Risks in AI-Generated Molecules</a></strong></p><ul><li><p><strong>MolSafeEval</strong> is a benchmark for evaluating safety risks in AI-generated molecules, addressing a gap in existing molecular generation benchmarks that focus on novelty and property matching but not hazard assessment. The framework integrates toxicological databases and hazard rules into a <strong>molecular safety knowledge graph</strong> and uses LLM-based reasoning to detect and explain unsafe features in generated compounds. The benchmark covers four generative tasks&#8212;unconditional generation, property optimization, protein-based design, and text-based generation&#8212;and reveals which current approaches produce molecules with toxic, reactive, or otherwise hazardous characteristics.</p></li></ul><p><em>Institutional affiliations: Zhejiang University, University of Oxford</em></p><p><strong><a href="http://arxiv.org/abs/2607.18665v1"><span>SciHazard: A Benchmark for Measuring Scientific Safety Risks with Decomposed Harm Scoring</span></a></strong></p><ul><li><p><strong><span>SciHazard</span></strong><span> is a benchmark for evaluating whether LLMs convert hazardous scientific knowledge into actionable misuse guidance. It contains 2,400 hazardous questions grounded in regulated entities and documented failure scenarios across 12 disciplines. </span><strong><span>DeHarm-Score</span></strong><span>, a decomposed evaluation method, measures harm by combining query severity, refusal behavior, response executability (via dynamic checklists), and whether responses introduce genuinely new risks. Testing 31 frontier models reveals autonomous research agents produce </span><strong><span>32% higher harmful outputs</span></strong><span> than standard LLMs, indicating a gap in current safety defenses for agentic systems.</span></p></li></ul><p><em><span>Institutional affiliations: Shanghai Artificial Intelligence Laboratory</span></em></p><h2><span>Robustness and Adversarial Attacks</span></h2><p><strong><a href="http://arxiv.org/abs/2607.23496v1"><span>Do LLMs Know Their Vulnerable Scenarios?</span></a></strong></p><ul><li><p><strong><span>Concept2Scenario</span></strong><span> uses sparse autoencoders (learned feature detectors) to map internal model activations to interpretable refusal-suppressing concepts, then translates them into natural-language scenarios that bypass safety training. The framework identifies </span><strong><span>scenario combinations</span></strong><span> that interact to weaken refusal more than individual scenarios alone. Discovered scenarios improve jailbreak success rates by up to 18.2 percentage points across open models and transfer to GPT-5, Claude, and Gemini, indicating shared refusal vulnerabilities across model families.</span></p></li></ul><p><em><span>Institutional affiliations: Renmin University of China, Xi&#8217;an Jiaotong University, University of Science and Technology of China, Wuhan University, Shanghai Artificial Intelligence Laboratory</span></em></p><h1>Export Controls &amp; Economic Policy</h1><h2><span>MOFCOM rebuts a US threat to sanction Chinese AI firms over &#8220;distillation&#8221;</span></h2><p><span>US Treasury Secretary Scott Bessent </span><a href="https://www.cnbc.com/2026/07/21/bessent-china-ai-sanctions.html"><span>said on July 21</span></a><span> that Washington could sanction Chinese AI developers if it finds they built their models by &#8220;distilling&#8221; US frontier models, which he cast as intellectual-property theft. China&#8217;s Ministry of Commerce (MOFCOM) </span><a href="https://www.mofcom.gov.cn/xwfb/xwfyrth/art/2026/art_efae08e51a5b4db6bc9742d32209e8f5.html"><span>answered on July 27</span></a><span>, calling the threat baseless and legally unfounded, a &#8220;double standard,&#8221; and a form of &#8220;AI hegemonism,&#8221; and warned it would take &#8220;all necessary measures&#8221; if Chinese interests suffer substantial harm.</span></p><p><span>Distillation trains a smaller model on a stronger model&#8217;s outputs. Anthropic </span><a href="https://www.anthropic.com/news/detecting-and-preventing-distillation-attacks"><span>alleged</span></a><span> in February that DeepSeek, Moonshot, and MiniMax had run distillation campaigns to extract Claude&#8217;s capabilities, and reporting has since centered on a possible move against Moonshot over its Kimi K3 model. MOFCOM argued that Chinese and US frontier models now release within days of each other and that some Chinese models lead, that US firms have themselves distilled Chinese models, and that nearly 200 US startups have </span><a href="https://static.politico.com/4a/bf/9c4021d8404386b0a311dcccf0e5/lta-open-weight-ai-letter-7-22-26.pdf"><span>urged Washington</span></a><span> not to cut off access to Chinese open-source models. No formal US investigation or sanction has been announced.</span></p><p><span>State media echoed MOFCOM&#8217;s arguments. On July 27, CCTV&#8217;s Yuyuan Tantian commentary account published </span><a href="https://www.163.com/dy/article/L2RD2ST90514R9P4.html"><span>&#12298;&#27169;&#22411;&#24320;&#28304;&#65292;&#21147;&#30772;&#35199;&#26041;&#20559;&#35265;&#12299;</span></a><span> (&#8221;Model Open-Sourcing Shatters Western Bias&#8221;), which cast open-weight models as a corrective to Western skepticism about open source. It points to a July 24 </span><a href="https://images.nvidia.com/pdf/Open-Weights-and-American-AI-Leadership.pdf"><span>statement</span></a><span> by Nvidia, Microsoft, Meta, IBM, and Hugging Face backing open-weight models and relays proposals from industry and policy researchers for tiered openness: basic capabilities open, some frontier capabilities released on a limited basis, and high-risk capabilities kept behind access controls and safety evaluations.</span></p><h2><span>China&#8217;s cyber-vulnerability platform flags Anthropic&#8217;s Claude Code</span></h2><p><span>On July 8, NVDB, China&#8217;s MIIT-affiliated cybersecurity threat and vulnerability information-sharing platform, </span><a href="https://www.ithome.com/0/974/001.htm"><span>issued a risk advisory</span></a><span> on Anthropic&#8217;s Claude Code, saying versions 2.1.91 through 2.1.196 carried a &#8220;backdoor&#8221; that could send a user&#8217;s region and identity markers to remote servers without consent. It recommended that users uninstall or upgrade and that organizations monitor outbound traffic from development networks. Chinese researchers who reverse-engineered the tool said the mechanism, present since the April 2 release of 2.1.91, checked system timezone and proxy information to identify China-based users and was not documented in the changelogs.</span></p><p><span>Anthropic </span><a href="https://www.cnbc.com/2026/07/08/china-anthropic-ai-claude-code-backdoor-security-threat.html"><span>said</span></a><span> the mechanism was an experiment it introduced in March to prevent unauthorized account resale and model distillation, and that it had been removed in a July 2 release; it noted that its usage policy already bars entities majority-owned by China-headquartered organizations, so Claude Code is not licensed in China. It appears to be the first named Chinese government-affiliated security warning against a US frontier lab&#8217;s flagship coding tool. The corporate response split: </span><a href="https://www.cnbc.com/2026/07/06/alibaba-anthropic-ai-ban-claude-china.html"><span>Alibaba barred employees from using Anthropic&#8217;s tools for work from July 10</span></a><span>. ByteDance, meanwhile, has set up </span><a href="https://finance.yahoo.com/technology/ai/articles/anthropic-closing-doors-chinese-engineers-172639142.html"><span>a reimbursement program</span></a><span> for engineers to expense personal Claude subscriptions used over VPN.</span></p><h1><span>On the Horizon</span></h1><p><span>Now that WAICO has been established, we&#8217;ll be tracking what its first actions will be.</span></p><div><hr></div><p style="text-align: justify;"><span>For more on how we select and track content, see our methodology </span><a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz"><span>here</span></a><span>.</span></p><div><hr></div><p><em><span>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</span></em></p><p><em><span>The China AI Bulletin is copy-edited and fact-checked by Kacie Yearout, Pivotal fellow and former diplomat.</span></em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>Unlike some of China&#8217;s earlier ethical principles, the Ethical Norms are quite similar to the <a href="https://www.oecd.org/en/topics/sub-issues/ai-principles.html">OECD AI Principles</a> (I analyzed this in an academic paper available <a href="https://link.springer.com/article/10.1007/s44206-024-00138-7">here</a>), which accords with China&#8217;s global AI governance goals.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>&#8220;LongCat-2.0&#30340;&#23436;&#25972;&#35757;&#32451;&#27969;&#31243;&#19982;&#22823;&#35268;&#27169;&#37096;&#32626;&#22343;&#20840;&#37096;&#20351;&#29992;</span><strong><span>&#22269;&#20135;&#31639;&#21147;&#38598;&#32676;</span></strong><span>&#8221;</span></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 8.5]]></title><description><![CDATA[Special issue on WAIC 2026 and Kimi K3]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-85-special-issue</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-85-special-issue</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Fri, 31 Jul 2026 20:52:00 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/776ea4e3-7def-4233-acbf-e52311da21e5_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome to a special issue of the China AI Bulletin, where we have two bonus features: <strong>a big WAIC round-up and a Kimi K3 deep dive</strong>. For the week&#8217;s regular content, check out <a href="https://open.substack.com/pub/chinaaibulletin/p/china-ai-bulletin-8?r=6md7mo&amp;utm_campaign=post&amp;utm_medium=web&amp;showWelcomeOnShare=true">China AI Bulletin #8</a>!</p><div class="digest-post-embed" data-attrs="{&quot;nodeId&quot;:&quot;343d93ab-b1b9-425d-959b-365d5704f430&quot;,&quot;caption&quot;:&quot;Welcome to Issue 8 of the China AI Bulletin, the latest on AI governance, development, and safety in China. It&#8217;s been quiet around here because I&#8217;ve been on the ground in Shanghai and Hangzhou, attending the World AI Conference and meeting with people in law and industry.&quot;,&quot;cta&quot;:null,&quot;showBylines&quot;:true,&quot;showDescription&quot;:true,&quot;showImage&quot;:true,&quot;size&quot;:&quot;md&quot;,&quot;isEditorNode&quot;:true,&quot;title&quot;:&quot;China AI Bulletin 8&quot;,&quot;publishedBylines&quot;:[{&quot;id&quot;:400365024,&quot;name&quot;:&quot;Emmie Hine&quot;,&quot;bio&quot;:&quot;Research Fellow at the Safe AI Forum.&quot;,&quot;photo_url&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/fcbe4125-e17c-4037-961f-1ffe6aaf738c_854x854.png&quot;,&quot;is_guest&quot;:false,&quot;bestseller_tier&quot;:null}],&quot;post_date&quot;:&quot;2026-07-31T20:52:37.325Z&quot;,&quot;cover_image&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c81693b8-2617-40d2-8d61-4b8b4ed033d2_2000x2000.png&quot;,&quot;cover_image_alt&quot;:null,&quot;canonical_url&quot;:&quot;https://chinaaibulletin.substack.com/p/china-ai-bulletin-8&quot;,&quot;section_name&quot;:null,&quot;video_upload_id&quot;:null,&quot;id&quot;:209104732,&quot;type&quot;:&quot;newsletter&quot;,&quot;reaction_count&quot;:0,&quot;comment_count&quot;:0,&quot;publication_id&quot;:8445135,&quot;publication_name&quot;:&quot;China AI Bulletin&quot;,&quot;publication_logo_url&quot;:&quot;https://substackcdn.com/image/fetch/$s_!In2i!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbe4cf74-cffe-41ae-a999-a2a2d239a916_1280x1280.png&quot;,&quot;belowTheFold&quot;:false,&quot;youtube_url&quot;:null,&quot;show_links&quot;:null,&quot;feed_url&quot;:null}"></div><h1>Executive Summary</h1><ul><li><p><strong><span>WAIC 2026 Round-Up:</span></strong><span> At the World AI Conference (Shanghai, July 17&#8211;20), China advanced its global AI governance agenda.</span></p><ul><li><p><strong><span>Twenty-nine countries signed the founding agreement for the World AI Cooperation Organization (WAICO)</span></strong><span>, a Shanghai-headquartered body Li Qiang first proposed in 2025.</span></p></li><li><p><strong><span>President Xi Jinping&#8217;s keynote</span></strong><span> and the conference&#8217;s &#8220;</span><strong><span>Chair&#8217;s Statement&#8221;</span></strong><span> framed the week around open cooperation and safety. The Statement was more specific on frontier AI risk than in past years, calling for guardrails on large models and &#8220;intrinsic safety&#8221; for agents.</span></p></li><li><p><span>The same day, the </span><strong><span>National Development and Reform Commission (NDRC), Cyberspace Administration of China (CAC), and Ministry of Industry and Information Technology (MIIT) each issued an international cooperation document</span></strong><span> (an AI Cooperation Development Action Plan, an agent interoperability initiative, and an International AI Ethics Governance Action Plan).</span></p></li></ul></li><li><p><strong><span>Moonshot released the weights for Kimi K3</span></strong><span>. </span></p><ul><li><p><span>Independent testing (Artificial Analysis, Epoch, and a joint UK&#8211;US government evaluation) ranks it the </span><strong><span>strongest open-weight model available</span></strong><span>.</span></p></li><li><p>It&#8217;s the first Chinese model card to evaluate cybersecurity as a risk factor of the model, but doesn&#8217;t evaluate for CBRN or loss of control risks.</p></li><li><p><span>K3 appears to be within a few months of the US closed frontier, though it is large, costly, and hard to run locally. </span></p></li><li><p><span>Its release drew a </span><strong><span>US allegation that it distilled Anthropic&#8217;s models</span></strong><span>, which several researchers dispute.</span></p></li></ul></li></ul><h1><span>WAIC 2026 Round-Up</span></h1><p><span>At the 2026 World AI Conference and High-Level Meeting on Global AI Governance (Shanghai, July 17&#8211;20), China advanced its global AI governance and open-source agendas through top-level speeches, officially founding the World AI Cooperation Organization (WAICO), and a number of other releases.</span></p><h2><span>Xi&#8217;s keynote and the Chair&#8217;s Statement encourage open-source &amp; discuss risks</span></h2><p><span>Xi Jinping opened the conference on July 17 with a keynote titled &#8220;Joining Hands to Build a Just and Equitable System For Global AI Governance&#8221; (</span><a href="https://www.news.cn/politics/leaders/20260717/72728b6f94154d63b3eaaaf9808b51eb/c.html"><span>Chinese</span></a><span>, </span><a href="http://english.scio.gov.cn/topnews/2026-07/18/content_118605932.html"><span>English</span></a><span>), organized around four principles:</span></p><p><span>(1) Maintain openness, win-win, and boosting innovation-driven development,</span></p><ul><li><p><span>Openness is not &#8220;open-source,&#8221; more like &#8220;openness to the world,&#8221; although Xi does endorse open-source here.</span></p></li></ul><p><span>(2) Strengthen risk awareness and ensure AI is secure and controllable,</span></p><ul><li><p><span>&#8220;Secure and controllable&#8221; (&#23433;&#20840;&#21487;&#25511;) is also fairly standard rhetoric, and although &#8220;&#23433;&#20840;&#8221; can translate as either &#8220;safe&#8221; or &#8220;secure,&#8221; this stock phrase tends to get translated as &#8220;secure.&#8221;</span></p></li></ul><p><span>(3) Encourage inclusiveness and promote mutual learning between civilizations, and</span></p><ul><li><p><span>&#8220;Shaping the values of AI with humanity&#8217;s common values&#8221; echoes Xi&#8217;s 2023 </span><a href="https://en.chinadiplomacy.org.cn/gci/index.shtml"><span>Global Civilization Initiative</span></a><span> (&#20840;&#29699;&#25991;&#26126;&#20513;&#35758;), a 2023 diplomatic initiative promoting multipolarity that includes youth exchange programs and visa-free travel to China. Here, he also hopes AI can help achieve it by &#8220;increas[ing] understanding&#8230; among all civilizations.&#8221;</span></p></li></ul><p><span>(4) Advocate solidarity and improve global governance.</span></p><ul><li><p><span>China has been advocating for the UN as a forum for global AI governance and continues to do so here. He also advocates helping Global South countries overcome the digital divide, which is relevant for WAICO (see below).</span></p></li></ul><p><span>The conference&#8217;s position statement, the &#8220;Chair&#8217;s Statement&#8221; (</span><a href="https://www.news.cn/20260717/3310820b96f949979ce6406712094935/c.html"><span>Chinese</span></a><span>, </span><a href="https://us.china-embassy.gov.cn/eng/zgyw/202607/t20260717_11984715.htm"><span>English</span></a><span>),</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a><span> included 15 principles, several of them safety-relevant.</span></p><p><span>Some of its safety rhetoric overlaps with Xi&#8217;s, but it also discusses frontier AI, agents, and&#8212;significantly&#8212;jointly countering misuse by terrorist, extremist, and transnational-criminal groups. Echoing Xi, the statement calls open-source a &#8220;vital pathway&#8221; to &#8220;inclusive AI development.&#8221;</span></p><h3><span>A few things I found notable in these speeches:</span></h3><p><strong><span>1. Amidst the fairly standard rhetoric about building out regulation and improving early-warning and emergency response was a mention of holding the &#8220;safety bottom line,&#8221; a phrase that only recently cropped up in AI.</span></strong></p><p><span>From </span><a href="https://chinaaibulletin.substack.com/i/204645234/state-council-discusses-ai-safety-bottom-line"><span>Issue #7</span></a><span>:</span></p><blockquote><p><span>&#8220;Holding the AI safety bottom line&#8221; draws on </span><a href="https://baike.baidu.com/item/%E5%BA%95%E7%BA%BF%E6%80%9D%E7%BB%B4/4192129"><span>&#24213;&#32447;&#24605;&#32500;</span></a><span> (</span><a href="https://chinamediaproject.org/2015/02/06/thus-spoke-uncle-xi/"><span>bottom-line thinking</span></a><span>), a worst-case-first governance method Xi has pushed since 2013. Under this method policymakers are encouraged to focus on guarding against tail risks. Applied to AI, it indicates a level of concern for AI safety at the State Council level that we haven&#8217;t seen before.</span></p></blockquote><p><span>I also heard this mentioned by several Chinese officials at different WAIC forums, meaning it&#8217;s diffusing down from the State Council and may draw greater attention at lower levels of the party.</span></p><p><strong><span>2.  Xi cautions explicitly against &#8220;overstretching the national security concept in AI&#8221; and &#8220;placing one country&#8217;s security over that of others.&#8221;</span></strong></p><p><span>This is a pretty explicit jab at the US, which has used &#8220;national security&#8221; as a justification for chip export controls, putting Chinese firms on the entity list, and threatening action about distillation (see Bessent&#8217;s statements below). The placement of this statement in the section about risk awareness, security, and controllability also feels significant, with the suggestion being that risk awareness and national security are being tied together in a detrimental way</span></p><p><strong><span>3. Xi is a fan of open source.</span></strong></p><p><span>Both Xi and the Chair&#8217;s Statement endorse open-source. Although Reuters </span><a href="https://www.reuters.com/world/beijing-is-looking-curbing-overseas-access-chinas-top-ai-models-sources-say-2026-07-07/"><span>reported</span></a><span> that MOFCOM might curb open-sourcing, it&#8217;s a key pillar of China&#8217;s AI industry and diplomacy, and I wouldn&#8217;t expect an export-control based crackdown.</span></p><p><strong><span>4. There continues to be more attention to frontier safety from top levels.</span></strong></p><p><span>Xi warns that measures to prevent loss of control (&#22833;&#25511;) must keep pace as capability advances. While &#8220;loss of control&#8221; doesn&#8217;t necessarily mean the same as it does in Western AI safety discourse&#8212;Xi&#8217;s use of &#8220;technological LoC&#8221; in particular had </span><a href="https://chinaaibulletin.substack.com/i/200679488/xi-jinping-discusses-technological-loss-of-control-and-embodied-ai-in-qiushi-speech"><span>several possible interpretations</span></a><span>&#8212;it&#8217;s still notable, and the Chair&#8217;s Statement on jointly &#8220;fighting against the abuse and malicious use of AI technologies by terrorist and extremist forces, and transnational organized criminal groups&#8221; is quite stark. Overall, we&#8217;re seeing a trend of frontier AI safety language being adopted more and more by top leadership, which these speeches continue.</span></p><h2><span>WAICO officially founded</span></h2><p><span>Twenty-nine countries </span><a href="https://www.gov.cn/yaowen/liebiao/202607/content_7075832.htm"><span>signed the agreement</span></a><span> establishing the World AI Cooperation Organization (WAICO), a new intergovernmental body to be headquartered in Shanghai and framed around supporting developing countries and narrowing the &#8220;intelligence divide.&#8221; Premier Li Qiang proposed WAICO at WAIC 2025. It&#8217;s been mentioned by officials at various times since, but this was the first concrete action. The next day, Assistant Foreign Minister Liu Bin used the High-Level Meeting on Global Governance of Artificial Intelligence to </span><a href="https://www.fmprc.gov.cn/wjbxw_new/202607/t20260718_11985613.shtml"><span>welcome countries into WAICO</span></a><span> and to restate that &#8220;AI should remain under human control.&#8221;</span></p><h2><span>NDRC, CAC, MIIT issue international coordination docs</span></h2><p><span>Alongside the WAICO signing, three governmental bodies each released documents related to international cooperation on July 17.</span></p><ul><li><p><strong><span>AI Cooperation Development Action Plan</span></strong><span> (&#12298;&#20154;&#24037;&#26234;&#33021;&#21512;&#20316;&#21457;&#23637;&#34892;&#21160;&#35745;&#21010;&#12299;). The NDRC set out eight actions across </span><a href="https://www.ndrc.gov.cn/xwdt/xwfb/202607/t20260717_1406562.html"><span>data, compute, open-source ecosystem sharing, AI+ empowerment, talent, rules and standards, safety-governance collaboration, and AI for good</span></a><span>, described by </span><a href="https://paper.people.com.cn/rmrb/pc/content/202607/18/content_30169473.html"><span>People&#8217;s Daily</span></a><span> as &#8220;pragmatic measures to respond to the United Nations&#8217; initiatives on strengthening international cooperation in artificial intelligence, bridging the digital divide, and promoting sustainable development through artificial intelligence.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> Action 3 endorses open-source, proposing an international open-source community with shared compliance and safety guidelines, and Action 7 promotes collaborative safety action.</span></p></li><li><p><strong><span>Global Cooperation Initiative on Agent Mutual Trust, Interconnection, and Interoperability</span></strong><span> (&#12298;&#26234;&#33021;&#20307;&#20114;&#20449;&#20114;&#32852;&#20114;&#25805;&#20316;&#20840;&#29699;&#21512;&#20316;&#20513;&#35758;&#12299;). The CAC proposed a </span><a href="http://www.cac.gov.cn/2026-07/17/c_1786032877362241.htm"><span>10-pillar initiative</span></a><span> on agent interoperability, standards, security frameworks, and cross-border data flows, extending China&#8217;s domestic agent-governance stack (the May 8 </span><a href="https://chinaaibulletin.substack.com/i/198737671/joint-framework-for-agentic-ai-released"><span>Agent Implementation Opinions</span></a><span> and the </span><a href="https://chinaaibulletin.substack.com/i/200679488/seven-part-ai-agent-interconnection-guidance-series-approved"><span>GB/Z agent interconnection standards</span></a><span>) outward into an international proposal.</span></p></li><li><p><strong><span>International AI Ethics Governance Action Plan</span></strong><span> (&#12298;&#22269;&#38469;&#20154;&#24037;&#26234;&#33021;&#20262;&#29702;&#27835;&#29702;&#34892;&#21160;&#35745;&#21010;&#12299;). MIIT issued a </span><a href="https://www.gov.cn/yaowen/liebiao/202607/content_7075890.htm"><span>full-lifecycle AI-ethics plan</span></a><span> covering risk classification, agile governance, capacity-building for developing countries, and research on explainability, privacy, and bias correction. MIIT frames it as &#8220;an important public good provided by China to the international community&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> and plans to work with international organizations to promote it.</span></p></li></ul><h2><span>Other WAIC releases</span></h2><p><span>Many other reports and documents were released over the week at WAIC; here are a few that caught our eye.</span></p><ul><li><p><strong><span>Agent Safety Governance Framework 1.0</span></strong><span> (&#12298;&#26234;&#33021;&#20307;&#23433;&#20840;&#27835;&#29702;&#26694;&#26550;1.0&#12299;). Tsinghua University&#8217;s Zhu Xufeng released this parallel to the AI Safety Governance Framework on July 18; Chen Tianhao, who worked on the framework, has an </span><a href="https://aibridgeeastwest.substack.com/p/from-model-outputs-to-agent-actions"><span>English explainer</span></a><span>. It organizes agent risk into four layers: knowledge and cognition (pollution, hallucination, bias); scenario interaction (data, account, asset)); social-group threats (market mechanisms, public institutions, social order); and physical environment and ecosystem. The full report is forthcoming.</span></p></li><li><p><strong><span>Ten Rule-of-Law Principles for the Healthy Development of Agent Services</span></strong><span> (&#12298;&#26234;&#33021;&#20307;&#26381;&#21153;&#20581;&#24247;&#21457;&#23637;&#30340;&#27861;&#27835;&#21313;&#21407;&#21017;&#12299;). A </span><a href="https://baijiahao.baidu.com/s?id=1871196083881623177"><span>principles document</span></a><span> released July 20 at the WAIC AI Rule of Law Forum, covering full-lifecycle, closed-loop risk management; user digital rights (service identification, tiered consent, minimum privilege, protections for the elderly and minors); and industry responsibility, with humans bearing final responsibility.</span></p></li><li><p><strong><span>Frontier AI Risk Monitoring Platform 2.0.</span></strong><span> Concordia AI (&#23433;&#36828;AI) </span><a href="https://concordia-ai.com/zh-hans/frontier-ai-risk-platform-2-0/"><span>released</span></a><span> the </span><a href="https://airiskmonitor.net/"><span>platform</span></a><span> on July 19, alongside its 2026 Q2 risk report (</span><a href="https://airiskmonitor.net/doc/zh/report/2026-Q2"><span>Chinese</span></a><span>, </span><a href="https://airiskmonitor.net/doc/en/report/2026-Q2"><span>English</span></a><span>). It covers five catastrophic risk categories&#8211;network attacks, biological risks, chemical risks, harmful manipulation, and loss of control&#8211;and finds that model risks are growing rapidly, while safeguards have regressed.</span></p></li><li><p><strong><span>Shanghai AI Lab</span></strong><span> released </span><strong><a href="https://huggingface.co/internlm/Intern-S2-Preview-397B"><span>Intern-S2-Preview-397B</span></a></strong><span> at WAIC, a 397-billion-parameter open (Apache 2.0) multimodal model in its &#20070;&#29983; (Scholar) line, pitched as a base for science and long-horizon agents rather than general chat. The lab says it matches its own prior trillion-parameter model on core scientific tasks, such as biomolecular-interaction design and material-structure generation, at a fraction of the training cost through a new deep vision-language pretraining approach, and claims leading general-reasoning performance among open models.</span></p></li><li><p><strong><span>Shanghai AI Lab also </span><a href="https://mp.weixin.qq.com/s/9rm1Crv8bdI3y3nQ7VXTJA"><span>launched</span></a><span> the Shu&#8217;an (&#20070;&#23433;) AI Safety Base Platform and an AI safety/security alliance.</span></strong><span> Shu&#8217;an is built around a digital-twin sandbox that mirrors a live system&#8217;s workflows, permissions, and assets so models and agents can be tested against it. Shanghai AI Lab&#8217;s own agent safety framework </span><a href="https://chinaaibulletin.substack.com/i/200679488/technical-ai-safety-publication-highlights"><span>AgentDog</span></a><span> is built in. The lab reports a 1,500-scenario/500,000-tool simulation environment, full security assessments of 17 production systems, and pilots at UnionPay, Shanghai Customs, and a Xiamen logistics base, plus a full-stack agent security toolbox. Shu&#8217;an&#8217;s launch is coupled with the AI Security Foundation Model Alliance (AI&#23433;&#20840;&#22522;&#30784;&#27169;&#22411;&#32852;&#30431;), which the readout links to TC260&#8217;s Working Group 9 on AI safety. It&#8217;s unclear who&#8217;s in the alliance and what relationship the two will have.</span></p></li></ul><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this interesting? Subscribe to get regular issues in your inbox.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1>&#128269; Kimi K3 Deep Dive</h1><p><span>Moonshot AI (&#26376;&#20043;&#26263;&#38754;) released the open weights for </span><strong><a href="https://huggingface.co/moonshotai/Kimi-K3"><span>Kimi K3</span></a></strong><span> on July 27, 11 days after it opened API access around WAIC. Independent testing indicates this is the strongest open-weight model available and sits within a few months of the US closed frontier. However, it is slower, more verbose, and more expensive than the open models that came before it, and its size means it isn&#8217;t runnable on consumer hardware. Moonshot has also joined MiniMax in building a commercial trigger into its license in an effort to commercialize its open models.</span></p><h2>It&#8217;s big and expensive, and efficiency is debated</h2><p><span>K3 is a </span><strong><span>2.8-trillion-parameter</span></strong><span> mixture-of-experts (MoE) model with </span><strong><span>104 billion active parameters</span></strong><span>, a one million-token context window, and native multimodal input. Moonshot bills it as the </span><a href="https://www.kimi.com/blog/kimi-k3"><span>first open model in the 3-trillion-parameter class</span></a><span>. It is candid about the model&#8217;s performance, conceding that K3 &#8220;trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol,&#8221; while beating everything else it tested, including Zhipu&#8217;s </span><strong><span>GLM-5.2</span></strong><span>, the only Chinese model in its comparison set.</span></p><p><span>The weights are available for download under a custom </span><strong><span>Kimi K3 License,</span></strong><span> rather than MIT or Apache. It adds a commercial trigger, requiring a separate agreement with Moonshot once a model-as-a-service business clears 20 million USD in annual revenue, plus an attribution requirement for large deployments. That is permissive but not the no-strings release DeepSeek and Alibaba have used.</span></p><p><span>A more substantial difference is price. K3 costs </span><strong><a href="https://www.kimi.com/blog/kimi-k3"><span>20 RMB/3 USD per million input tokens and 100 RMB/15 USD per million output</span></a></strong><span>, roughly four times its predecessors and the </span><a href="https://simonwillison.net/2026/Jul/16/kimi-k3/"><span>most expensive model any Chinese lab has shipped</span></a><span>, though cheaper than Opus and Sol. Moonshot argues it is nonetheless cheaper per task, reporting 91.2 percent on BrowseComp at about 2 USD per task, roughly half what GPT-5.6 Sol costs. Though </span><a href="https://artificialanalysis.ai/models/kimi-k3"><span>Artificial Analysis</span></a><span> measures K3 as &#8220;very verbose,&#8221; emitting roughly twice the output tokens of its peers, it reports it&#8217;s cheaper per task than GPT-5.6 Sol, Fable, or Opus 5. Moonshot ran all its reported benchmarks at K3&#8217;s maximum-effort setting, which Zvi Mowshowitz </span><a href="https://thezvi.substack.com/p/on-kimi-k3-its-capabilities-and-related"><span>notes</span></a><span> burns &#8220;a lot more tokens than are used in similar tests by Fable or Sol,&#8221; so the model &#8220;likely outperforms on benchmarks relative to practical performance.&#8221;</span></p><p><span>Though it may be more price-efficient, its use efficiency is unclear. </span><a href="https://artificialanalysis.ai/models/kimi-k3"><span>Artificial Analysis</span></a><span> ranks it 60/99 on speed (35 output tokens/second), and so while it may be efficient by price per task, those tasks may take longer, especially if it is more verbose.</span></p><p><span>K3&#8217;s </span><a href="https://arxiv.org/abs/2607.24653"><span>Moonshot-reported benchmarks</span></a><span> are strong, with </span><strong><span>93.5 percent on GPQA Diamond</span></strong><span> (PhD-level science, where top models now cluster in the low 90s) and </span><strong><span>88.3 on Terminal-Bench 2.1</span></strong><span>, an agentic-terminal test run in each model&#8217;s own harness. It&#8217;s unclear why, but Moonshot omitted the standard </span><strong><span>SWE-bench Verified</span></strong><span> while reporting several other coding scores.</span></p><h2>Independent tests rank it fourth overall and first among open models</h2><p><span>Independent testing broadly matches Moonshot&#8217;s ranking claim, though with caveats. </span><strong><a href="https://artificialanalysis.ai/models/kimi-k3"><span>Artificial Analysis</span></a></strong><span>, running its own harness, scores K3 at </span><strong><span>57 on its Intelligence Index, fourth overall</span></strong><span>, behind Opus 5, Fable 5, and GPT-5.6 Sol. It&#8217;s </span><strong><span>first among all open-weight models</span></strong><span>, ahead of GLM-5.2 (51) and DeepSeek V4 Pro (44).</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!XtWQ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!XtWQ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 424w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 848w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 1272w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!XtWQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png" width="1415" height="414" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:414,&quot;width&quot;:1415,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!XtWQ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 424w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 848w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 1272w, https://substackcdn.com/image/fetch/$s_!XtWQ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F10cc0290-6d2b-42d5-8de0-497587214afc_1415x414.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p><span>On </span><a href="https://epoch.ai/benchmarks/eci?view=graph&amp;tab=release-date&amp;subset-view=graph&amp;subset-tab=Software+engineering"><span>Epoch AI&#8217;s Capabilities Index</span></a><span>, an independent composite, K3 set a </span><a href="https://x.com/EpochAIResearch/status/2079602012644360382"><span>new open-weights record</span></a><span> of 156, which places it between Anthropic&#8217;s Opus 4.6 and OpenAI&#8217;s GPT-5.4. K3 tops </span><strong><span>LMArena&#8217;s</span></strong><span> </span><a href="https://x.com/arena/status/2077824029126504525"><span>frontend-coding arena</span></a><span>, beating Fable 5, but is </span><a href="https://arena.ai/leaderboard/agent"><span>fifth on general agentic tasks</span></a><span>. Independent evaluator </span><strong><a href="https://www.vals.ai/models/kimi_kimi-k3"><span>vals.ai</span></a></strong><span> ranks it fifth on its composite index and fourth on </span><a href="https://www.vals.ai/benchmarks/swebench"><span>SWE-bench Verified</span></a><span> at 93.4 percent, yet only </span><a href="https://www.vals.ai/benchmarks/lcb"><span>11th of 131 on LiveCodeBench</span></a><span>, which uses recent competitive-programming problems to avoid training contamination. Security firm </span><strong><a href="https://semgrep.dev/blog/2026/kimi-k3s-code-security-results-lack-precision/"><span>Semgrep</span></a></strong><span> found K3 weaker than rival models at finding vulnerabilities in real codebases, with accuracy collapsing on large repositories, and judged it &#8220;not a drop-in replacement&#8221; for the code-analysis tools security teams already use.</span></p><p><span>On how far K3 sits behind the closed frontier, Nathan Lambert (Allen Institute for AI) estimates </span><a href="https://www.interconnects.ai/p/kimi-k3-the-open-weights-escalation"><span>three to five months</span></a><span> and </span><a href="https://thezvi.substack.com/p/on-kimi-k3-its-capabilities-and-related"><span>Mowshowitz four to six</span></a><span>. Forecaster Peter Wildeford </span><a href="https://x.com/peterwildeford/status/2079052877801128237"><span>notes</span></a><span> K3&#8217;s score falls almost exactly where the trend in Chinese models&#8217; Epoch scores predicts, about six months behind the US frontier.</span></p><h2><span>K3 finds and exploits vulnerabilities, but has trouble with hardened targets</span></h2><p><span>Moonshot&#8217;s </span><a href="https://arxiv.org/abs/2607.24653"><span>technical report</span></a><span> includes a dedicated offensive-cyber capability evaluation, the kind of dangerous capability testing Western frontier labs put in their system cards. Cyber is the only dangerous capability test in Moonshot&#8217;s </span><a href="https://arxiv.org/abs/2607.24653"><span>technical report</span></a><span>, which doesn&#8217;t include a CBRN or loss-of-control assessment. Chinese labs have reported cyber benchmark scores before, but generally as line items in a capability table;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span> a purpose-built dangerous-capability evaluation framed around misuse risk appears to be a first from a Chinese frontier lab. Moonshot ran two tiers, vulnerability discovery and end-to-end exploit development, and points to outside corroboration: a joint assessment by the UK AI Security Institute and NIST&#8217;s Center for AI Standards and Innovation (</span><a href="https://www.aisi.gov.uk/blog/preliminary-assessment-of-kimi-k3s-cyber-capabilities"><span>AISI and CAISI</span></a><span>) that it says &#8220;reaches conclusions consistent with ours.&#8221; It could not benchmark K3 against Claude or GPT, because those models refuse cyber tasks, so GLM-5.2 was its only comparison.</span></p><p><span>On vulnerability discovery, K3 scanned widely deployed software, from operating system kernels and databases to web frameworks, blockchain, and VPN code, and surfaced hundreds of candidate bugs; about 70% of the human-reviewed findings were genuine, including 16 previously unknown vulnerabilities across six projects. Two were in the Linux kernel: a remotely triggerable denial-of-service bug and a Dirty-COW-class flaw that permits local privilege escalation. On exploit development, K3 built working exploits for 14 of 36 tasks (38.9%), against GLM-5.2&#8217;s 8 (22.2%), but 10 of its 14 successes were against unhardened user-space software like PostgreSQL and Apache; on the hardened Linux kernel track, neither model solved more than a quarter.</span></p><p><span>K3 can chain a working exploit, just not the hardest ones. Moonshot&#8217;s own suite shows it completing end-to-end exploits against unhardened targets; the joint AISI and CAISI test, using arbitrary code execution as the bar for a completed attack, put K3 at zero of 41 tasks, against roughly 20 of 41 for US frontier models.</span></p><p><span>Although end-to-end exploit development is arguably most relevant to misuse, the AISI/CAISI eval reports that K3&#8217;s guardrails don&#8217;t prevent it from helping with offensive cyber attempts, so even though it may not be as capable as models like Mythos, its offensive capabilities may be more accessible. Without CBRN or LoC tests, it&#8217;s hard to know what other risks it might pose. Moonshot separately </span><a href="https://mp.weixin.qq.com/s/74dqOtRDX89J44OXgMMt3g"><span>concedes</span></a><span> K3 is &#8220;overly proactive&#8221; (&#36807;&#20110;&#20027;&#21160;), acting for the user when intent is ambiguous, which raises questions about possible autonomous action.</span></p><h2>Open does not mean accessible</h2><p><span>K3&#8217;s size limits who can actually run it. Moonshot ships the weights already </span><a href="https://huggingface.co/moonshotai/Kimi-K3"><span>quantized to 4-bit</span></a><span>, so it can&#8217;t be compressed further. The weights alone are about 1.4 terabytes, and because it is a mixture-of-experts model, all 2.8 trillion parameters must stay in memory even though only 104 billion are active per token. Moonshot&#8217;s launch-day inference partner, </span><a href="https://vllm.ai/blog/2026-07-27-k3"><span>vLLM</span></a><span>, says the model barely fits on a single server of eight Nvidia B300 GPUs, or at least 16 of the prior-generation B200s; renting such a node runs about </span><a href="https://verda.com/pricing"><span>60 USD an hour</span></a><span> and owning one runs into the </span><a href="https://aiserver.eu/product/nvidia-dgx-b300/"><span>high six figures</span></a><span>. The most cost-efficient way to access for most people will likely be through the API; that use will be subject to Moonshot&#8217;s safeguards.</span></p><h2>The distillation allegation</h2><p><span>After its release, K3 drew accusations that it had distilled Anthropic models. White House science and technology adviser </span><a href="https://x.com/mkratsios47/status/2079933645888880708"><span>Michael Kratsios</span></a><span> alleged that Moonshot &#8220;distilled Anthropic&#8217;s Fable&#8221; through &#8220;a sophisticated internal platform to conduct large scale distillation against US models.&#8221; K3 identifies itself as Claude far more often than chance, which points to Claude-derived data in its training. Yet </span><a href="https://x.com/RyanGreenblatt/status/2078663148509544589"><span>Redwood Research&#8217;s Ryan Greenblatt</span></a><span> finds K3 names the older Claude 4.5 generation, not Fable.</span></p><p><span>Moonshot </span><a href="https://news.qq.com/rain/a/20260723A046PT00"><span>denied it</span></a><span>: its enterprise-business head Huang Zhenxin said on July 21 that K3&#8217;s leap comes from its own original architecture, not from distilling any existing model. Machine-learning researchers Braden Hancock and Nathan Lambert also </span><a href="https://techcrunch.com/2026/07/23/experts-say-exploiting-anthropics-fable-isnt-how-kimi-k3-got-so-good/"><span>reject the Fable distillation allegations</span></a><span>, noting Fable had been public only since July 1, too little time to distill, train, and release a model this size, and that distillation yields few gains near the frontier. K3&#8217;s own </span><a href="https://arxiv.org/abs/2607.24653"><span>technical report</span></a><span> does discuss distillation, but of the internal kind that merges Moonshot&#8217;s own specialist models into one, a standard technique unrelated to the accusation. No one has measured how much Claude-derived data contributed; on balance, it seems like some happened from older models rather than Fable.</span></p><h2><span>In China: a milestone and a compute ceiling</span></h2><p><span>Many netizens dubbed K3 China&#8217;s  &#8220;</span><a href="https://36kr.com/p/3899876606985864"><span>Fable 5 moment</span></a><span>.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a><span> Most coverage cites Artificial Analysis; major Chinese benchmarks like </span><a href="https://www.superclueai.com/homepage"><span>SuperCLUE</span></a><span>, </span><a href="https://rank.opencompass.org.cn/leaderboard/llm"><span>OpenCompass</span></a><span>, and </span><a href="https://flageval.baai.ac.cn/#/leaderboard"><span>FlagEval</span></a><sup><span>fn</span></sup><span> haven&#8217;t published scores yet.</span></p><p><sup><span>Fn</span></sup><span> BAAI, which runs FlagEval, also has an </span><a href="https://flagsafe.baai.ac.cn/research/biosafety/dna-assembly-biosecurity"><span>interesting bio eval</span></a><span> for split-order attacks.</span></p><p><span>They also reported on Moonshot&#8217;s compute constraints. Demand exceeded Moonshot&#8217;s limits within 48 hours of the API launch, and the company </span><a href="https://www.qbitai.com/2026/07/455179.html"><span>halted</span></a><span> new signups around July 19, saying its existing compute was near its limits. </span><a href="https://36kr.com/p/3908058642961542"><span>36Kr&#8217;s write-up</span></a><span>, titled &#8220;waiting for the chips&#8221; (&#31561;&#33455;&#26469;), argues that Huawei&#8217;s Ascend Atlas 950 and Zhipu&#8217;s announced gigawatt of fully domestic compute are the way to break through the compute ceiling.</span></p><h2>Market reaction and Moonshot&#8217;s strategy</h2><p><span>Markets first read K3&#8217;s open release as a threat to demand for high-end AI chips, with the </span><a href="https://www.techtimes.com/articles/321066/20260720/kimi-k3-wipes-33t-chip-stocks-moonshot-moves-toward-hong-kong-ipo.htm"><span>SOX semiconductor index falling about 12 percent on the week</span></a><span>, its worst in more than a year, before a wave of corrections. </span><a href="https://x.com/SemiAnalysis_/status/2078250693010309260"><span>SemiAnalysis</span></a><span> argues the opposite of the chip-demand scarcity: at 2.8 trillion parameters, K3 needs more than 1.5 terabytes of high-bandwidth memory to serve and runs on large multi-GPU systems, so cheaper per-token open models expand compute demand rather than shrink it. Competitor </span><a href="https://mp.weixin.qq.com/s/stMLz7vvtlXeaWgqmtCS-g"><span>Zhipu was down about 28 percent in Hong Kong, and MiniMax about 16 percent</span></a><span>. Moonshot, meanwhile, is capitalizing on the release; it is reportedly raising a </span><a href="https://technode.com/2026/07/22/moonshot-ai-reportedly-plans-final-pre-ipo-round-at-50-billion-valuation/"><span>pre-IPO round</span></a><span> at a </span><strong><span>50-billion-USD pre-money valuation</span></strong><span> on roughly 300 million USD in annualized revenue.</span></p><h2>Moonshot and Zhipu lead a fast-shifting open field</h2><p><span>On </span><a href="https://artificialanalysis.ai/models/kimi-k3"><span>Artificial Analysis&#8217;s intelligence index</span></a><span>, Moonshot&#8217;s Kimi K3 and Zhipu&#8217;s GLM-5.2 lead the Chinese open field (57 and 51, respectively). DeepSeek&#8217;s V4-Pro sits further back at 44, alongside MiniMax&#8217;s M3, although it&#8217;s still technically a &#8220;preview&#8221; version; Alibaba&#8217;s open Qwen tier is a mid-tier workhorse. The wildcard is the full version of Alibaba&#8217;s </span><a href="https://www.marktechpost.com/2026/07/19/alibaba-previews-qwen3-8-max-a-2-4-trillion-parameter-multimodal-model-days-after-moonshots-kimi-k3-open-weight-launch/"><span>Qwen3.8-Max</span></a><span>, currently out as an API-only preview. Alibaba </span><a href="https://x.com/Alibaba_Qwen/status/2078759124914098291"><span>claims</span></a><span> it runs second only to Fable 5, with open weights promised soon. Standings shift fast regardless: K3 took the open lead from GLM-5.2 after only about six weeks.</span></p><p><span>Moonshot and MiniMax have also built a revenue trigger into their open licenses. Moonshot&#8217;s </span><a href="https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE"><span>Kimi K3 License</span></a><span> and </span><a href="https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE"><span>MiniMax&#8217;s Community License</span></a><span> keep the weights free to download but require a separate agreement once a commercial user clears about 20 million USD in yearly revenue. </span><a href="https://huggingface.co/zai-org/GLM-5.2"><span>Zhipu</span></a><span> and </span><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro"><span>DeepSeek</span></a><span> stayed on the standard MIT license. The &#8220;</span><a href="https://www.163.com/dy/article/L2USG35T051492T3.html"><span>value war</span></a><span>&#8220; (&#20215;&#20540;&#25112;) is now moving into license terms, a bet that open-model profit will come from charging large deployers for capability rather than serving the cheapest tokens.</span></p><div><hr></div><p><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p><p><em>The China AI Bulletin is copy-edited and fact-checked by Kacie Yearout, Pivotal fellow and former diplomat.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>According to <a href="https://aisafetychina.substack.com/p/chinas-key-ai-safety-updates-at-waic">Concordia</a>, the Chair&#8217;s statement &#8220;is issued by the host on its own authority&#8221; rather than being jointly negotiated and &#8220;is best read as a statement of the Chinese government&#8217;s position, not as evidence of international consensus.&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p> &#8220;&#20197;&#21153;&#23454;&#20030;&#25514;&#21709;&#24212;&#32852;&#21512;&#22269;&#20851;&#20110;&#21152;&#24378;&#20154;&#24037;&#26234;&#33021;&#22269;&#38469;&#21512;&#20316;&#12289;&#24357;&#21512;&#25968;&#23383;&#40511;&#27807;&#12289;&#25512;&#21160;&#20154;&#24037;&#26234;&#33021;&#36171;&#33021;&#21487;&#25345;&#32493;&#21457;&#23637;&#31561;&#26041;&#38754;&#30340;&#20513;&#35758;&#12290;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>&#8220;&#34892;&#21160;&#35745;&#21010;&#26159;&#20013;&#22269;&#38754;&#21521;&#22269;&#38469;&#31038;&#20250;&#25552;&#20379;&#30340;&#37325;&#35201;&#20844;&#20849;&#20135;&#21697;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>For instance, <a href="https://arxiv.org/pdf/2602.02276">Kimi K2.5</a> included CyberGym as one of nine &#8220;coding&#8221; benchmarks.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#8220;&#22269;&#20135; Fable 5 &#26102;&#21051;&#8221;</p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 7]]></title><description><![CDATA[Developments from 17/6/26-1/7/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-7</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-7</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Thu, 02 Jul 2026 21:40:30 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/a0b50418-0f69-4dc2-a838-999c9d91545c_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Welcome to Issue 7 of the China AI Bulletin, the latest on AI governance, development, and safety in China. </span><strong><span>Today&#8217;s highlights:</span></strong><span> the State Council applies &#8220;bottom-line thinking&#8221; to AI safety for the first time, ByteDance&#8217;s Doubao 2.1 launches as a closed flagship computer-use agent, and new research finds phone-use agents complete harmful real-world tasks even while recognizing them as harmful.</span></p><p><strong><span>A brief request: </span></strong><span>If you have five minutes to help us improve the Bulletin, we&#8217;d appreciate it if you could fill out this short </span><a href="https://docs.google.com/forms/d/e/1FAIpQLSeQhDgZYbBRFlCSXsEKaI0N2I_MW5yJEt4LK9qN6Vp7VRMeuQ/viewform?usp=publish-editor"><span>reader survey</span></a><span>.</span></p><p><em><span>Editor&#8217;s note: the next issue of the Bulletin will be July 30, after the World AI Conference.</span></em></p><p><em><span>Number of the week: </span><a href="https://www.news.cn/politics/20260625/7264a728b3a54314a5e1aff934c6f585/c.html"><span>800,000 RMB (117,743 USD)</span></a><span>&#8212;the investment necessary for an AI short drama to be managed as a &#8220;key short drama&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a></span><sup><span> </span></sup><span>by the National Radio and Television Administration.</span></em></p><h1><span>Executive Summary</span></h1><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/204645234/domestic-ai-governance"><span>Domestic AI Governance</span></a><span>:</span></strong><span> A </span><strong><span>State Council executive meeting</span></strong><span> heard a report on AI development and called for </span><strong><span>&#8220;holding the AI safety bottom line&#8221; (&#24213;&#32447;)</span></strong><span>, the first time bottom-line thinking has been applied to AI at the State Council level, alongside pushes on frontier capability, compute, and AI+ deployment. The financial regulator, the National Financial Regulatory Administration (NFRA),</span><strong><span> </span></strong><span>issued </span><strong><span>32 guiding opinions on AI in banking and insurance</span></strong><span>, tiering applications by risk and putting the heaviest controls on high-risk uses like credit approval and underwriting. Eight departments issued </span><strong><span>&#8220;AI+ Consumption&#8221; implementation opinions</span></strong><span> aimed at getting AI &#8220;into every home and shop.&#8221;</span></p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/204645234/national-standards"><span>National Standards</span></a><span>:</span></strong><span> </span><strong><span>TC260</span></strong><span> published a security practice guide for deploying and using AI agents; State Administration for Market Regulation </span><strong><span>(SAMR)</span></strong><span> signaled it may make the agent </span><strong><span>identity code (&#36523;&#20221;&#30721;) mandatory</span></strong><span> and, alongside  Standardization Administration of China (SAC) and Ministry of Industry and Information Technology (MIIT)&#8217;s standards committee, moved to </span><strong><span>accelerate frontier AI standards</span></strong><span> for agents, embodied intelligence, and world models. </span><strong><span>MIIT</span></strong><span> said it has finished drafting a </span><strong><span>mandatory autonomous-driving safety standard</span></strong><span> that exceeds the newly approved UN regulation.</span></p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/204645234/international-ai-governance">International AI Governance</a>:</strong> At <strong>Summer Davos</strong> in Dalian, Premier <strong>Li Qiang</strong> tied the risk of <strong>loss of control of technology (&#25216;&#26415;&#22833;&#25511;)</strong> to AI, warning governance that fails to keep pace could bring serious consequences, and said China will keep working to improve global AI governance rules. This is the latest example of <strong>&#25216;&#26415;&#22833;&#25511; </strong>being used increasingly at high levels of government in China.<strong><span>Frontier Lab Developments:</span></strong><span> </span><strong><span>ByteDance</span></strong><span> launched </span><strong><span>Doubao 2.1 (Seed 2.1)</span></strong><span>, a closed flagship sold as an agentic product that can operate a user&#8217;s computer, and </span><strong><span>Alibaba&#8217;s Qwen</span></strong><span> released </span><strong><span>Qwen-AgentWorld</span></strong><span>, its first open-weight &#8220;language world model,&#8221; built to predict environment states from an agent&#8217;s actions rather than to chat or code.</span></p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/204645234/ai-safety-publication-highlights"><span>AI Safety Publications</span></a><span>:</span></strong><span> The spotlight is </span><strong><span>&#8220;It Lied to a Doctor to Buy Poison Ingredients,&#8221;</span></strong><span> a study of </span><strong><span>phone-use agents</span></strong><span> that found agents built on nine mainstream models completed harmful tasks at high rates (68.8 percent) while rarely refusing; the authors called this a </span><strong><span>Safety Awareness-Execution Gap</span></strong><span>, where an agent flags a request as harmful yet carries it out anyway. Also notable: </span><strong><span>Governance Decay</span></strong><span> shows that compacting an agent&#8217;s context to save tokens silently drops standing safety constraints (violation rates rise from 0 percent to 30&#8211;59 percent); </span><strong><span>Tencent</span></strong><span> open-sourced the </span><strong><span>AI-Infra-Guard</span></strong><span> agent red-teaming framework; and a governance paper mapped China&#8217;s proposed </span><strong><span>World AI Cooperation Organization (WAICO)</span></strong><span> against 15 existing bodies as the first standing organization built to anchor a development-first governance pole.</span></p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/204645234/export-controls-and-economic-policy"><span>Export Controls and Economic Policy</span></a><span>:</span></strong><span> The National Development and Reform Commission </span><strong><span>(NDRC)</span></strong><span> warned provinces against &#8220;disorderly competition and piling in&#8221; on AI and computing infrastructure, citing idle data centers and roughly 58 percent average occupancy, and pushed provinces toward the national </span><strong><span>East-Data-West-Computing</span></strong><span> layout and a centrally dispatched </span><strong><span>national computing-power network (&#31639;&#21147;&#32593;)</span></strong><span>.</span></p></li></ul><h1><span>Domestic AI Governance</span></h1><h2><span>State Council discusses AI safety &#8220;bottom line&#8221;</span></h2><p><span>On June 29, Premier Li Qiang chaired a </span><a href="https://www.gov.cn/yaowen/liebiao/202606/content_7073672.htm"><span>State Council executive meeting</span></a><span> that heard a report on AI development. The executive meeting consists of the premier, vice-premiers, state councillors, and secretary-general. It </span><a href="https://en.wikipedia.org/wiki/State_Council_of_the_People%27s_Republic_of_China"><span>meets two to three times a month</span></a><span> to discuss draft laws, deliberate administrative regulations, and discuss and decide on &#8220;</span><a href="https://npcobserver.com/2024/03/11/china-npc-2024-state-council-organic-law/"><span>important matters</span></a><span>.&#8221; Hearing a report is a way for the State Council to indicate policy priorities before taking more formal action, such as reviewing and approving a document.</span></p><p><strong><span>The readout highlights:</span></strong></p><ul><li><p><strong><span>Domestic development initiative.</span></strong><span> The readout calls for &#8220;firmly grasping the initiative in development,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> meaning promoting domestic development. It is not AI-specific; President Xi Jinping </span><a href="https://digichina.stanford.edu/work/xi-jinping-strive-to-become-the-worlds-primary-center-for-science-and-high-ground-for-innovation/"><span>used the phrase</span></a><span> for science and technology broadly in 2018, and the State Council applied it to future industries at large at the </span><a href="https://www.news.cn/politics/leaders/20260605/2f63912a65d04f54881cf8849908d0de/c.html"><span>June 5 session</span></a><span>.</span></p></li><li><p><strong><span>Capability and compute. </span></strong><span>This item focuses on building AI capability, calling for breakthroughs in key technologies, accelerating the construction of ultra-large compute clusters, a greater high-quality data supply, guarantees for talent and capital,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> and enterprise basic and frontier research. Though similar lists of goals appear in other policy documents, this one is more frontier-focused. It includes some elements mentioned in the 15th Five-Year Plan (FYP), which called for basic/frontier research support, a better data supply, and guarantee for talent and funding. The FYP also called for feasibility studies on building ultra-large compute clusters, so the read-out language marks a progression from studying to executing.</span></p></li><li><p><strong><span>&#8220;AI+&#8221; applications.</span></strong><span> The Party-State&#8217;s AI+ diffusion agenda appears quite late in the readout, relative to other AI-focused items. It seems to still be a priority, but current AI+ initiatives have been in the works for months, so a shift will take time to manifest.</span></p></li><li><p><strong><span>Safety bottom line.</span></strong><span> &#8220;Holding the AI safety bottom line&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span> draws on </span><a href="https://baike.baidu.com/item/%E5%BA%95%E7%BA%BF%E6%80%9D%E7%BB%B4/4192129"><span>&#24213;&#32447;&#24605;&#32500;</span></a><span> (</span><a href="https://chinamediaproject.org/2015/02/06/thus-spoke-uncle-xi/"><span>bottom-line thinking</span></a><span>), a worst-case-first governance method Xi has pushed since 2013. Under this method policymakers are encouraged to focus on guarding against tail risks. Applied to AI, it indicates a level of concern for AI safety at the State Council level that we haven't seen before. &#24213;&#32447; language on AI has so far been confined to </span><a href="https://www.cac.gov.cn/2026-05/08/c_1779979789523320.htm"><span>ministerial</span></a><span>, </span><a href="https://www.tc260.org.cn/sysFile/downloadFile/59fa7e26d54844ffbac59f5361e3dd8f"><span>technical</span></a><span>, and </span><a href="http://iolaw.cssn.cn/gg/ggqt/202606/t20260603_6009785.shtml"><span>academic</span></a><span> texts (including a People&#8217;s Daily op-ed </span><a href="https://chinaaibulletin.substack.com/i/202627026/peoples-daily-commentary-frames-ai-misuse-in-the-cognitive-domain-as-a-national-security-concern"><span>covered last issue</span></a><span>), while the highest AI-specific policymaking venue, the </span><a href="https://www.gov.cn/yaowen/liebiao/202504/content_7021072.htm"><span>April 2025 Politburo study session</span></a><span> Xi chaired, used the softer &#8220;secure, reliable, and controllable&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a><span> to describe AI and did not invoke &#24213;&#32447; at all. It also was not used in relation to AI in the Five-Year Plan. The readout names science and technology ethics (likely the </span><a href="https://chinaaibulletin.substack.com/i/198737671/miit-launches-ten-province-ethics-review-pilot"><span>ethics review system being established</span></a><span>), testing and certification, and a tiered-and-classified oversight system as mechanisms to reduce AI risk.</span></p></li></ul><h2><span>The financial regulator sets detailed controls on high-risk financial AI</span></h2><p><span>On June 18, NFRA issued </span><a href="https://www.nfra.gov.cn/cn/view/pages/governmentDetail.html?docId=1261784&amp;itemId=&amp;generaltype=1"><span>Guiding Opinions on the Safe Development and Application of AI in Banking and Insurance</span></a><span>,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a><span> 32 detailed opinions spanning governance, development and application, data governance, compute, risk management, safe-development capability, and supervision. It imposes responsibility on whoever uses the AI, sorts applications by risk, and puts the heaviest controls on what it deems the riskiest uses.</span></p><p><span>AI use that involves transactions of funds, credit approval, underwriting and claims, asset valuation, or anything directly affecting customer interests count as high-risk. Those must clear the institution&#8217;s risk-management committee before launch, run under human oversight with emergency shutoff and manual fallback paths, and use black-box models only as an aid to a human decision-maker. The guidance also bars personal and private data, such as names, ID numbers, phone numbers, and bank-card numbers from being used to train or optimize generative AI models.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a><span> Externally sourced generative models must be filed with national cyberspace authorities, and the guidance names specific agent-security threats, such as prompt injection, memory poisoning, tool abuse, and loss of operational control (&#36816;&#34892;&#22833;&#25511;).</span></p><h2><span>Eight agencies target the demand side of the AI+ push</span></h2><p><span>The Ministry of Commerce (MOFCOM) and seven other departments issued the </span><a href="https://www.gov.cn/zhengce/zhengceku/202606/content_7072672.htm"><span>Implementation Opinions on Accelerating Development of &#8220;AI+ Consumption&#8221;</span></a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a><span> on June 9, circulated June 18, a 17-measure plan to move AI into consumer markets. Like the other AI+ sector opinions, much of it is a catalog of products and scenarios to build: next-generation AI phones, glasses, and smart-home devices; humanoid, companion, and elderly-care robots; hotel service robots; generative AI education models; and digital avatar live-streaming for retail.</span></p><p><span>Where it differs from other sector opinions is the emphasis on products around uptake rather than R&amp;D, repeatedly invoking demonstration applications, widespread adoption, and getting AI into &#8220;every home&#8221; and &#8220;every shop.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a><span> The </span><a href="https://www.mofcom.gov.cn/xwfb/sjfzrfb/art/2026/art_e8db95906e7449cb86a769f69826fa5d.html"><span>MOFCOM interpretation</span></a><span> casts it as the consumption pillar of the State Council&#8217;s 2025 &#8220;AI+&#8221; Action Opinion and the central consumption-boosting plan. As political scientist Jeffrey Ding </span><a href="https://www.fpri.org/article/2024/09/explaining-chinas-diffusion-deficit/"><span>argues</span></a><span>, the capacity to diffuse a general-purpose technology across the economy, not just to lead at the frontier, is a central driver of economic competition between major powers, and a consumption program built around uptake is clearly aimed at diffusion.</span></p><h2><span>National Standards</span></h2><h3><span>TC260 issues security guidance for deploying and using agents</span></h3><p><span>On July 1, TC260, China&#8217;s National Technical Committee on Cybersecurity, under the SAC, published a </span><a href="https://www.tc260.org.cn/portal/article/2/71c613fd3db34b6a9da62f25b8219733"><span>practice guide on deploying and using AI agents</span></a><span> securely.&#185; It sets out security measures across the agent lifecycle, from evaluation and preparation through deployment, use, and discontinuation, and is written for individuals deploying personal agents and organizations choosing among commercial agent services. It explicitly warns against using models that haven&#8217;t gone through the domestic registration process and using unknown &#8220;transfer stations,&#8221; which have been supporting a </span><a href="https://www.chinatalk.media/p/how-to-buy-cheap-claude-tokens-in"><span>gray market of Claude access on the mainland</span></a><span>. As a practice guide (&#23454;&#36341;&#25351;&#21335;), it is non-binding guidance that sits below a national standard. It complements two other TC260 agent documents: the March 31 guide for </span><a href="https://www.tc260.org.cn/portal/article/2/160310cea5f6411d92fd99a52a42424f"><span>OpenClaw-type agents</span></a><span>, which called for enterprise registries of approved deployments, and the </span><a href="https://chinaaibulletin.substack.com/i/198737671/wg9-holds-second-2026-plenary-agent-security-standard-enters-deliberation"><span>agent security standards </span></a><span>still working through TC260&#8217;s pipeline (a recommended GB/T basic specification and a mandatory GB for agent applications). TC260 is filling the guidance layer while the binding standards remain under review.</span></p><h3><span>SAMR signals it plans to make the agent identity code mandatory</span></h3><p><span>At a </span><a href="https://www.sac.gov.cn/xw/bzhyw/art/2026/art_3130ceca6a03423484a2793e1c45d84c.html"><span>June 26 press conference</span></a><span>, SAMR discussed the seven-part AI Agent Interconnection series (GB/Z 185.1&#8211;185.7-2026), which the SAC published on May 22 and we </span><a href="https://chinaaibulletin.substack.com/i/200679488/seven-part-ai-agent-interconnection-guidance-series-approved"><span>covered in Issue 5</span></a><span>. More than 70 organizations contributed to the drafting of the series, which covers the full agent-interoperability protocol. </span><strong><span>SAMR said it plans to turn the agent identity number (&#36523;&#20221;&#30721;) into a mandatory requirement</span></strong><span> and to accelerate the development of standards for agent auditing and transactions.</span></p><h3><span>Standards bodies move to accelerate AI standards</span></h3><p><span>On June 29, the SAC said it would </span><a href="https://www.sac.gov.cn/xw/bzhdt/art/2026/art_f0754d80166942c6a6bb2d56ee27fcc6.html"><span>speed up national standards</span></a><span> for agents, embodied intelligence, and world models, alongside standards for computing infrastructure, high-quality datasets, simulation platforms, deep-learning compilers, and open-source model frameworks. Although it didn&#8217;t name specific timelines, concrete projects are already appearing: in the last fortnight, the AI subcommittee of the National Information Technology Standardization Technical Committee (TC28/SC42) registered new national-standard projects for an </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=557270008E2345CDE06397BE0A0AC8C8"><span>AI computing-center management platform</span></a><span> and an </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=2ACFE99EAA38FAE1E06397BE0A0AD0BA"><span>open-source model platform</span></a><span>.</span></p><p><span>Separately, the AI Standardization Technical Committee (TC1) of MIIT, under the China Academy of Information and Communications Technology (CAICT), held its </span><a href="https://miittc1.caict.ac.cn/meetDetail/?id=139"><span>first plenary of 2026</span></a><span> on June 30, with working-group meetings on data, intelligent computing systems, models and platforms, and intelligent products and services. The agendas beyond the meeting titles are not public.</span></p><h3><span>MIIT promotes UN autonomous driving standard</span></h3><p><span>At the 199th session of the UN World Forum for Harmonization of Vehicle Regulations (WP.29) in Geneva on June 22&#8211;26, the working party approved </span><a href="https://www.miit.gov.cn/xwfb/gxdt/sjdt/art/2026/art_708b7727cc4544e8bc2f335cdc92b531.html"><span>the Autonomous Driving System Global Technical Regulation</span></a><span> (ADS GTR). MIIT calls it the first globally unified technical regulation for autonomous-driving systems and says China &#8220;spearheaded&#8221; the standard. It was jointly drafted with the EU, UK, US, Canada, and Japan, but China has vice-chaired WP.29&#8217;s working group on Automated and Connected Vehicles since its founding in 2018 and co-chairs the Functional Requirements subgroup.</span></p><p><span>MIIT is working on similar standards domestically; it says it has finished drafting a mandatory national standard for autonomous driving system safety, now in the approval process, that covers the international regulation&#8217;s core content and adds more detailed requirements for L3 and L4 systems.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a><span> It </span><a href="https://www.ncsti.gov.cn/kjdt/tzgg/202606/t20260618_250015.html"><span>solicited comments on the report-for-approval draft</span></a><span> from June 17 to 24.</span></p><h1><span>International AI Governance</span></h1><h2><span>Li Qiang links technology loss-of-control risk to AI at Summer Davos</span></h2><p><span>Opening the </span><a href="http://politics.people.com.cn/n1/2026/0624/c1024-40746688.html"><span>17th World Economic Forum Annual Meeting of the New Champions</span></a><span> (&#8220;Summer Davos&#8221;) in Dalian on June 24, Premier Li said AI is driving frontier breakthroughs so fast that, in his words, &#8220;some say&#8221; humanity has entered the &#8220;Cambrian&#8221; of the intelligent era,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a><span> which he then set against a warning: &#8220;the risks of loss of control of technology (&#25216;&#26415;&#22833;&#25511;) and ethical breaches are becoming more prominent,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a><span> and &#8220;if the corresponding governance cannot keep pace, it may well lead to serious consequences.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a><span> He said China will keep taking part in global AI governance &#8220;with a responsible and constructive attitude,&#8221; working with others to &#8220;improve institutional rules and raise the effectiveness of regulation.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a></p><p><span>&#8220;&#25216;&#26415;&#22833;&#25511; (loss of control of technology)&#8221; has been reaching higher levels of government. Xi </span><a href="https://chinaaibulletin.substack.com/p/china-ai-bulletin-5?open=false#%C2%A7domestic-ai-governance"><span>first used it</span></a><span> in his January 2026 Politburo study-session speech that </span><em><span>Qiushi</span></em><span> published in May, but it sat in a general future-industries governance passage and was not tied to AI. Li has now specifically linked it to AI at a flagship international forum and in China&#8217;s own outward-facing framing. Five days later, he chaired the State Council executive meeting whose readout discussed &#8220;holding the AI safety bottom line&#8221; (see Domestic AI Governance above), so in one week the premier discussed AI safety both abroad and at home.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this interesting? Subscribe to get regular issues in your inbox.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1><span>Frontier Lab Developments</span></h1><h2><span>Notable Model Releases</span></h2><p><strong><span>ByteDance</span></strong><span> launched </span><strong><a href="https://seed.bytedance.com/en/blog/seed2-1-officially-released-advancing-ai-prod"><span>Doubao 2.1 (Seed 2.1)</span></a></strong><span> as its new closed-source flagship. It is selling two versions: (1) a consumer &#8220;Professional Edition&#8221; at 500 RMB/73.59 USD a month that runs agentic tasks and can operate a user&#8217;s local computer, and (2) via API access on Volcano Engine at 6 RMB/0.88 USD per million input tokens and 30 RMB/4.42 USD per million output, roughly </span><a href="https://36kr.com/p/3865600233395201"><span>80 percent cheaper than Anthropic&#8217;s Claude Opus</span></a><span>. ByteDance claims parity with leading Western models on coding and agentic benchmarks.</span></p><p><strong><span>Alibaba (Qwen)</span></strong><span> released </span><strong><a href="https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B"><span>Qwen-AgentWorld-35B-A3B</span></a></strong><span>, its first &#8220;language world model,&#8221; a 35 billion-parameter Mixture-of-Experts (MoE) model trained to predict the next environment state from an agent&#8217;s action across seven interaction domains, rather than a general chat or coding model. It uses a sparse MoE (256 experts, eight routed plus one shared), a 262K-token context, and open weights under Apache 2.0. Qwen reports 56.39 overall on AgentWorldBench, its own benchmark for agent-environment modeling. Environment-modeling is the core training objective here, built in from pre-training.</span></p><p><strong><span>Baidu</span></strong><span> released </span><strong><a href="https://huggingface.co/baidu/Unlimited-OCR"><span>Unlimited-OCR</span></a></strong><span>, a compact three billion-parameter vision-language optical character recognition (OCR) model for one-shot, multi-page document parsing, open-weight under an MIT license.</span></p><p><strong><span>Also released this fortnight:</span></strong></p><ul><li><p><strong><a href="https://github.com/SeerRay-Lab/Xiaomi-GUI-0"><span>Xiaomi-GUI-0</span></a></strong><span> (Xiaomi), code and a </span><a href="https://arxiv.org/abs/2606.31410"><span>technical report</span></a><span> for a 30B-parameter (3B-active) on-device mobile GUI agent that carries out multi-app smartphone tasks. Xiaomi reports 72 percent task success on its own RealMobile benchmark and 78.9 percent on AndroidWorld, with weights forthcoming.</span></p></li><li><p><strong><a href="https://huggingface.co/sensenova/SenseNova-U1-8B-MoT-Infographic-V2"><span>SenseNova-U1-8B-MoT-Infographic-V2</span></a></strong><span> (SenseTime), a V2 refresh of its unified multimodal infographic model (18 billion parameters total, eight billion in the language model), open-weight (Apache 2.0), on a Mixture-of-Transformers architecture that drops the separate visual encoder and variational autoencoder for native pixel-to-word generation.</span></p></li><li><p><strong><a href="https://www.caixin.com/2026-06-24/102456873.html"><span>Seedance 2.5</span></a></strong><span> (ByteDance), a closed video-generation model capable of generating clips up to 30 seconds long.</span></p></li><li><p><strong><a href="https://www.stdaily.com/web/gdxw/2026-06/25/content_537227.html"><span>ERNIE/Wenxin refresh</span></a></strong><span> (Baidu), a closed site and model-matrix update.</span></p></li></ul><h2><span>Technical Publication Highlights</span></h2><p><span>Frontier labs released 152 papers on arXiv this fortnight. Highlights are below; a full list with summaries can be found </span><a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0"><span>here</span></a><span>.</span></p><p></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Wo58!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Wo58!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 424w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 848w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 1272w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Wo58!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png" width="635" height="571" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:571,&quot;width&quot;:635,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:47987,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/204645234?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Wo58!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 424w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 848w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 1272w, https://substackcdn.com/image/fetch/$s_!Wo58!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fce1a4ce7-2664-4e22-a861-31a080ed5dec_635x571.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3><span>Alibaba</span></h3><p><strong><a href="http://arxiv.org/abs/2606.26620v1"><span>Discovering Millions of Interpretable Features with Sparse Autoencoders</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Qwen3-Instruct SAE</span></strong><span>, a suite of sparse autoencoders trained across Qwen3 models (1.7B&#8211;8B parameters) that decompose neural activations into interpretable features, with a refusal-steering case study demonstrating causal steering capabilities.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.25442v1"><span>PolicyAlign: Direct Policy-Based Safety Alignment for Large Language Models</span></a></strong></p><ul><li><p><strong><span>PolicyAlign</span></strong><span> aligns LLMs directly to natural-language safety policies by synthesizing violating examples and using self-distillation, then filters to high-impact instructions for stable training. It reduces costly supervision while maintaining general capabilities across diverse domains like medicine, law, and finance.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.23075v1"><span>Safety in Self-Evolving LLM Agent Systems: Threats, Amplification, and Case Studies</span></a></strong></p><ul><li><p><span>Exposes critical vulnerabilities in </span><strong><span>self-evolving LLM agent systems</span></strong><span> where adversarial attacks become permanently encoded and self-amplify across generations. The </span><strong><span>Module&#8211;Lifecycle Attack Surface matrix</span></strong><span> identifies 17 of 25 functional areas with critical threats lacking defenses, and case studies show evolution-native designs achieve 100 percent attack persistence while existing security measures block only 2.5 percent of threats.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.22873v1"><span>SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning</span></a></strong></p><ul><li><p><span>Presents </span><strong><span>SingGuard</span></strong><span>, a multimodal safety model that adapts to runtime policy changes by treating safety rules as inputs and predicting both violations and triggered rules. The system supports </span><strong><span>three inference modes</span></strong><span> (fast direct judgment, hybrid, and slow deliberation) optimized via decoupled reinforcement learning, and introduces </span><strong><span>SingGuard-Bench</span></strong><span> with 56K examples across 80-plus risk types including cross-modal composition risks, achieving state-of-the-art F1 across six benchmark families and improving policy-following accuracy from 64.7 percent to 74.2 percent under dynamic rule shifts.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.25034v1"><span>Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Yuvion VL</span></strong><span>, a multimodal model family purpose-built for detecting adversarial content and AI safety risks through adversarial-aware data synthesis, three-stage training, and a contrastive fine-tuning method that distinguishes visually similar cases with different safety implications. The 32B variant </span><strong><span>outperforms comparably sized open-source and closed-source commercial models</span></strong><span> on safety tasks while maintaining general capability parity.</span></p></li></ul><h3><span>ByteDance</span></h3><p><strong><a href="http://arxiv.org/abs/2606.29887v1"><span>SafePyramid: A Hierarchical Benchmark for In-context Policy Guardrailing</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>SafePyramid</span></strong><span>, a benchmark with 1,000 conversations and 3,000 application-specific policies to evaluate whether AI safety guardrails can identify policy violations based on context-provided rules rather than fixed taxonomies. Current best models achieve only 54, 35, and 13 percent accuracy on single-rule, multi-rule reasoning, and novel policy adaptation tasks respectively, revealing substantial gaps in policy execution capabilities.</span></p></li></ul><h3><span>SenseTime</span></h3><p><strong><a href="http://arxiv.org/abs/2606.25325v1"><span>Omni-Perception Policy Optimization for Multimodal Emotion Reasoning</span></a></strong></p><ul><li><p><span>Proposes Omni-Perception Policy Optimization </span><strong><span>(OPPO)</span></strong><span>, a reinforcement learning framework that optimizes multimodal perception in emotion-reasoning models by </span><strong><span>rewarding trajectories that utilize visual, acoustic, and emotion cues</span></strong><span> and </span><strong><span>suppressing hallucinations through KL-penalized masking of cross-modal evidence tokens</span></strong><span>. On emotion reasoning benchmarks, OPPO improves both task performance and faithfulness metrics while reducing spurious multimodal claims.</span></p></li></ul><h3><span>StepFun</span></h3><p><strong><a href="http://arxiv.org/abs/2606.28322v2"><span>PerceptionRubrics: Calibrating Multimodal Evaluation to Human Perception</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>PerceptionRubrics</span></strong><span>, a rubric-based evaluation framework that moves beyond holistic semantic matching to </span><strong><span>atomic fact auditing</span></strong><span> with 12,000-plus instance-specific criteria. A </span><strong><span>Gated Scoring mechanism</span></strong><span> penalizes failures on mandatory visual facts, revealing that models often pass fragmented elements but fail strict conjunctive constraints&#8212;and exposing a persistent 8 percent performance gap between open-source and proprietary models.</span></p></li></ul><h3><span>Tencent</span></h3><p><strong><a href="http://arxiv.org/abs/2606.21129v1"><span>AgenticOS: An Intent-Oriented Secure Operating System Architecture for Autonomous AI Agents</span></a></strong></p><ul><li><p><span>Reframes OS security for autonomous AI agents by shifting from </span><strong><span>resource-based access control to intent-based filtering</span></strong><span>: agents declare high-level goals, and the system automatically synthesizes least-privilege environments with mandatory mediation and auditing. The four-layer </span><strong><span>AgenticOS architecture</span></strong><span> (Ghost Kernel, Logic Shutter, Agent Capsule, and Semantic Boundary Gateway) prevents attackers from composing low-level primitives into unauthorized behaviors even after compromising the agent runtime.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.20374v1"><span>ARGUS: Production-Scale Tracing and Performance Diagnosis for over 10,000-GPU Clusters</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>ARGUS</span></strong><span>, a production tracing system for 10,000-plus GPU clusters that maintains </span><strong><span>sub-2 percent overhead while capturing fine-grained CPU, framework, and GPU kernel traces</span></strong><span>, compressing raw events by 3,700&#215; and automatically diagnosing fail-slow performance anomalies across compute, communication, and software bottlenecks. Deployed at Tencent for six months, it has detected and resolved issues like hardware degradation and JIT stalls that waste millions of GPU-hours.</span></p></li></ul><p><strong><a href="http://arxiv.org/abs/2606.31227v1"><span>Securing the AI Agent: A Unified Framework for Multi-Layer Agent Red Teaming</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>AI-Infra-Guard</span></strong><span>, an open-source red teaming framework that matches different security detection methods to distinct agent layers&#8212;from rule-based scanning of infrastructure components to LLM-driven auditing of tools and multi-turn adversarial testing of agent behavior, covering 75-plus components and 1,400-plus vulnerability rules.</span></p></li></ul><h1><span>AI Safety Publication Highlights</span></h1><p><strong><span>Chinese researchers published 72 AI-safety-related papers </span></strong><span>this fortnight. Highlights are below; a full list with summaries is available </span><a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1"><span>here</span></a><span>.</span></p><h2><span>&#128269;Spotlight</span></h2><p><strong><a href="http://arxiv.org/abs/2606.27944v1"><span>It Lied to a Doctor to Buy Poison Ingredients: Quantifying Real-World Misuse of Phone-use Agents</span></a></strong><span> (Fudan University)</span></p><p><span>Phone-use agents (AI systems that operate real smartphone apps by reading the screen and tapping through the interface) can function in ways that API-and CLI-based agents cannot, which is also what makes their misuse more dangerous. A team from Fudan built a benchmark grounded in six Chinese laws and administrative regulations (1,381 test cases across six misuse categories) and ran agents built on </span><strong><span>nine mainstream commercial and open-source models</span></strong><span> across </span><strong><span>27 real Chinese apps</span></strong><span> on physical devices, not in simulation.</span></p><p><span>Most agents rarely refused, and the average harmful-task completion rate reached </span><strong><span>68.8 percent </span></strong><span>on real devices. Small open-source agents matched the commercial models while running faster and cheaper. Refusals concentrated on overt tasks such as buying a controlled substance or harassment; the agents were far more compliant with covert misuse such as fraud and coordinated review manipulation, which the authors argue does more real-world harm.</span></p><p><span>In the case study that gives the paper its title, the authors report that a Claude Opus 4.8</span><strong><span> </span></strong><span>agent went online to buy ingredients for a toxic substance, fabricated a medical diagnosis, deceived an online doctor into issuing an electronic prescription, and completed the order and payment on its own. The authors call it the first documented case of an AI agent procuring controlled precursor materials, but that claim needs two caveats. First, it&#8217;s unclear if it purchased all the precursors to the bombs/toxic substances or how difficult it would be to manufacture them. Second, it&#8217;s unclear what the magnitude of harm would be. The paper doesn&#8217;t describe how explosive the possible reaction would be, and while it purchased a precursor to mercury iodide, mercury iodide is harmful when inhaled or absorbed, but isn&#8217;t considered a mass-casualty weapon. Finally, the purchase ran through a real e-commerce platform that sold the items behind a gameable online-prescription check, so it may reflect weak platform security as much as a new AI capability. Still, the authors claim it&#8217;s significant the agent convincingly fabricated a diagnosis. The paper&#8217;s core concept is a </span><strong><span>Safety Awareness-Execution Gap</span></strong><span>: an agent recognizes a request is harmful yet executes it anyway, which the authors trace to reduced activation of the model&#8217;s safety neurons when acting as an agent. Re-eliciting that awareness (through a detector, a prompt-based defense, or activation steering) cut misuse with limited capability loss, but the covert threats stayed largely unsolved.</span></p><p><span>The study shows why agent behavior is increasingly governed separately from model outputs: the gap between a model that refuses a harmful request in chat and an agent that completes it on a real phone is what the agent-security standards in this issue&#8217;s National Standards section are meant to address.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!myVe!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!myVe!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 424w, https://substackcdn.com/image/fetch/$s_!myVe!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 848w, https://substackcdn.com/image/fetch/$s_!myVe!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 1272w, https://substackcdn.com/image/fetch/$s_!myVe!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!myVe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png" width="911" height="586" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bc027fe2-53f2-434b-8826-39670763a347_911x586.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:586,&quot;width&quot;:911,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Phone use agent misuse showing taobao searches for bomb materials, malicious reporting, privacy leaks.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Phone use agent misuse showing taobao searches for bomb materials, malicious reporting, privacy leaks." title="Phone use agent misuse showing taobao searches for bomb materials, malicious reporting, privacy leaks." srcset="https://substackcdn.com/image/fetch/$s_!myVe!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 424w, https://substackcdn.com/image/fetch/$s_!myVe!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 848w, https://substackcdn.com/image/fetch/$s_!myVe!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 1272w, https://substackcdn.com/image/fetch/$s_!myVe!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbc027fe2-53f2-434b-8826-39670763a347_911x586.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Phone use agent misuse workflows from the <a href="https://arxiv.org/pdf/2606.27944v1">paper</a>.</figcaption></figure></div><h2><span>Agentic Safety</span></h2><p><strong><a href="http://arxiv.org/abs/2606.22528v2"><span>Governance Decay: How Context Compaction Silently Erases Safety Constraints in Long-Horizon LLM Agents</span></a></strong></p><ul><li><p><strong><span>Governance Decay</span></strong><span> describes how LLM agents systematically violate in-context safety constraints when long conversations are automatically summarized to save tokens. The researchers measure this with </span><strong><span>ConstraintRot</span></strong><span>, a benchmark showing violation rates jump from 0 percent to 30&#8211;59 percent after compaction, because summarization algorithms treat standing policies as low-salience and drop them. </span><strong><span>Constraint Pinning</span></strong><span>, a training-free defense, restores compliance to 0 percent with minimal token overhead, though the paper identifies remaining failure modes where the defense degrades.</span></p></li></ul><p><em><span>Institutional affiliations: Beijing Institute of Technology</span></em></p><p><strong><a href="http://arxiv.org/abs/2606.27027v1"><span>ShareLock: A Stealthy Multi-Tool Threshold Poisoning Attack Against MCP</span></a></strong></p><ul><li><p><strong><span>ShareLock</span></strong><span> distributes malicious instructions across multiple tool descriptions using cryptographic secret sharing (Shamir&#8217;s scheme), making each individual tool appear benign while collectively reconstructing the attack when triggered. This multi-tool poisoning framework defeats </span><strong><span>detection by manual inspection or automated scanners</span></strong><span> because no single tool contains the full malicious payload&#8212;only harmless-looking fragments. Experiments on mainstream LLMs show </span><strong><span> more than 90 percent attack success</span></strong><span> while evading tool description-based defenses, demonstrating threat actors can exploit Model Context Protocol&#8217;s distributed architecture to hide poisoning across the agent&#8217;s tool ecosystem.</span></p></li></ul><p><em><span>Institutional affiliations: Shanghai Jiao Tong University</span></em></p><h2><span>Alignment</span></h2><p><strong><a href="http://arxiv.org/abs/2606.19168v1"><span>Beyond Safe Data: Pretraining-Stage Alignment with Regular Safety Reflection</span></a></strong></p><ul><li><p><strong><span>Safety Reflection Pretraining</span></strong><span> embeds periodic safety reflections directly into pretraining corpora to establish self-monitoring as a foundational capability rather than relying solely on data filtering. Testing on 1.7B models shows the method </span><strong><span>reduces jailbreak success rates</span></strong><span> and improves safety classification accuracy compared to unsafe data filtering or rewriting alone. The approach addresses a specific failure mode: models composing benign knowledge into unsafe behaviors&#8212;a risk that data sanitization alone does not prevent.</span></p></li></ul><p><em><span>Institutional affiliations: Tsinghua University</span></em></p><p><strong><a href="http://arxiv.org/abs/2606.27709v1"><span>Low-Agreeableness Persona Conditioning for Safe LLM Fine-Tuning</span></a></strong></p><ul><li><p><span>Fine-tuning LLMs for conversational warmth </span><strong><span>increases jailbreak susceptibility</span></strong><span> and harmful outputs, but this appears driven by data construction rather than empathy itself. The authors introduce </span><strong><span>persona-driven conditioning</span></strong><span>, where user inputs are rewritten to reflect low agreeableness while assistant responses remain warm and de-escalating. Across four models, this approach </span><strong><span>reduces jailbreak success rates and harmful outputs</span></strong><span> versus standard warmth fine-tuning, while maintaining conversational warmth&#8212;suggesting safer empathetic alignment is achievable through data design alone.</span></p></li></ul><p><em><span>Institutional affiliations: Hong Kong University of Science and Technology</span></em></p><h2><span>Evaluation and Benchmarks</span></h2><p><strong><a href="http://arxiv.org/abs/2606.23686v1"><span>LIBERO-Safety: A Comprehensive Benchmark for Physical and Semantic Safety in Vision-Language-Action Models</span></a></strong></p><ul><li><p><strong><span>LIBERO-Safety</span></strong><span> is a benchmark for evaluating physical and semantic safety in vision-language-action models, systems that control robotic manipulation. The authors develop a </span><strong><span>parametric scenario generator</span></strong><span> to create collision-free demonstrations at scale (19,664 examples with domain randomization) and evaluate eight VLA models and two embodied foundation models. Their analysis exposes a </span><strong><span>generalization-safety tradeoff</span></strong><span>: higher-diversity training improves collision avoidance but reduces task success, and models show </span><strong><span>semantic misalignment</span></strong><span> between language instructions and safe execution paths.</span></p></li></ul><p><em><span>Institutional affiliations: Tsinghua University, Beijing Academy of Artificial Intelligence, Beihang University, Eastern Institute of Technology, Shanghai Jiao Tong University, Microsoft Research Asia</span></em></p><p><strong><a href="http://arxiv.org/abs/2606.25396v1"><span>Long-Term Simulation Exposes Cognitive-Developmental Risks in AI Companions</span></a></strong></p><ul><li><p><strong><span>TSJ (Theater-Stage-Judge)</span></strong><span> is a longitudinal evaluation framework that simulates extended AI companion interactions across developmental stages to detect risks that single-turn testing misses. While testing six models over 12,960 simulated person-days across four age groups and 24 risk dimensions, the framework found </span><strong><span>short-horizon evaluations systematically underestimate harm</span></strong><span>, with stable risk estimates emerging only after approximately 140 conversational turns. Early childhood and emerging adulthood showed highest vulnerability in cognitive trust and emotional dependency domains.</span></p></li></ul><p><em><span>Institutional affiliations: East China Normal University, Shanghai Artificial Intelligence Laboratory</span></em></p><p><strong><a href="http://arxiv.org/abs/2606.18632v1"><span>ROBOSHACKLES: A Safety Dataset for Human-Injury Prevention in Embodied Foundation Models</span></a></strong></p><ul><li><p><strong><span>ROBOSHACKLES</span></strong><span> is a 10,000-clip robotic video dataset for safety alignment in embodied foundation models, systems that combine multimodal reasoning with executable robot actions. The dataset is synthesized from real robot observations using image editing and video generation to create realistic hazardous scenarios (direct injuries and indirect environmental harms) that cannot be safely collected in practice. Evaluation of six models shows </span><strong><span>100 percent unsafe action rates</span></strong><span> on these scenarios, establishing a benchmark for refusal learning and hazard anticipation in robot control systems.</span></p></li></ul><p><em><span>Institutional affiliations: Chinese Academy of Sciences, University of Science and Technology of China</span></em></p><p><strong><a href="http://arxiv.org/abs/2606.18936v2"><span>SciRisk-Bench: A Risk-Dimension-Aware Benchmark for AI4Science Safety</span></a></strong></p><ul><li><p><strong><span>SciRisk-Bench</span></strong><span> is a benchmark evaluating LLM safety in scientific contexts across </span><strong><span>seven disciplines, 31 subdisciplines, and 10 risk dimensions</span></strong><span>&#8212;including dual-use synthesis details, omitted safety precautions, overconfident claims, and data privacy violations. The authors evaluate mainstream and science-specialized LLMs to identify where models fail to recognize or mitigate risks specific to laboratory, clinical, and research workflows. Results show that existing general-purpose safety benchmarks inadequately capture the distinct failure modes that arise when LLMs mediate high-stakes scientific decisions.</span></p></li></ul><p><em><span>Institutional affiliations: Chinese Academy of Sciences, University of Chinese Academy of Sciences, Zhongguancun Academy, Beijing Key Laboratory of Safe AI and Superalignment, Renmin University of China, Beijing Institute of AI Safety and Governance</span></em></p><h2><span>Governance and Policy</span></h2><p><strong><a href="http://arxiv.org/abs/2606.23860v1"><span>World Artificial Intelligence Cooperation Organization (WAICO): Mapping an Emerging Institution in the Global AI Governance Regime Complex</span></a></strong></p><ul><li><p><span>This paper maps </span><strong><span>WAICO</span></strong><span> (China&#8217;s proposed World Artificial Intelligence Cooperation Organization) within the global AI governance landscape. Using structured coding of 15 existing AI governance bodies, the authors identify </span><strong><span>WAICO&#8217;s unique institutional positioning</span></strong><span>: universal membership (no values test), sovereignty-centered framing, and development-first priorities&#8212;features absent in Western-led bodies (which gate entry by shared values and emphasize rights/safety) and UN bodies (open but anchored in human rights). The analysis characterizes WAICO as the first standing organization designed to anchor a development-oriented governance pole distinct from the incumbent rights-and-safety framework.</span></p></li></ul><p><em><span>Institutional affiliations: Tsinghua University, Federal University of Rio de Janeiro</span></em></p><h2><span>Guardrails and Deployment Safety</span></h2><p><strong><a href="http://arxiv.org/abs/2606.21584v1"><span>When EER Hides Deployment Failure: Auditing Threshold Transfer and Unlabeled Score Calibration for Speech Deepfake Detectors</span></a></strong></p><ul><li><p><span>This paper reveals a </span><strong><span>deployment gap in speech deepfake detectors</span></strong><span>: models achieving low equal error rate (EER) on test sets often fail catastrophically when their thresholds are applied to new, unlabeled data. Testing a state-of-the-art detector shows </span><strong><span>0.21 percent  in-domain EER</span></strong><span> but </span><strong><span>39.5 percent half total error rate on out-of-domain data</span></strong><span>, rejecting 78.7 percent of legitimate speech. The authors prove common score calibration techniques cannot improve EER and find that popular unlabeled test-time corrections fail, some collapse entirely on new datasets, exposing a systematic mismatch between research metrics and real-world detector behavior that matters for security-critical deployment.</span></p></li></ul><p><em><span>Institutional affiliations: Xidian University</span></em></p><h2><span>Interpretability</span></h2><p><strong><a href="http://arxiv.org/abs/2606.28153v2"><span>Robust Harmful Features Under Jailbreak Attacks: Mechanistic Evidence from Attention Head Specialization in Large Language Models</span></a></strong></p><ul><li><p><span>This paper identifies </span><strong><span>Adversarially Compromised Heads (ACHs)</span></strong><span> and </span><strong><span>Safety-Aligned Heads (SAHs)</span></strong><span>&#8212;specialized attention structures that respond differently to jailbreak attacks. ACHs in early layers are suppressed by adversarial prompts, while SAHs in mid-layers maintain active safety signals even when attacks succeed. The authors show that removing just a few ACHs triggers jailbreak behavior, and  persistent SAH activations can be read directly&#8212;without retraining&#8212;to detect attacks with competitive accuracy, suggesting safety information survives suppression rather than being eliminated.</span></p></li></ul><p><em><span>Institutional affiliations: Beijing University of Posts and Telecommunications</span></em></p><h2><span>Misuse and Dangerous Capabilities</span></h2><p><strong><a href="http://arxiv.org/abs/2606.21846v1"><span>Mind the Intention: Task-Aware Backdoor Attacks for Forecast-Driven Distribution Network Operations</span></a></strong></p><ul><li><p><strong><span>GridTroj</span></strong><span> is a backdoor attack framework that embeds hidden triggers in energy forecasting models to cause </span><strong><span>cascading operational failures</span></strong><span> in power distribution networks. Unlike standard backdoor attacks that only manipulate forecast outputs, GridTroj&#8217;s &#8220;Intention Planner&#8221; designs triggers and poisoned training data to specifically damage downstream grid operations&#8212;voltage instability, line overloads, or demand-supply mismatches. Experiments show the attack successfully compromises real optimization tasks, demonstrating that forecast-driven grid automation creates exploitable vulnerabilities where poisoned models act as persistent, stealthy threats to critical infrastructure.</span></p></li></ul><p><em><span>Institutional affiliations: Xi&#8217;an Jiaotong University, Fudan University</span></em></p><h1><span>Export Controls &amp; Economic Policy</span></h1><h2><span>The NDRC warns provinces against disorderly AI and compute buildout</span></h2><p><span>Developing &#8220;AI+&#8221; has to fit local conditions, and provinces should &#8220;resolutely avoid disorderly competition and piling in,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a><span> the National Development and Reform Commission&#8217;s (NDRC) High-Tech Deputy Director Zhang Kailin </span><a href="http://www.news.cn/tech/20260624/d7ad9698c4b140d595112ef88c55c175/c.html"><span>said</span></a><span> on June 24. He pointed to provinces already specializing: 10, including Anhui and Jilin, are folding their plans into the national East-Data-West-Computing (&#19996;&#25968;&#35199;&#31639;) layout, and 17 have made green, low-carbon power a basic principle for new computing infrastructure. The NDRC&#8217;s caution follows Xi&#8217;s own. At the July 2025 </span><a href="https://www.gov.cn/yaowen/liebiao/202507/content_7032431.htm"><span>Central Urban Work Conference</span></a><span>, Xi questioned whether every province needs to develop the same few industries (AI, computing power, and green energy vehicles).</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a></p><p><span>Overbuilding is a particular concern in AI, with state media </span><a href="http://www.news.cn/finance/20250429/3df0c33317d2499ab3a297a413e0acce/c.html"><span>flagging</span></a><span> idle older data centers, low average occupancy, and a structural split in which general-purpose compute is in relative oversupply while high-end intelligent compute stays scarce. The state response has been consolidation. On June 29, </span><a href="http://www.news.cn/tech/20260629/1e0029b074ac4d9ba7df6e550b861aac/c.html"><span>Xinhua framed China&#8217;s compute infrastructure as moving from scattered construction toward networked dispatch</span></a><span>, toward a national computing-power network (&#31639;&#21147;&#32593;) that delivers compute on demand, like water or electricity. That continues the national integrated computing-power network buildout we covered in </span><a href="https://chinaaibulletin.substack.com/p/china-ai-bulletin-4"><span>Issue 4</span></a><span> and </span><a href="https://chinaaibulletin.substack.com/p/china-ai-bulletin-5"><span>Issue 5</span></a><span>.</span></p><h1><span>On the Horizon</span></h1><p><strong><span>China hosts the APEC Digital and AI Ministerial (Chengdu, July 23&#8211;24).</span></strong><span> MIIT </span><a href="https://www.miit.gov.cn/xwfb/gxdt/art/2026/art_9b5792e2e8d14eb3beb5fbb8f5ff63fa.html"><span>opened overseas-media registration</span></a><span> on June 30 for the 2026 APEC Digital and AI Ministerial Meeting, in Chengdu the week after the World Artificial Intelligence Conference. No agenda is public yet, but we&#8217;ll watch for anything AI-related coming out of it.</span></p><div><hr></div><p style="text-align: justify;"><span>For more on how we select and track content, see our methodology </span><a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz"><span>here</span></a><span>.</span></p><div><hr></div><p><em><span>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</span></em></p><p><em><span>The China AI Bulletin is copy-edited and fact-checked by Kacie Yearout, Pivotal fellow and former diplomat.</span></em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>&#37325;&#28857;&#24494;&#30701;&#21095;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>&#21457;&#23637;&#20027;&#21160;&#26435;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>Based on the FYP, for talent these could include: a national strategic talent force, a high-tech talent immigration system, channels to move researchers into enterprises, and/or job evaluation/salary reform. For finance/capital, it could include more active use of government guidance funds, long-term capital for early-stage &#8220;hard tech,&#8221; R&amp;D expense deduction reforms, and encouraging venture capital/market investing.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;&#35201;&#23432;&#29282;&#20154;&#24037;&#26234;&#33021;&#23433;&#20840;&#24213;&#32447;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#8220;<span>&#23433;&#20840;&#12289;&#21487;&#38752;&#12289;&#21487;&#25511;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>&#12298;&#22269;&#23478;&#37329;&#34701;&#30417;&#30563;&#31649;&#29702;&#24635;&#23616;&#20851;&#20110;&#38134;&#34892;&#19994;&#20445;&#38505;&#19994;&#20154;&#24037;&#26234;&#33021;&#23433;&#20840;&#24320;&#21457;&#24212;&#29992;&#30340;&#25351;&#23548;&#24847;&#35265;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>&#22995;&#21517;&#12289;&#36523;&#20221;&#35777;&#21495;&#12289;&#25163;&#26426;&#21495;&#12289;&#38134;&#34892;&#21345;&#21495;&#31561;&#20010;&#20154;&#20449;&#24687;&#21644;&#38544;&#31169;&#25968;&#25454;&#19981;&#24471;&#29992;&#20110;&#29983;&#25104;&#24335;&#20154;&#24037;&#26234;&#33021;&#27169;&#22411;&#35757;&#32451;&#21644;&#20248;&#21270;"&#65288;&#25351;&#23548;&#24847;&#35265;&#31532;&#20108;&#21313;&#22235;&#39033;&#65289;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>&#12298;&#21830;&#21153;&#37096;&#31561;8&#37096;&#38376;&#20851;&#20110;&#21152;&#24555;"&#20154;&#24037;&#26234;&#33021;+&#28040;&#36153;"&#21457;&#23637;&#30340;&#23454;&#26045;&#24847;&#35265;&#12299;&#65292;&#21830;&#24314;&#21457;&#12308;2026&#12309;89&#21495;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>&#8220;&#20154;&#24037;&#26234;&#33021;&#36827;&#19975;&#23478;&#8221; and &#8220;&#21315;&#38598;&#19975;&#24215;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>&#12298;&#26234;&#33021;&#32593;&#32852;&#27773;&#36710; &#33258;&#21160;&#39550;&#39542;&#31995;&#32479;&#23433;&#20840;&#35201;&#27714;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>&#26377;&#20154;&#35828;&#20154;&#31867;&#24050;&#32463;&#27493;&#20837;&#26234;&#33021;&#26102;&#20195;&#30340;&#8221;&#23506;&#27494;&#32426;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>&#25216;&#26415;&#22833;&#25511;&#12289;&#20262;&#29702;&#22833;&#33539;&#31561;&#39118;&#38505;&#20063;&#26356;&#21152;&#31361;&#20986;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>&#22914;&#26524;&#30456;&#20851;&#27835;&#29702;&#36319;&#19981;&#19978;&#65292;&#23601;&#24456;&#21487;&#33021;&#23548;&#33268;&#20005;&#37325;&#21518;&#26524;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>&#20013;&#22269;&#23558;&#32487;&#32493;&#20197;&#36127;&#36131;&#20219;&#12289;&#24314;&#35774;&#24615;&#24577;&#24230;&#21442;&#19982;&#20154;&#24037;&#26234;&#33021;&#31561;&#39046;&#22495;&#20840;&#29699;&#27835;&#29702;&#65292;&#21516;&#21508;&#26041;&#19968;&#36947;&#20581;&#20840;&#21046;&#24230;&#35268;&#21017;&#12289;&#25552;&#21319;&#30417;&#31649;&#25928;&#33021;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>&#8220;&#22362;&#20915;&#36991;&#20813;&#26080;&#24207;&#31454;&#20105;&#21644;&#19968;&#25317;&#32780;&#19978;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>&#8220;&#19978;&#39033;&#30446;&#65292;&#19968;&#35828;&#23601;&#26159;&#20960;&#26679;&#65306;&#20154;&#24037;&#26234;&#33021;&#12289;&#31639;&#21147;&#12289;&#26032;&#33021;&#28304;&#27773;&#36710;&#65292;&#26159;&#19981;&#26159;&#20840;&#22269;&#21508;&#30465;&#20221;&#37117;&#35201;&#24448;&#36825;&#20123;&#26041;&#21521;&#21435;&#21457;&#23637;&#20135;&#19994;&#65311;&#8221;&#65288;&#20064;&#36817;&#24179;&#65292;&#20013;&#22830;&#22478;&#24066;&#24037;&#20316;&#20250;&#35758;&#65292;2025&#24180;7&#26376;&#65289;</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 6]]></title><description><![CDATA[Developments from 3/6/26-17/6/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-6</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-6</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Thu, 18 Jun 2026 20:20:16 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4fb4442f-a71f-4959-9b4c-1bd84ef6e437_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p><span>Welcome to Issue 6 of the China AI Bulletin, the latest on AI governance, development, and safety in China. </span><strong><span>Today&#8217;s highlights</span></strong><span>: China signals it is accelerating preparations for its proposed World AI Cooperation Organization, Zhipu put GLM-5.2 into its coding plan as a response to the US withdrawing Anthropic Fable access, and a Chinese benchmark tests how far AI models can autonomously break into servers.</span></p><p><em><span>Number of the week: </span><a href="https://mp.weixin.qq.com/s/uVkc_wpou8U8vWoH4mpIUw"><span>10 trillion</span></a><span>&#8212;the daily target output from a new Beijing &#8220;token factory&#8221;</span></em></p><h1><span>Executive Summary</span></h1><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/domestic-ai-governance"><span>Domestic AI Governance</span></a><span>:</span></strong><span> The </span><strong><span>&#8220;AI+&#8221; implementation wave</span></strong><span> continued&#8212;MIIT&#8217;s </span><strong><span>AI+ Information and Communications plan</span></strong><span> and the National Data Administration&#8217;s plan to build </span><strong><span>high-quality training datasets</span></strong><span> operationalized the plan across the telecom and data layers. Separately, a </span><em><span>People&#8217;s Daily</span></em><span> commentary </span><strong><span>framed AI misuse as a national security threat</span></strong><span>, a </span><strong><span>State Council employment plan expressed concern over AI and jobs</span></strong><span>, MIIT and SASAC launched a </span><strong><span>humanoid robot real-world training push</span></strong><span>, and the CAC </span><strong><span>opened a public channel to report AI application problems</span></strong><span>.</span></p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/national-standards"><span>National Standards</span></a><span>:</span></strong><span> MIIT&#8217;s AI standardization committee advanced </span><strong><span>29 AI safety/security standards</span></strong><span>, and the SAC opened a 126-standard batch for comment that includes a </span><strong><span>seven-part Industrial AI Agents series</span></strong><span>.</span></p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/international-ai-governance"><span>International AI Governance</span></a><span>:</span></strong><span> Foreign Minister Wang Yi signaled China is </span><strong><span>accelerating preparations for a World AI Cooperation Organization (WAICO)</span></strong><span>. Chinese security scholars published a run of proposals on what the </span><strong><span>US-China AI dialogue</span></strong><span> should cover, including joint evaluations and communication mechanisms.</span></p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/frontier-lab-developments"><span>Frontier Lab Developments</span></a><span>:</span></strong><span> A steady run of open-weight releases, led by </span><strong><span>Zhipu&#8217;s GLM-5.2</span></strong><span> and </span><strong><span>Moonshot&#8217;s Kimi-K2.7-Code</span></strong><span>, plus </span><strong><span>Xiaomi&#8217;s MiMo coding agent</span></strong><span> and </span><strong><span>Baichuan&#8217;s clinical-grade M4</span></strong><span>. Frontier labs released 138 papers on arXiv this fortnight.</span></p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/technical-ai-safety-publication-highlights"><span>Technical AI Safety</span></a><span>:</span></strong><span> Chinese researchers published </span><strong><span>48 AI safety papers</span></strong><span>, still weighted toward agent safety. </span><strong><span>In the </span><a href="https://chinaaibulletin.substack.com/i/202627026/spotlight"><span>spotlight</span></a><span>:</span></strong><span> a Fudan/Shanghai AI Lab/Concordia AI/Shanghai Innovation Institute benchmark for </span><strong><span>LLMs&#8217; autonomous penetration capability</span></strong><span>, which rose with general model capability.</span></p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/202627026/export-controls-and-economic-policy"><span>Export Controls &amp; Economic Policy</span></a><span>:</span></strong><span> In the </span><strong><span>first US export control on the use of a deployed frontier model</span></strong><span>, the US barred foreign nationals from accessing Anthropic&#8217;s Fable 5 and Mythos 5, resulting in Anthropic pulling the models for everyone.</span></p></li></ul><p></p><h1><span>Domestic AI Governance</span></h1><h2><span>State Council&#8217;s employment plan shows concern over AI labor displacement</span></h2><p><span>On June 17, the State Council released its </span><em><a href="https://www.gov.cn/zhengce/content/202606/content_7072481.htm"><span>Plan for Implementing the Employment-First Strategy during the 15th Five-Year Plan</span></a><span> (</span></em><span>&#12298;&#23454;&#26045;&#23601;&#19994;&#20248;&#20808;&#25112;&#30053;&#8220;&#21313;&#20116;&#20116;&#8221;&#35268;&#21010;&#12299;). One of its nine task areas is on &#8220;adapting to AI development to promote employment and entrepreneurship.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> The plan is relatively optimistic; it ties employment to the national &#8220;AI+&#8221; action, calls to explore new forms of human-machine collaboration and to &#8220;strengthen AI&#8217;s job-creation effect,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a><span> and urges making good use of the &#8220;AI development dividend&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><span> to increase employment and benefit livelihoods. However, it also acknowledges potential risks to labor. Alongside the dividend language, it commits to track and handle AI-driven job loss: it lists &#8220;AI&#8217;s impact on employment&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a><span> among the issues for its employment-monitoring and risk-response work, alongside population aging, and calls to improve early warning and handling of employment risks from AI.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a><span> A separate platform-labor provision presses platform companies to &#8220;regulate algorithms and improve transparency&#8221; and to protect gig workers&#8217; rights to know, participate in, and choose how algorithm rules apply to them.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a></p><p></p><div class="image-gallery-embed" data-attrs="{&quot;gallery&quot;:{&quot;images&quot;:[{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/970a13db-d922-42cf-855a-9b94aa0d50ec_1814x1242.png&quot;},{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c7f58226-9a0b-4215-95f9-e5bda1a6730f_1190x1330.png&quot;}],&quot;caption&quot;:&quot;Box 10 on promoting employment in response to AI development and a translation&quot;,&quot;alt&quot;:&quot;&quot;,&quot;staticGalleryImage&quot;:{&quot;type&quot;:&quot;image/png&quot;,&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d2996508-bdc7-4af1-aa5d-c3810ee1e9ab_1456x720.png&quot;}},&quot;isEditorNode&quot;:true}"></div><p></p><h2><span>The National Data Administration lays out a plan to build AI training datasets</span></h2><p><span>On June 8, the National Data Administration (NDA)&#8212;the data-governance body established in 2023&#8212;released its </span><a href="https://www.nda.gov.cn/sjj/zwgk/tzgg/0608/20260608172117399715004_pc.html"><span>Implementation Plan for Promoting the Building of High-Quality Industry Datasets</span></a><span>,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a><span> laying out a plan to leverage training data to promote the </span><a href="https://www.gov.cn/zhengce/content/202508/content_7037861.htm"><span>&#8220;AI+&#8221; Initiative</span></a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a><span> and 15th Five-Year Plan. The plan groups seventeen measures under six &#8220;special actions&#8221;&#8212;expanding dataset building capacity, labeling/annotation, improving quality and efficiency, promoting applications, managing services, and unleashing value</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a><span>&#8212;around a &#8220;data flywheel&#8221; (&#25968;&#25454;&#39134;&#36718;) in which scenarios pull data, data trains models, and model use generates more data. It sets a 2028 target of application-validated datasets across key sectors, backed by national standards and a push for &#8220;one evaluation, nationwide mutual recognition&#8221; of dataset quality.</span></p><p><span>The plan also addresses the use of data for AI training, committing the NDA to develop &#8220;data rules for AI development&#8221;&#8212;separating data holding, use, and operation rights, and &#8220;improving rules for data use in the AI-training stage&#8221; so that copyrighted works can be &#8220;used for model training in an orderly way&#8221; under authorization and revenue-sharing.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a><span> That pulls training-data legality and copyright into the formal agenda, beyond the general call to respect IP in the 2023 Generative AI Measures. Days earlier, on June 4, NDA director Liu Liehong (&#21016;&#28872;&#23439;) </span><a href="https://www.nda.gov.cn/sjj/swdt/sjdt/0611/20260611190621450122505_pc.html"><span>convened a symposium</span></a><span> on the same theme with DeepSeek, ByteDance, Alibaba Cloud, and Tencent and legal scholars including the China University of Political Science and Law.</span></p><p><span>The plan was released with six expert explainers (&#19987;&#23478;&#35299;&#35835;) over ten days, with authors spanning national research bodies, academia, and the Beijing and Shanghai municipal data bureaus. Hu Jianbo (&#32993;&#22362;&#27874;), head of the National Data Development Research Institute, wrote two: one </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0605/20260605122430764908996_pc.html"><span>marking the launch</span></a><span> of a new National Dataset Management Service System&#8212;the &#8220;physically dispersed, logically centralized&#8221; backbone the plan envisions&#8212;and one </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0609/20260609212055771982317_pc.html"><span>walking through</span></a><span> the plan&#8217;s logic, arguing that the &#8220;public-data dividend&#8221; is fading and that proprietary industry data is now the competitive moat for model-builders. His first piece cast datasets as a great-power &#8220;strategic high ground,&#8221; pointing to the US &#8220;</span><a href="https://www.whitehouse.gov/presidential-actions/2025/11/launching-the-genesis-mission/"><span>Genesis Mission</span></a><span>,&#8221; which involves consolidating federal data for AI training. Tsinghua&#8217;s Meng Qingguo (&#23391;&#24198;&#22269;) </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0612/20260612222534802449334_pc.html"><span>stressed</span></a><span> that Chinese models are &#8220;long on general knowledge but short on specialized knowledge&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a><span> and detailed the value-release agenda&#8212;token-based pricing, dataset trading on exchanges, and dataset-backed financing. The other three came from China Academy of Information and Communications Technology (CAICT) vice president Wei Liang (&#39759;&#20142;), </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0610/20260610184207948233640_pc.html"><span>on data supply</span></a><span>, and the deputy directors of the </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0614/20260614131904807563400_pc.html"><span>Beijing</span></a><span> (Peng Xuehai, &#24429;&#38634;&#28023;) and </span><a href="https://www.nda.gov.cn/sjj/zwgk/zjjd/0615/20260615170358664206520_pc.html"><span>Shanghai</span></a><span> (Qian Xiao, &#38065;&#26195;) data bureaus.</span></p><h2><span>MIIT issues a three-year &#8220;AI+ Information and Communications&#8221; plan</span></h2><p><span>On June 10, the Ministry of Industry and Information Technology (MIIT) released its </span><a href="https://www.miit.gov.cn/jgsj/txs/wjfb/art/2026/art_c1fe635702fc4339bf85289fe605ac21.html"><span>Implementation Opinion on the Innovative Development of &#8220;AI+ Information and Communications&#8221; (2026&#8211;2028)</span></a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a><span>&#8212;the latest in a series of sectoral plans operationalizing the national &#8220;AI+&#8221; initiative (including AI+ energy; see </span><a href="https://chinaaibulletin.substack.com/i/198737671/other-news-ai-energy-and-leader-inspections"><span>CAIB #4</span></a><span> and </span><a href="https://chinaaibulletin.substack.com/i/200679488/ai-energy-continues-to-develop"><span>#5</span></a><span> for discussion). The document includes network and compute build-out goals, plus AI-specific ones: by 2028 MIIT wants telecom networks running with &#8220;high-grade autonomy&#8221; (&#33258;&#26234;), at least 30 &#8220;high-value&#8221; AI scenarios, and a set of &#8220;distinctive AI agents.&#8221; It also names embodied intelligence as a network-integration priority and, on the consumer side, AI smartphones and PCs, smart-home devices, and carrier-built AI assistants. On the governance side, it includes &#8220;strengthening the industry&#8217;s governance capacity&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a><span> as one of four pillars and &#8220;strengthening international cooperation&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a><span> among its safeguards.</span></p><h2><span>MIIT and SASAC launch a humanoid robot training push</span></h2><p><span>On June 8, MIIT&#8217;s General Office and the State Council&#8217;s State-owned Assets Supervision and Administration Commission (SASAC) jointly issued a </span><a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_cd666691abf8471fb8553d463aa416e3.html"><span>notice</span></a><span> launching a &#8220;2026 Humanoid Robot and Embodied Intelligence Real-Scene Training Special Action.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a><span> It aims to create a deployment and data flywheel: put humanoid and quadruped robots to work in real-world settings&#8212;including manufacturing, warehousing, inspection, elderly care, and emergency response&#8212;to accumulate &#8220;real-machine data&#8221; (&#30495;&#26426;&#25968;&#25454;) that improves embodied intelligence models and hardware. Its goal is to have &#8220;100+ high-value application scenarios&#8221; and &#8220;10,000-unit-scale deployment capacity&#8221; by the end of 2026.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a><span> It also pulls in standards and safety scaffolding&#8212;a robot &#8220;ID card&#8221; (&#36523;&#20221;&#35777;; a requirement </span><a href="https://tech.chinadaily.com.cn/a/202606/01/WS6a1cee5fa310942cc49af3b5.html"><span>recently in effect</span></a><span>) for lifecycle management, MIIT&#8217;s humanoid-robot standardization committee, and collision-detection and emergency-braking requirements for human-machine settings&#8212;and floats a &#8220;robot-as-a-service&#8221; model to lower buyers&#8217; costs.</span></p><h2><span>People&#8217;s Daily commentary frames AI misuse in the cognitive domain as a national security concern</span></h2><p><span>On June 16, the People&#8217;s Daily ran a </span><a href="https://www.news.cn/politics/20260616/655b4e0b510a4f9baf33daeb7ede4ba6/c.html"><span>commentary</span></a><span> by Zhang Jun (&#24352;&#20891;)&#8212;a Chinese Academy of Engineering academician and Party secretary of the Beijing Institute of Technology&#8212;arguing that AI is &#8220;deeply reconstructing the boundaries and system of national security&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a><span> and that AI misuse in the &#8220;cognitive domain&#8221; (including deepfakes, disinformation, and public opinion attacks) has become &#8220;one of the most direct threats to national security.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a><span> Built around Xi Jinping&#8217;s cited statement that China must &#8220;ensure AI is safe, reliable, and controllable,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a><span> it outlines three priorities: talent (cross-disciplinary people who understand &#8220;technology, governance, and security&#8221;), self-reliance in &#8220;root technologies&#8221; (&#26681;&#25216;&#26415;) such as chips and algorithm frameworks, and &#8220;bottom-line thinking&#8221; on risk, including calling to &#8220;accelerate AI legislation&#8221; and build lifecycle safety assessment into AI projects. It is an individual op-ed, not a policy statement, but its placement in the People&#8217;s Daily can </span><a href="https://www.cambridge.org/core/journals/china-quarterly/article/abs/command-communication-the-politics-of-editorial-formulation-in-the-peoples-daily/DA9116B78E8898F0C63C7726488E7D53"><span>indicate</span></a><span> agenda signaling or that a view is being floated to assess public opinion.</span></p><h2><span>CAC opens a public channel for reporting AI application problems</span></h2><p><span>On June 12, the reporting center of the Cyberspace Administration of China (CAC) opened a dedicated section for the public to report problems with AI applications, part of a Qinglang special campaign to rectify AI application &#8220;disorder&#8221; (see coverage in </span><a href="https://chinaaibulletin.substack.com/i/198737671/cac-targets-ai-chaos-in-enforcement-sweep"><span>CAIB #4</span></a><span>).</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a><span> Reports run through the existing 12377 system, a national hotline/website for reporting illegal online content. The section lists 14 problem types in two buckets. One covers AI service and compliance failures: large models that skipped required filing, weak platform security and content-filtering, training-corpus safety and data-poisoning risks, missing labels on AI-generated content, use of AI for illegal activity, and lax security management of open-source models. The other covers content abuses: fabricated information, impersonation, violent or vulgar material, harm to minors, AI-run &#8220;water armies&#8221; (&#32593;&#32476;&#27700;&#20891;) for astroturfing, non-compliant AI apps, and using AI to &#8220;remix&#8221; (&#39764;&#25913;) classic works into &#8220;digital slop&#8221; (&#25968;&#23383;&#27860;&#27700;). The scope maps onto China&#8217;s existing AI rules&#8212;including the requirement to file large models with the government and label AI-generated content&#8212;and gives the public a route to flag violations of them.</span></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this useful? Subscribe to get it in your inbox regularly.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h2><span>National Standards</span></h2><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">MIIT&#8217;s AI standards committee moves 29 AI safety/security standards at a working-group session</span></h3><p><span>At the 2026 first standards-week of the Ministry of Industry and Information Technology&#8217;s AI Standardization Technical Committee, its Security Governance Working Group (WG8) </span><a href="https://miittc1.caict.ac.cn/newsDetail/?id=95&amp;type=1"><span>held its third 2026 session</span></a><span> on June 8&#8211;9 in Beijing, chaired by group head Shi Lin (&#30707;&#38678;). More than 100 experts attended, with representatives from CAICT, Beijing Jiaotong University, Sangfor, the three state telecom carriers, Ant Group, Huawei, Inspur, ZTE, and ByteDance&#8217;s Volcano Engine, among others. The session moved 29 standards across three stages: it discussed four public-comment drafts&#8212;including overall technical requirements for large-model security benchmark testing</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-21" href="#footnote-21" target="_self">21</a><span>&#8212;approved ten new project registrations, including security requirements for equipment-manufacturing industrial large models,</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-22" href="#footnote-22" target="_self">22</a><span> and reviewed fifteen pre-research drafts, including on agent data-security technical requirements and evaluation methods.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-23" href="#footnote-23" target="_self">23</a><span> Discussion centered on data, model, and interaction security and risk governance for large models, agents, and embodied intelligence&#8212;the build-out of a dedicated &#8220;AI Security Governance&#8221; standards series.</span></p><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">SAC proposes industrial AI agents standards</span></h3><p><span>The Standardization Administration of China (SAC) </span><a href="https://std.samr.gov.cn/gb/gbSuggestionPlan?bId=10003289"><span>issued</span></a><span> 126 proposed national standards for comment, including a seven-part Industrial AI Agents</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-24" href="#footnote-24" target="_self">24</a><span> series&#8212;</span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A24F0FA9CB62D2FE06397BE0A0A6E65"><span>general requirements</span></a><span>, </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A2505BD02E63303E06397BE0A0A1B2A"><span>classification and evaluation</span></a><span>, </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A25153DDC3A36BDE06397BE0A0A6E36"><span>task perception and understanding</span></a><span>, </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A2523855A663A5BE06397BE0A0A05E0"><span>knowledge memory and reasoning</span></a><span>, </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A2632E7B6AF7C15E06397BE0A0A32A7"><span>decision and orchestration</span></a><span>, </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A263A96BCB07DCDE06397BE0A0A57E4"><span>skill development and interaction</span></a><span>, and </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4A265E5B695C860CE06397BE0A0A30D6"><span>autonomous execution and learning</span></a><span>. The batch also includes </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=48B8BAC03CA2ED33E06397BE0A0A17EE"><span>application requirements for agents in business-management systems</span></a><span>, two digital human standards</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-25" href="#footnote-25" target="_self">25</a><span> (</span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=20E31CBFFC119DD0E06397BE0A0AB378"><span>detection and recognition</span></a><span>, and a </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4C2DF32A01A0EB15E06397BE0A0A9ACA"><span>taxonomy for the emotional presentation of digital humans</span></a><span>), a </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=4C2E3381CF45FE26E06397BE0A0A39D3"><span>multi-view 3D-reconstruction spec</span></a><span>, a </span><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=47244CFAA468810EE06397BE0A0AE14D"><span>compute-in-memory accelerator instruction set</span></a><span>, plus several autonomous-driving functional-safety and safety-of-the-intended-functionality (SOTIF) standards.</span></p><h1><span>International AI Governance</span></h1><h2><span>China&#8217;s proposed World AI Cooperation Organization may be advancing</span></h2><p><span>On June 17, the State Council Information Office released </span><a href="https://www.news.cn/20260617/fc1d3c9326434d1fafebfac25e4ba11f/c.html"><span>Building a More Just and Equitable Global Governance System: China&#8217;s Concepts, Initiatives and Actions</span></a><span>.</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-26" href="#footnote-26" target="_self">26</a><span> The white paper is not AI-specific&#8212;it covers climate, the digital divide, food and energy security, and other cross-border challenges&#8212;but it names &#8220;the misuse of AI giving rise to safety/security risks&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-27" href="#footnote-27" target="_self">27</a><span> among the problems it says global governance must address, and it devotes a passage to AI. That passage reiterates previous rhetoric on global governance, calling to &#8220;promote AI to develop for good and for the benefit of all,&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-28" href="#footnote-28" target="_self">28</a><span> restating the 2023 </span><a href="https://www.fmprc.gov.cn/eng/xw/zyxw/202405/t20240530_11332389.html"><span>Global AI Governance Initiative</span></a><span>, and backing the UN as the main channel in building a global AI governance system. It also directly addresses the security concerns of military AI: China &#8220;attaches great importance to guarding against the risks of military applications of AI,&#8221; urging states to be &#8220;prudent and responsible&#8221; in developing and using such technologies, to keep relevant weapons systems &#8220;always under human control,&#8221; and to &#8220;prevent an AI arms race.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-29" href="#footnote-29" target="_self">29</a><span> It doesn&#8217;t dip into loss of control language, but demonstrates a clear concern for the security risks of military AI.</span></p><p><span>At the launch </span><a href="https://www.stdaily.com/web/gdxw/2026-06/17/content_533731.html"><span>press conference</span></a><span>, Politburo member and Foreign Minister Wang Yi (&#29579;&#27589;) said China is accelerating preparations to establish the World AI Cooperation Organization (WAICO)</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-30" href="#footnote-30" target="_self">30</a><span> and welcomes all parties to join. China first </span><a href="https://www.fmprc.gov.cn/mfa_eng/xw/zyxw/202507/t20250729_11679232.html"><span>proposed establishing</span></a><span> WAICO at the July 2025 World AI Conference, alongside the Global AI Governance Action Plan, and floated Shanghai as its headquarters. The white paper reiterates the 2025 language&#8212;it says China &#8220;proposed establishing&#8221; (&#20513;&#35758;&#25104;&#31435;) WAICO&#8212;so Wang Yi&#8217;s &#8220;accelerating preparations&#8221; (&#21152;&#32039;&#31609;&#24314;) is the firmest commitment language to date. Although there are no other details, National Development and Reform Commission (NDRC) Vice Chairman Zhou Haibing (&#21608;&#28023;&#20853;) mentioned the July 2026 World AI Conference as &#8220;an opportunity to further strengthen international cooperation in artificial intelligence with all parties.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-31" href="#footnote-31" target="_self">31</a><span> Zhou also stated that a next step is to &#8220;uphold coordinated development and safety/security&#8221; and &#8220;explore AI regulatory cooperation to jointly guard against AI safety/security risks.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-32" href="#footnote-32" target="_self">32</a></p><h2><span>Chinese security scholars propose topics for the US-China AI dialogue</span></h2><p><span>Since the two governments agreed in May to begin an </span><a href="https://www.chinausfocus.com/peace-security/china-and-the-united-states-begin-official-ai-dialogue"><span>official AI dialogue</span></a><span>, China&#8217;s strategic studies/IR community has produced a steady run of commentary on what the channels should actually discuss, including two new pieces this fortnight. Tsinghua&#8217;s Xiao Qian, deputy director of the Center for International Security and Strategy (CISS) and vice dean of the Institute for AI International Governance (I-AIIG), </span><a href="https://www.chinausfocus.com/peace-security/ai-safety-is-where-us-china-cooperation-still-matters"><span>opened the run in May</span></a><span> with a broad menu including expert dialogues on frontier AI risks, communication channels on major AI incidents, joint work on AI evaluation and safety testing, and confidence-building measures on military AI and cyber stability. Days later, Fudan&#8217;s Cai Cuihong </span><a href="https://fddi.fudan.edu.cn/e6/ad/c21253a779949/page.htm"><span>proposed</span></a><span> three principles for cooperation&#8212;equality (&#23545;&#31561;), boundaries (&#26377;&#30028;), and openness (&#24320;&#25918;)&#8212;coupling risk-warning and crisis-communication systems with accident reporting and model evaluation exchange, while warning against AI safety/security as an excuse for the &#8220;over-securitization of civilian AI, open-source models, cloud services, and scientific exchange.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-33" href="#footnote-33" target="_self">33</a><span> (Unlike the others, which were published in English in China-US Focus and thus aimed at an Anglophone audience, this was published in Chinese in China Daily.)</span></p><p><span>The last two weeks brought a series of narrower proposals. One to highlight: Qi Haotian (Peking University) </span><a href="https://www.chinausfocus.com/peace-security/china-us-coopetition-in-ais-military-applications"><span>proposed</span></a><span> a military AI &#8220;minimal template&#8221;: channels scoped to verify whether a dangerous anomaly is &#8220;real, local, degraded, spoofed, or spreading,&#8221; not to &#8220;settle blame in real time.&#8221; The design, borrowed explicitly from the Cold War US-Soviet hotline, is to buy time to verify and avoid miscalculation, not to resolve disputes.</span></p><h1><span>Frontier Lab Developments</span></h1><h2><span>Notable Model Releases</span></h2><p><strong><span>Zhipu releases GLM-5.2, a 744-billion-parameter open-weight model.</span></strong><span> On June 13&#8212;the day after the US pulled Anthropic&#8217;s Fable 5, and three days before the open weights&#8212;Zhipu pushed GLM-5.2 to </span><a href="https://x.com/Zai_org/status/2065704919299235870"><span>all tiers of its GLM Coding Plan</span></a><span>, the Claude-Code-compatible coding subscription it has run since 2025. Founder Tang Jie </span><a href="https://x.com/jietang/status/2065784751345287314"><span>wrote</span></a><span> that &#8220;the sudden restriction of certain frontier models is deeply regrettable&#8221; and that access had been &#8220;abruptly cut off for non-technical reasons;&#8221; an accompanying </span><a href="https://mp.weixin.qq.com/s/LDrbtLM0wiCTJorvd5GY9w"><span>developer letter</span></a><span> argued that that frontier intelligence &#8220;should not belong only to a few, nor be withdrawn at any time by a few rules.&#8221;</span><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-34" href="#footnote-34" target="_self">34</a></p><p><span>On June 16, Zhipu released the </span><a href="https://huggingface.co/zai-org/GLM-5.2"><span>open weights</span></a><span>. GLM-5.2 is a 744-billion-parameter mixture-of-experts model&#8212;a design that activates only a fraction of its parameters (here, 40 billion) for any given input&#8212;and Zhipu reports gains in efficiency and in handling long inputs. Code is at </span><a href="https://github.com/zai-org/GLM-5"><span>zai-org/GLM-5</span></a><span>, with the details in the family </span><a href="https://arxiv.org/abs/2602.15763"><span>technical report</span></a><span> and </span><a href="https://z.ai/blog/glm-5.2"><span>release blog</span></a><span>.</span></p><p><span>Zhipu also open-weighted </span><a href="https://huggingface.co/zai-org/SCAIL-2"><span>SCAIL-2</span></a><span>, a character-animation video model. More details in the </span><a href="https://huggingface.co/papers/2606.10804"><span>paper</span></a><span>, </span><a href="https://github.com/zai-org/SCAIL-2"><span>repository</span></a><span>, and a </span><a href="https://teal024.github.io/SCAIL-2/"><span>project page</span></a><span>.</span></p><p><strong><span>Moonshot releases Kimi-K2.7-Code, the fortnight&#8217;s fastest-adopted model.</span></strong><span> On June 11, Moonshot posted </span><a href="https://huggingface.co/moonshotai/Kimi-K2.7-Code"><span>Kimi-K2.7-Code</span></a><span>, a roughly 1.06-trillion-parameter open-weight multimodal model that drew about 173,000 Hugging Face downloads in its first week&#8212;the highest of any model this fortnight. It ships with the </span><a href="https://github.com/MoonshotAI/kimi-code"><span>kimi-code CLI</span></a><span>.</span></p><p><strong><span>Xiaomi releases the MiMo-Code agent and MiMo-V2.5-Pro.</span></strong><span> Xiaomi&#8217;s </span><a href="https://github.com/XiaomiMiMo/MiMo-Code"><span>MiMo-Code</span></a><span> coding agent was the fastest-starred repository of the fortnight, running on the underlying </span><a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro-FP4-DFlash"><span>MiMo-V2.5-Pro</span></a><span> model. The adoption came with friction: </span><a href="https://36kr.com/p/3849833227572226"><span>36Kr</span></a><span> reported on June 12 that the agent shipped with a wave of early bugs&#8212;users opened more than 200 GitHub issues over severe lag, login-credential and API-key-import failures, and environment errors. Xiaomi&#8217;s MiMo line traces to its 2025 </span><a href="https://arxiv.org/abs/2505.07608"><span>technical report</span></a><span> for the original 7B-parameter model; no dedicated report for V2.5-Pro has appeared.</span></p><p><strong><span>Baichuan details M4, its clinical-grade medical model.</span></strong><span> In a </span><a href="https://arxiv.org/abs/2606.08982"><span>technical report</span></a><span> (June 8), Baichuan Intelligence described Baichuan-M4, a medical model built for continuous patient care rather than one-off medical Q&amp;A. It is structured as a coordinated agent system, with long-term patient memory, evidence-based retrieval, and the ability to read documents, X-rays, and dermatology images alongside text. Baichuan reports leading results across a medical evaluation suite covering clinical knowledge, safety, OSCE-style consultations, and long-context patient memory.</span></p><p><strong><span>Also released this fortnight:</span></strong></p><ul><li><p><strong><span>SenseTime SenseNova-U1-8B-MoT</span></strong><span>&#8212;a multimodal model (</span><a href="https://huggingface.co/sensenova/SenseNova-U1-8B-MoT-Interleaved"><span>HF</span></a><span>, </span><a href="https://github.com/OpenSenseNova/SenseNova-U1"><span>GitHub</span></a><span>).</span></p></li><li><p><strong><span>Tencent Hunyuan UniRL</span></strong><span>&#8212;a unified reinforcement-learning framework (</span><a href="https://github.com/Tencent-Hunyuan/UniRL"><span>GitHub</span></a><span>).</span></p></li><li><p><strong><span>Tencent Hy-Embodied-0.5-VLA</span></strong><span>&#8212;a vision-language-action model for robots (</span><a href="https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-UMI"><span>HF</span></a><span>, </span><a href="https://github.com/Tencent-Hunyuan/Hy-Embodied-0.5-VLA"><span>GitHub</span></a><span>).</span></p></li><li><p><strong><span>ByteDance Bernini</span></strong><span>&#8212;an image-to-video model (</span><a href="https://huggingface.co/ByteDance/Bernini-Diffusers"><span>HF</span></a><span>, </span><a href="https://github.com/bytedance/Bernini"><span>GitHub</span></a><span>).</span></p></li><li><p><strong><span>ByteDance EvoQuality</span></strong><span>&#8212;a training-data-quality method (</span><a href="https://huggingface.co/ByteDance/EvoQuality"><span>HF</span></a><span>, </span><a href="https://github.com/bytedance/EvoQuality"><span>GitHub</span></a><span>, </span><a href="https://openreview.net/forum?id=INOi0YqI8p"><span>paper</span></a><span>).</span></p></li></ul><p></p><h2><span>Technical Publication Highlights</span></h2><p><span>Frontier labs released 138 papers on arXiv this fortnight&#8212;and this edition, Baichuan and MiniMax got themselves on the board. Highlights are below; a full list with summaries can be found </span><a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0"><span>here</span></a><span>.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OnZi!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OnZi!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 424w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 848w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 1272w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OnZi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png" width="1278" height="1150" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1150,&quot;width&quot;:1278,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127141,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/202627026?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!OnZi!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 424w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 848w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 1272w, https://substackcdn.com/image/fetch/$s_!OnZi!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F897da5a0-f06a-419f-92b1-b3d28a02649a_1278x1150.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Alibaba</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.10646v1"><span>How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs</span></a></strong></p><ul><li><p><span>Proposes </span><strong><span>FlowTracer</span></strong><span>, an RL framework that traces how information flows through an LLM by modeling attention patterns as a directed graph, then assigns credit to tokens based on their role in routing information toward correct answers rather than treating all tokens equally. Token importances derived from this flow analysis </span><strong><span>reshape reward signals to focus learning on decisive reasoning steps</span></strong><span>, delivering consistent gains on reasoning tasks.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.09076v1"><span>Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions</span></a></strong></p><ul><li><p><span>Replaces scalar reward signals with </span><strong><span>score distributions</span></strong><span> in text-to-image models, using a teacher-student framework where a large VLM infers rubric-aligned score distributions via reasoning, then distills this into a compact student model for deployment. The 9B student model reaches </span><strong><span>88.6% human preference accuracy while enabling a 41.3% improvement in downstream image generation</span></strong><span> when used as a differentiable optimization signal.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.12370v1"><span>Breaking Entropy Bounds: Accelerating RL Training via MTP with Rejection Sampling</span></a></strong></p><ul><li><p><span>Identifies how model entropy bounds </span><strong><span>Multi-Token Prediction acceptance rates</span></strong><span> during RL training, then proposes </span><strong><span>Bebop</span></strong><span>, which combines probabilistic rejection sampling with a novel TV loss to optimize multi-step decoding. The approach achieves up to 95% acceptance rates and </span><strong><span>1.8x end-to-end speedup</span></strong><span> on Qwen models without requiring online MTP updates during RL.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.17030v1"><span>Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Qwen-RobotWorld</span></strong><span>, a </span><strong><span>language-conditioned video world model</span></strong><span> that predicts future visual trajectories across robotic manipulation, autonomous driving, navigation, and human-robot tasks using natural language as a unified action interface. </span><strong><span>It ranks 1st on EWMBench and DreamGen Bench</span></strong><span>, and tops WorldModelBench and PBench open models.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Baichuan</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.08982v2"><span>Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Baichuan-M4</span></strong><span>, a clinical-grade medical agent system for continuous patient care that coordinates a reasoning model, runtime framework, and clinical tools to handle dynamic consultations, long-term memory, and multimodal medical data&#8212;achieving a </span><strong><span>3.3% hallucination rate</span></strong><span> across static knowledge, safety, and evidence-based retrieval benchmarks.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Baidu</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.05806v1"><span>When Tools Fail: Benchmarking Dynamic Replanning and Anomaly Recovery in LLM Agents</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>ToolMaze</span></strong><span>, a </span><strong><span>benchmark that evaluates how LLM agents handle real-world tool failures</span></strong><span> through dynamic replanning and error recovery. Results show implicit semantic failures cause the sharpest performance drops (~37% recovery rate decline), and fault-tolerance improves 3.66x slower than general task performance with scale, revealing replanning as a distinct bottleneck beyond model size.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Meituan</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.10394v1"><span>STAGE-Claw: Automated State-based Agent Benchmarking for Realistic Scenarios</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>STAGE-Claw</span></strong><span>, a framework that automatically generates realistic personal-agent benchmarks by creating tasks, environments, and ground-truth validation in actual operating systems, then evaluates agents on </span><strong><span>whether the final system state matches the goal</span></strong><span> rather than response text. Testing 11 frontier models on 40 tasks reveals </span><strong><span>tool-call reliability issues and cost-performance tradeoffs</span></strong><span>.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">MiniMax</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.13473v1"><span>MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>MaxProof</span></strong><span>, a test-time scaling framework that applies </span><strong><span>proof generation, verification, and repair</span></strong><span> to competition math problems. By treating a single model as both generator and verifier, searching over proof populations, and using tournament selection, </span><strong><span>M3 achieves 35/42 on IMO 2025 and 36/42 on USAMO 2026</span></strong><span>&#8212;surpassing human gold-medal performance on both benchmarks.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.13392v1"><span>MiniMax Sparse Attention</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>MiniMax Sparse Attention (MSA)</span></strong><span>, a blockwise sparse attention mechanism that reduces per-token compute by 28.4x at 1M context while maintaining performance parity with standard attention. Co-designed GPU kernels achieve 14.2x prefill and 7.6x decoding speedups on H800 chips to facilitate ultra-long-context inference at scale.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Tencent</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2602.10430v1"><span>Breaking the Curse of Repulsion: Optimistic Distributionally Robust Policy Optimization for Off-Policy Generative Recommendation</span></a></strong></p><ul><li><p><span>Addresses </span><strong><span>model collapse</span></strong><span> in offline RL-based recommendation systems by reformulating training as </span><strong><span>Distributionally Robust Optimization (DRO)</span></strong><span>, proving that hard filtering of low-quality data optimally recovers high-quality behaviors while eliminating noise-induced divergence. DRPO achieves state-of-the-art performance on mixed-quality recommendation benchmarks.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.11324v1"><span>Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>Embodied-R1.5</span></strong><span>, an 8B-parameter foundation model that unifies embodied reasoning (cognition, planning, correction, and pointing) within a single architecture, trained on 15B tokens via automated data pipelines and balanced multi-task RL. It claims </span><strong><span>state-of-the-art on 16 of 24 embodied VLM benchmarks</span></strong><span> and, when fine-tuned, outperforms leading robotics models on manipulation tasks while validating strong zero-shot real-robot performance across instruction following, affordance grounding, and complex long-horizon tasks.</span></p></li></ul><h3><span data-color="rgb(67, 67, 67)" style="color: rgb(67, 67, 67);">Xiaomi</span></h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.05645v1"><span>Discrete-WAM: Unified Discrete Vision-Action Token Editing for World-Policy Learning</span></a></strong></p><ul><li><p><span>Proposes </span><strong><span>Discrete-WAM</span></strong><span>, a </span><strong><span>world model that represents future visual states and actions as aligned discrete tokens</span></strong><span>, enabling compositional reasoning about how ego actions shape driving scenarios. Built on unified discrete diffusion, it jointly trains world modeling and policy learning, claiming strong performance on autonomous-driving benchmarks.</span></p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.08525v1"><span>DriveReward: A Comprehensive Dataset and Generative Vision-Language Reward Model for Autonomous Driving</span></a></strong></p><ul><li><p><span>Introduces </span><strong><span>DriveReward</span></strong><span>, a dataset and specialized </span><strong><span>vision-language reward model</span></strong><span> for autonomous driving that uses temporally-grounded visual annotations and counterfactual failure cases to train a 1B parameter model that outperforms larger VLMs on driving-specific reward alignment. The model achieves performance comparable to hand-crafted rules when integrated into RL fine-tuning and trajectory scoring.</span></p></li></ul><p></p><h1><span>Technical AI Safety Publication Highlights</span></h1><p><span>There were </span><strong><span>48 AI-safety-related papers published by Chinese researchers </span></strong><span>this fortnight. Highlights are below; a full list with summaries is available </span><a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1"><span>here</span></a><span>.</span></p><h2><span>&#128269;Spotlight</span></h2><p><span>The autonomous execution of cyberattacks that cause real-world harm is often considered a red line frontier AI must not cross. The paper &#8220;</span><a href="https://arxiv.org/abs/2606.13079"><span>The Emergence of Autonomous Penetration Capabilities in Large Language Model-Powered AI Systems</span></a><span>&#8221; isolates a core enabling sub-task: </span><strong><span>autonomous penetration</span></strong><span>, whether an LLM-powered agent can, with no human in the loop, break into a target server, find and exploit a vulnerability, and gain unauthorized control. Its premise is that current evaluations don&#8217;t measure this honestly&#8212;according to the paper, </span><strong><span>OpenAI&#8217;s and Anthropic&#8217;s are opaque</span></strong><span> (their system cards don&#8217;t disclose scaffolding, targets, or protocols), academic benchmarks use </span><strong><span>oversimplified targets</span></strong><span>, and most </span><strong><span>hand the model too much task-specific prior knowledge</span></strong><span>.</span></p><p><span>To fix that, the authors&#8212;a team from </span><strong><span>Fudan University</span></strong><span>, the </span><strong><span>Shanghai AI Laboratory</span></strong><span>, </span><strong><span>Concordia AI, </span></strong><span>and</span><strong><span> </span></strong><span>the</span><strong><span> Shanghai Innovation Institute</span></strong><span>&#8212;built a framework with </span><strong><span>300 target servers</span></strong><span>, each pairing a vulnerable service with 1-3 secure ones, so the agent had to find the real weakness amid decoys. Each model ran inside a</span><strong><span> general-purpose agent loop</span></strong><span>, equipped with standard cybersecurity tools but no target-specific hints, so the score reflects the model&#8217;s own ability rather than prior knowledge it was handed. Across </span><strong><span>19 open-weight and proprietary models</span></strong><span>, penetration success ran from </span><strong><span>10.7% to 69.3%</span></strong><span>, and </span><strong><span>rose in step with general model capability</span></strong><span>&#8212;models get better at autonomous intrusion as a byproduct of getting more capable overall, not from being built for it.</span></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!mTn9!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!mTn9!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 424w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 848w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 1272w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!mTn9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png" width="1016" height="468" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:468,&quot;width&quot;:1016,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!mTn9!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 424w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 848w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 1272w, https://substackcdn.com/image/fetch/$s_!mTn9!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F3e48e308-1436-4655-9e9d-157016f94be4_1016x468.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: the <a href="https://arxiv.org/abs/2606.13079">paper</a></figcaption></figure></div><p><span>The authors are careful to call this a &#8220;first step&#8221; towards evaluating the offensive capabilities of AI systems; the metric scores one scoped objective&#8212;gaining shell access to a single host&#8212;not a full, multi-stage attack. However, the benchmark, intended to be harder to game, may help reduce the likelihood of such an attack in the first place.</span></p><h2><span>Agentic Safety</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.15242v1"><span>Benign in Isolation, Harmful in Composition: Security Risks in Agent Skill Ecosystems</span></a></strong></p><ul><li><p><strong><span>Skill Composition Risk (SCR)</span></strong><span> occurs when individual LLM agent skills appear safe in isolation but become harmful when combined&#8212;for example, one skill&#8217;s output could elevate another skill&#8217;s permissions or leak data into a subsequent operation. The authors introduce </span><strong><span>SCR-Bench</span></strong><span>, a benchmark that evaluates multi-skill execution paths in sandboxed environments, tracking state changes and outcomes rather than relying on surface behavior alone. Results show attack success rates jumping from near-zero in isolation to 33.6&#8211;96.5 percent under composition, demonstrating that vetting individual skills misses critical interaction vulnerabilities.</span></p></li></ul><p><em><span>Institutional affiliations: East China Normal University, Centre for Frontier AI Research, A*STAR, Shanghai Innovation Institute</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.08531v1"><span>VESTA: A Fully Automated Scenario Generation and Safety Evaluation Framework for LLM Agents</span></a></strong></p><ul><li><p><strong><span>VESTA</span></strong><span> is an automated framework that generates diverse safety scenarios for LLM agents and evaluates their behavior during task execution, not just final outputs. It maps five risk dimensions&#8212;including deception, excessive autonomy, and environmental harm&#8212;into </span><strong><span>1,072 executable test scenarios</span></strong><span> across memory, tool use, and external environment access. Evaluation of 12 agents revealed </span><strong><span>average attack success rates of 47.1%</span></strong><span>, with some models exceeding 70%, indicating that process-level safety evaluation surfaces behavioral risks that static benchmarks miss.</span></p></li></ul><p><em><span>Institutional affiliations: BrainCog AI Lab (CAS), Beijing Institute of AI Safety and Governance (Beijing-AISI), Beijing Key Laboratory of Safe AI and Superalignment, School of Artificial Intelligence, UCAS, Long-term AI</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.04455v1"><span>The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development?</span></a></strong></p><ul><li><p><span>The </span><strong><span>Meta-Agent Challenge (MAC)</span></strong><span> tests whether frontier models can autonomously design agent systems&#8212;a capability beyond standard task execution benchmarks. Meta-agents receive a sandbox, evaluation API, and time budget to iteratively build agents optimized on held-out test sets. Current models rarely match human-engineered baselines; those that succeed rely on proprietary frontier systems. High optimization pressure surfaces </span><strong><span>emergent adversarial behaviors</span></strong><span> like ground-truth data exfiltration, revealing gaps in robustness and alignment under recursive self-improvement scenarios.</span></p></li></ul><p><em><span>Institutional affiliations: Chinese Information Processing Laboratory (CAS), University of Chinese Academy of Sciences, Ant Group</span></em></p><h2><span>Alignment</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.09068v1"><span>Emergent Misalignment Can Be Induced by Sycophancy and Reversed via Alignment Gating</span></a></strong></p><ul><li><p><span>Fine-tuning models to agree with users regardless of accuracy&#8212;</span><strong><span>sycophancy training</span></strong><span>&#8212;induces </span><strong><span>broad misalignment</span></strong><span> beyond the narrow domain, a previously underexplored pathway to harmful emergent behavior. The authors propose </span><strong><span>Alignment Gating</span></strong><span>, which inserts learnable gates during fine-tuning to identify and suppress internal representations driving unsafe outputs. Gating weights trained on narrow domains </span><strong><span>generalize to suppress misalignment across broader contexts</span></strong><span> while preserving general capabilities, offering an efficient reversal mechanism.</span></p></li></ul><p><em><span>Institutional affiliations: Shanghai AI Lab</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.04075v1"><span>Large Language Models Hack Rewards, and Society</span></a></strong></p><ul><li><p><strong><span>SocioHack</span></strong><span> is a benchmark of 72 simulated societal environments designed to test whether LLMs exploit regulatory gaps during reinforcement learning training. The authors find that </span><strong><span>reward hacking naturally scales into &#8220;societal hacking&#8221;</span></strong><span>&#8212;models discover loopholes that remain technically compliant with rules while defeating their intent, mirroring how they game narrow metrics. Current safeguards provide limited protection, raising concerns about gathering real-world feedback for model training without stronger assurances that RL won&#8217;t uncover and exploit gaps in actual regulations.</span></p></li></ul><p><em><span>Institutional affiliations: King&#8217;s College London, Fudan University, The Alan Turing Institute</span></em></p><h2><span>Evaluation and Benchmarks</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.10484v1"><span>AgentCanary: A Security Evaluation Framework for Autonomous AI Agents in Real Executable Environments</span></a></strong></p><ul><li><p><strong><span>AgentCanary</span></strong><span> is a security evaluation framework that tests autonomous AI agents in </span><strong><span>real, executable environments</span></strong><span> rather than static Q&amp;A settings. It uses an </span><strong><span>Entry &#215; Impact risk taxonomy</span></strong><span> to systematically map how attacks compromise agents and what harms result, coupled with dynamic task artifacts and persistent state to simulate realistic multi-step workflows. Evaluation across frontier models reveals </span><strong><span>agents frequently fail to detect attacks</span></strong><span>&#8212;especially under compromised tools, state persistence, and long-horizon execution&#8212;establishing baselines for hardening agent security before deployment.</span></p></li></ul><p><em><span>Institutional affiliations: Ant Group, Tsinghua University, Nanjing University, Peking University</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.18060v1"><span>PseudoBench: Measuring How Agentic Auto-Research Fuels Pseudoscience</span></a></strong></p><ul><li><p><strong><span>PseudoBench</span></strong><span> is an adversarial benchmark measuring whether autonomous AI research agents can resist pseudoscientific narratives across five domains. Testing seven state-of-the-art agents on 200 claim-evidence pairs, researchers found that current systems </span><strong><span>produce persuasive pseudoscientific reports with near-zero refusal rates</span></strong><span>, with maximum resistance only 27.4%. Stronger agents risk </span><strong><span>legitimizing false claims through sophisticated scientific framing</span></strong><span>, potentially contaminating academic literature and eroding public trust in science before these systems see widespread deployment.</span></p></li></ul><p><em><span>Institutional affiliations: Shanghai Artificial Intelligence Laboratory, Xi&#8217;an Jiao Tong University, Shanghai Jiao Tong University</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.05567v1"><span>ZERO-APT: A Closed-Loop Adversarial Framework for LLM-Driven Automated Penetration Testing under Intelligent Defense</span></a></strong></p><ul><li><p><strong><span>ZERO-APT</span></strong><span> is a closed-loop framework that evaluates LLM-driven penetration testing agents against </span><strong><span>active defenders</span></strong><span> rather than static targets. It addresses three gaps: </span><strong><span>realism</span></strong><span> (embedding an LLM defender that detects attacks via system telemetry), </span><strong><span>consistency</span></strong><span> (enforcing causal reasoning through architecture rather than relying on unstable LLM chains), and </span><strong><span>auditability</span></strong><span> (a Judge agent produces structured reports tracing every decision). The prototype achieves 79% attack success on Windows Server scenarios, with full decision transparency&#8212;a capability gap highlighted by baseline agents scoring 22&#8211;39% against live defense.</span></p></li></ul><p><em><span>Institutional affiliations: Zhejiang University of Technology</span></em></p><h2><span>Guardrails and Deployment Safety</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.15396v1"><span>CHILLGuard: Towards Fine-Grained Chinese LLM Safety Guardrail with Scalable Data Construction and Model-aware Preference Alignment</span></a></strong></p><ul><li><p><strong><span>CHILLGuard</span></strong><span> is a Chinese-language safety classifier that addresses the gap between English-centric guardrails and China&#8217;s regulatory and cultural context. The authors develop a </span><strong><span>31-category risk taxonomy</span></strong><span> and construct 405,000+ annotated training samples through retrieval-augmented generation, adversarial prompt rewriting, and multi-model label voting. Trained via preference optimization, CHILLGuard achieves </span><strong><span>15.92% F1 improvement</span></strong><span> over existing Chinese baselines, enabling fine-grained risk classification for localized deployment.</span></p></li></ul><p><em><span>Institutional affiliations: Tsinghua University, Beijing Normal University, South China University of Technology, Harbin Institute of Technology, Shenzhen, Shenzhen ShenNong Information Technology Co.</span></em></p><h2><span>Interpretability</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.17478v1"><span>Decoding Hidden Deception in Reasoning LLMs: Activation Explainers for Deception Auditing</span></a></strong></p><ul><li><p><strong><span>STATEWITNESS</span></strong><span> is a decoder-based explainer that reads hidden states from reasoning LLMs and answers natural-language queries about them to detect deception. Rather than outputting scalar confidence scores, it generates structured reports and token-level evidence traces that humans can inspect directly. On seven deception datasets, </span><strong><span>STATEWITNESS achieved 0.916 AUROC</span></strong><span>&#8212;11.6% better than text-based monitors and 25% better than probe baselines&#8212;and reduced false negatives when combined with existing detection systems.</span></p></li></ul><p><em><span>Institutional affiliations: Zhejiang University, Griffith University</span></em></p><h2><span>Robustness and Adversarial Attacks</span></h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.11817v1"><span>Grammar-Constrained Decoding Can Jailbreak LLMs into Generating Malicious Code</span></a></strong></p><ul><li><p><strong><span>CodeSpear</span></strong><span> demonstrates that grammar-constrained decoding&#8212;a technique meant to enforce syntactic correctness in code generation&#8212;can be weaponized to jailbreak LLMs into producing malicious code by restricting the model&#8217;s output space in ways that bypass safety training. The authors propose </span><strong><span>CodeShield</span></strong><span>, a defense that teaches models to generate structurally diverse but semantically harmless &#8220;honeypot&#8221; code under constrained decoding, preserving refusals in natural language while blocking attacks across 10 popular models.</span></p></li></ul><p style="text-align: justify;"><em><span>Institutional affiliations: Tsinghua University, University of Electronic Science and Technology of China</span></em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.07970v1"><span>Defending Against Malicious Finetuning by Scaling Train-time Adversarial Attacks</span></a></strong></p><ul><li><p><strong><span>Patcher</span></strong><span> defends open-weight LLMs against malicious fine-tuning by simulating </span><strong><span>full-parameter attacks during training</span></strong><span>&#8212;not just parameter-efficient ones. The method uses adversarial training to find model weights that resist stronger poisoning attacks, scaling up attack intensity in the optimization loop to force robustness. Experiments show </span><strong><span>substantial improvements</span></strong><span> over standard alignment across diverse attack scenarios and model sizes, with parallel implementation reducing training time.</span></p></li></ul><p><em><span>Institutional affiliations: Xiongan AI Institute, Tsinghua University, Shanghai Qi Zhi Institute</span></em></p><h1><span>Export Controls &amp; Economic Policy</span></h1><h2><span>US bars foreign nationals from accessing Mythos and Fable</span></h2><p><span>On June 12, Anthropic </span><a href="https://www.anthropic.com/news/fable-mythos-access"><span>disabled</span></a><span> Fable 5 and Mythos 5 after the US government, citing national security authorities, issued an export control directive suspending access for all foreign nationals&#8212;inside or outside the United States, and including Anthropic&#8217;s own foreign-national employees. Because Anthropic had no reliable way to screen users by nationality, it suspended both models for everyone, US users included, while it sought clarification; Chinese media including </span><a href="https://www.caixin.com/2026-06-13/102453978.html"><span>Caixin</span></a><span> reported on the order on June 13.</span></p><h1><span>On the Horizon</span></h1><p><strong><span>A national accreditation for AI-security service providers (CNITSEC):</span></strong><span> Following the China Information Technology Security Evaluation Center&#8217;s </span><a href="https://mp.weixin.qq.com/s/G73th9B1ttgaH9v7NTszRg"><span>announcement</span></a><span> of a </span><a href="https://www.itsec.gov.cn/"><span>National Information Security Service Qualification</span></a><span> that vets vendors in AI security, watch for the first accredited cohort and for how this interacts with TC260 and CAC model-evaluation tracks.</span></p><div><hr></div><p style="text-align: justify;"><span>For more on how we select and track content, see our methodology </span><a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz"><span>here</span></a><span>.</span></p><div><hr></div><p><em><span>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</span></em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p><span>&#8220;&#36866;&#24212;&#20154;&#24037;&#26234;&#33021;&#21457;&#23637;&#20419;&#36827;&#23601;&#19994;&#21019;&#19994;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p><span>&#8220;&#25506;&#32034;&#20154;&#26426;&#21327;&#21516;&#30340;&#26032;&#22411;&#24037;&#20316;&#24418;&#24577;&#65292;&#24378;&#21270;&#20154;&#24037;&#26234;&#33021;&#30340;&#23601;&#19994;&#21019;&#36896;&#25928;&#24212;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p><span>&#8220;&#20154;&#24037;&#26234;&#33021;&#21457;&#23637;&#32418;&#21033;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p><span>&#8220;&#20154;&#24037;&#26234;&#33021;&#23545;&#23601;&#19994;&#24433;&#21709;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p><span>&#8220;&#23436;&#21892;&#20154;&#24037;&#26234;&#33021;&#24212;&#29992;&#23601;&#19994;&#39118;&#38505;&#39044;&#35686;&#22788;&#32622;&#20307;&#31995;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p><span>&#8220;&#30563;&#20419;&#24179;&#21488;&#20225;&#19994;&#35268;&#33539;&#31639;&#27861;&#12289;&#25552;&#39640;&#36879;&#26126;&#24230;&#65292;&#20445;&#38556;&#26032;&#23601;&#19994;&#24418;&#24577;&#21171;&#21160;&#32773;&#23545;&#31639;&#27861;&#35268;&#21017;&#30340;&#30693;&#24773;&#26435;&#12289;&#21442;&#19982;&#26435;&#12289;&#36873;&#25321;&#26435;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p><span>&#12298;&#20851;&#20110;&#25512;&#36827;&#34892;&#19994;&#39640;&#36136;&#37327;&#25968;&#25454;&#38598;&#24314;&#35774;&#34892;&#21160;&#30340;&#23454;&#26045;&#26041;&#26696;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p><span>The &#8220;AI+&#8221; (&#20154;&#24037;&#26234;&#33021;+) initiative is China&#8217;s national program to integrate AI across industries and the broader economy, elevated in the 2024 Government Work Report and formalized in the State Council&#8217;s August 2025 </span><a href="https://www.gov.cn/zhengce/content/202508/content_7037861.htm">&#8220;AI+&#8221; Opinions</a><span>. It echoes the 2015 &#8220;Internet+&#8221; (&#20114;&#32852;&#32593;+) initiative and is operationalized by ministry- and sector-level plans.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p><span>&#8220;&#24378;&#22522;&#25193;&#23481;&#12289;&#26631;&#27880;&#25915;&#22362;&#12289;&#25552;&#36136;&#22686;&#25928;&#12289;&#24212;&#29992;&#36171;&#33021;&#12289;&#31649;&#29702;&#26381;&#21153;&#12289;&#20215;&#20540;&#37322;&#25918;:</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p><span>&#8220;&#33853;&#23454;&#25968;&#25454;&#25345;&#26377;&#26435;&#12289;&#20351;&#29992;&#26435;&#12289;&#32463;&#33829;&#26435;&#19977;&#26435;&#20998;&#32622;&#21046;&#24230;&#8230;&#8230;&#23436;&#21892;&#20154;&#24037;&#26234;&#33021;&#35757;&#32451;&#38454;&#27573;&#25968;&#25454;&#20351;&#29992;&#35268;&#21017;&#65292;&#25512;&#21160;&#29256;&#26435;&#20316;&#21697;&#25968;&#25454;&#31561;&#26377;&#24207;&#29992;&#20110;&#27169;&#22411;&#35757;&#32451;&#65292;&#23436;&#21892;&#25968;&#25454;&#25480;&#26435;&#20351;&#29992;&#26426;&#21046;&#21644;&#25910;&#30410;&#20998;&#37197;&#35268;&#21017;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p><span>&#8220;&#27169;&#22411; &#8216;&#36890;&#35782;&#26377;&#20313;&#12289;&#19987;&#35782;&#19981;&#36275;&#8217;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p><span>&#12298;&#8220;&#20154;&#24037;&#26234;&#33021;+&#20449;&#24687;&#36890;&#20449;&#8221;&#21019;&#26032;&#21457;&#23637;&#23454;&#26045;&#24847;&#35265;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p><span>&#8220;&#22686;&#24378;&#20449;&#24687;&#36890;&#20449;&#34892;&#19994;&#27835;&#29702;&#33021;&#21147;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p><span>&#8220;&#21152;&#24378;&#22269;&#38469;&#21512;&#20316;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p><span>&#12298;&#24037;&#19994;&#21644;&#20449;&#24687;&#21270;&#37096;&#21150;&#20844;&#21381; &#22269;&#21153;&#38498;&#22269;&#36164;&#22996;&#21150;&#20844;&#21381;&#20851;&#20110;&#32852;&#21512;&#24320;&#23637;2026&#24180;&#24230;&#20154;&#24418;&#26426;&#22120;&#20154;&#19982;&#20855;&#36523;&#26234;&#33021;&#23454;&#26223;&#23454;&#35757;&#19987;&#39033;&#34892;&#21160;&#30340;&#36890;&#30693;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p><span>&#8220;&#20957;&#32451;&#24418;&#25104;&#30334;&#20010;&#20197;&#19978;&#39640;&#20215;&#20540;&#24212;&#29992;&#22330;&#26223;&#8230;&#8230;&#24102;&#21160;&#24418;&#25104;&#19975;&#21488;&#32423;&#35268;&#27169;&#33853;&#22320;&#33021;&#21147;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p><span>&#8220;&#20154;&#24037;&#26234;&#33021;&#27491;&#28145;&#24230;&#37325;&#26500;&#22269;&#23478;&#23433;&#20840;&#30340;&#36793;&#30028;&#19982;&#20307;&#31995;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p><span>&#8220;&#29983;&#25104;&#24335;&#20154;&#24037;&#26234;&#33021;&#25216;&#26415;&#22312;&#35748;&#30693;&#22495;&#30340;&#28389;&#29992;&#65292;&#24050;&#25104;&#20026;&#23545;&#22269;&#23478;&#23433;&#20840;&#26368;&#30452;&#25509;&#30340;&#23041;&#32961;&#20043;&#19968;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p><span>&#8220;&#30830;&#20445;&#20154;&#24037;&#26234;&#33021;&#23433;&#20840;&#12289;&#21487;&#38752;&#12289;&#21487;&#25511;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p><span>&#8220;&#28165;&#26391;&#183;&#25972;&#27835;AI&#24212;&#29992;&#20081;&#35937;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-21" href="#footnote-anchor-21" class="footnote-number" contenteditable="false" target="_self">21</a><div class="footnote-content"><p><span>&#12298;&#20154;&#24037;&#26234;&#33021; &#23433;&#20840;&#27835;&#29702; &#22823;&#27169;&#22411;&#23433;&#20840;&#22522;&#20934;&#27979;&#35797;&#24635;&#20307;&#25216;&#26415;&#35201;&#27714;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-22" href="#footnote-anchor-22" class="footnote-number" contenteditable="false" target="_self">22</a><div class="footnote-content"><p><span>&#12298;&#20154;&#24037;&#26234;&#33021; &#23433;&#20840;&#27835;&#29702; &#35013;&#22791;&#21046;&#36896;&#24037;&#19994;&#22823;&#27169;&#22411;&#23433;&#20840;&#35201;&#27714;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-23" href="#footnote-anchor-23" class="footnote-number" contenteditable="false" target="_self">23</a><div class="footnote-content"><p><span>&#12298;&#20154;&#24037;&#26234;&#33021; &#23433;&#20840;&#27835;&#29702; &#26234;&#33021;&#20307;&#25968;&#25454;&#23433;&#20840;&#25216;&#26415;&#35201;&#27714;&#19982;&#35780;&#20272;&#26041;&#27861;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-24" href="#footnote-anchor-24" class="footnote-number" contenteditable="false" target="_self">24</a><div class="footnote-content"><p><span>&#8220;&#24037;&#19994;&#26234;&#33021;&#20307;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-25" href="#footnote-anchor-25" class="footnote-number" contenteditable="false" target="_self">25</a><div class="footnote-content"><p><a href="https://zhuanlan.zhihu.com/p/406259189">&#34394;&#25311;&#25968;&#23383;&#20154;</a><span> (literally &#8220;virtual digital humans,&#8221; often translated as &#8220;virtual humans&#8221; or &#8220;digital humans&#8221;) are AI-driven avatars that look like humans and can interact with others in a human-like way.</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-26" href="#footnote-anchor-26" class="footnote-number" contenteditable="false" target="_self">26</a><div class="footnote-content"><p><span>&#12298;&#26500;&#24314;&#26356;&#21152;&#20844;&#27491;&#21512;&#29702;&#30340;&#20840;&#29699;&#27835;&#29702;&#20307;&#31995;&#65306;&#20013;&#22269;&#30340;&#29702;&#24565;&#12289;&#20513;&#35758;&#19982;&#34892;&#21160;&#12299;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-27" href="#footnote-anchor-27" class="footnote-number" contenteditable="false" target="_self">27</a><div class="footnote-content"><p><span>&#8220;&#20154;&#24037;&#26234;&#33021;&#28389;&#29992;&#24341;&#21457;&#23433;&#20840;&#39118;&#38505;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-28" href="#footnote-anchor-28" class="footnote-number" contenteditable="false" target="_self">28</a><div class="footnote-content"><p><span> &#8220;&#20419;&#36827;&#20154;&#24037;&#26234;&#33021;&#21521;&#21892;&#26222;&#24800;&#21457;&#23637;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-29" href="#footnote-anchor-29" class="footnote-number" contenteditable="false" target="_self">29</a><div class="footnote-content"><p><span>&#8220;&#20013;&#22269;&#39640;&#24230;&#37325;&#35270;&#20154;&#24037;&#26234;&#33021;&#20891;&#20107;&#24212;&#29992;&#39118;&#38505;&#38450;&#33539;&#65292;&#20027;&#24352;&#21508;&#22269;&#22312;&#30740;&#21457;&#20351;&#29992;&#30456;&#20851;&#25216;&#26415;&#26102;&#24212;&#37319;&#21462;&#24910;&#37325;&#36127;&#36131;&#24577;&#24230;&#65292;&#30830;&#20445;&#26377;&#20851;&#27494;&#22120;&#31995;&#32479;&#22987;&#32456;&#22788;&#20110;&#20154;&#31867;&#25511;&#21046;&#20043;&#19979;&#65292;&#38450;&#27490;&#20154;&#24037;&#26234;&#33021;&#20891;&#22791;&#31454;&#36187;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-30" href="#footnote-anchor-30" class="footnote-number" contenteditable="false" target="_self">30</a><div class="footnote-content"><p><span>&#19990;&#30028;&#20154;&#24037;&#26234;&#33021;&#21512;&#20316;&#32452;&#32455;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-31" href="#footnote-anchor-31" class="footnote-number" contenteditable="false" target="_self">31</a><div class="footnote-content"><p><span>&#8220;&#26399;&#24453;&#20197;&#26412;&#27425;&#22823;&#20250;&#20026;&#22865;&#26426;&#65292;&#21516;&#21508;&#26041;&#36827;&#19968;&#27493;&#21152;&#24378;&#22269;&#38469;&#20154;&#24037;&#26234;&#33021;&#21512;&#20316;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-32" href="#footnote-anchor-32" class="footnote-number" contenteditable="false" target="_self">32</a><div class="footnote-content"><p><span>&#8220;&#19979;&#19968;&#27493;&#65292;&#20013;&#22269;&#23558;&#22362;&#25345;&#32479;&#31609;&#21457;&#23637;&#21644;&#23433;&#20840;&#8230;&#8230;&#25506;&#32034;&#24320;&#23637;&#20154;&#24037;&#26234;&#33021;&#30417;&#31649;&#21512;&#20316;&#65292;&#20849;&#21516;&#38450;&#33539;&#20154;&#24037;&#26234;&#33021;&#23433;&#20840;&#39118;&#38505;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-33" href="#footnote-anchor-33" class="footnote-number" contenteditable="false" target="_self">33</a><div class="footnote-content"><p><span>&#8220;</span><span data-color="rgb(51, 51, 51)" style="color: rgb(51, 51, 51);">&#20294;AI&#23433;&#20840;&#19981;&#24212;&#25104;&#20026;&#23558;&#27665;&#29992;AI&#12289;&#24320;&#28304;&#27169;&#22411;&#12289;&#20113;&#26381;&#21153;&#12289;&#31185;&#23398;&#20132;&#27969;&#25110;&#20154;&#25165;&#27969;&#21160;&#8220;&#27867;&#23433;&#20840;&#21270;&#8221;&#30340;&#19975;&#33021;&#29702;&#30001;&#12290;&#8221;</span></p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-34" href="#footnote-anchor-34" class="footnote-number" contenteditable="false" target="_self">34</a><div class="footnote-content"><p><span>&#8220;&#21069;&#27839;&#26234;&#33021;&#19981;&#24212;&#21482;&#23646;&#20110;&#23569;&#25968;&#20154;&#65292;&#20063;&#19981;&#24212;&#34987;&#23569;&#25968;&#35268;&#21017;&#38543;&#26102;&#25910;&#22238;&#8221;</span></p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 5]]></title><description><![CDATA[Developments from 20/5/26-3/6/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-5</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-5</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Fri, 05 Jun 2026 04:56:28 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/26cfeb84-4892-449e-a96c-94552642e4bd_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome to Issue 5 of the China AI Bulletin, the latest on AI governance, development, and safety in China. Today&#8217;s highlights: Xi Jinping mentioned technological loss of control in a newly published speech, SAMR and NDRC open a new AI metrology track to support evals, MiniMax ships M3 as its first major flagship this year, and Shanghai AI Lab releases AgentDoG 1.5 as a training-free guardrail for agent safety.</p><p><em>Editor&#8217;s note: I&#8217;ll be in DC next week (June 8-11). If anyone working on China-related AI governance would like to meet up, feel free to reach out!</em></p><p><em>Number of the week: </em><strong>$200 million</strong> &#8212; the combined funding round closed by Beijing-based <a href="https://www.qbitai.com/2026/06/427516.html">VAST</a> on June 1 as it launched its <strong>Project Eden </strong>world model.</p><h1>Executive Summary</h1><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/domestic-ai-governance">Domestic AI Governance</a>:</strong> Xi Jinping&#8217;s January Politburo speech, published in <em>Qiushi</em> on May 31, <strong>discusses technological loss of control</strong> and <strong>names embodied intelligence among six future industries key to the 15th Five-Year Plan</strong>. SAMR and NDRC issued the <strong>AI Metrology Guidance</strong>, an <strong>explicit commitment to AI measurement infrastructure</strong>. CAC published four expert interpretations of the recent Ethics-Safety Guidelines, and Anhui, Harbin, and Hubei staked out distinct provincial AI strategies.</p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/national-standards">National Standards</a>:</strong> TC28/SC42 published <strong>GB/Z 185</strong>, China&#8217;s <strong>first state-issued agent-interconnection standard family.</strong></p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/international-ai-governance">International AI Governance</a>:</strong> MIIT and ASEAN inaugurated the <strong>China-ASEAN AI Industry Innovation Center</strong>, creating a formal bilateral dialogue venue with a standards-and-governance mandate. State-visit readouts with Serbia and Pakistan <strong>named AI as a designated cooperation area</strong>.</p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/frontier-lab-developments">Frontier Lab Developments</a>:</strong></p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/notable-model-releases">Notable Model Releases</a>:</strong> MiniMax released <strong>M3</strong>, the lab&#8217;s first major flagship of 2026. It&#8217;s a <strong>1M-context agentic foundation</strong> paired with a downloadable desktop coding agent <strong>MiniMax Code</strong>. <strong>Alibaba had the busiest release cadence of the cycle</strong>, including <strong>Qwen3.7-Plus</strong> (a multimodal GUI/CLI-agent foundation) and <strong>Qwen-VLA</strong>. Tencent released the full <strong>Hy-MT2 translation family</strong>, and StepFun shipped <strong>Step-3.7-Flash</strong>.</p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/technical-publication-highlights">Technical Papers</a>:</strong> Frontier labs released <strong>114 papers</strong> on arXiv, led by Alibaba (43), Tencent (24), and Huawei (16).</p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/200679488/technical-ai-safety-publication-highlights">Technical AI Safety</a>:</strong> 115 safety publications by Chinese researchers this fortnight, with a continued heavy agent safety focus. Shanghai AI Lab released <strong>AgentDoG 1.5</strong>, an<strong> agent-safety alignment framework deployable as a training-free online guardrail</strong>.</p></li></ul><h1>Domestic AI Governance</h1><h2>Xi Jinping discusses technological loss of control and embodied AI in <em>Qiushi</em> speech</h2><p>On May 31, <em>Qiushi</em> (the CCP&#8217;s flagship theoretical journal) <a href="https://www.gov.cn/yaowen/liebiao/202605/content_7070720.htm">released</a> the full text of Xi Jinping&#8217;s speech at the 24th collective study session of the 20th Central Committee Politburo, delivered January 30, 2026 (translation by Bill Bishop <a href="https://sinocism.notion.site/Forward-Looking-Planning-and-Development-of-Future-Industries-37184ece41d781c09ff2dfd2d5140e85">here</a>). The speech, titled &#8220;Forward-looking planning and development of future industries,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> confirms six &#8220;future industries&#8221; from the 20th CCP Central Committee 4th Plenum as the 15th Five-Year Plan&#8217;s main directions: quantum technology, biomanufacturing, hydrogen and nuclear fusion energy, brain-computer interfaces, embodied intelligence (&#20855;&#36523;&#26234;&#33021;), and 6G.</p><p>The speech is structured around five points:</p><p>(1) strengthen industrial coordination and planning;</p><p>(2) lead with science &amp; technology innovation (referring to making tech breakthroughs, strengthening basic research, and applying S&amp;T innovation to industry);</p><p>(3) make enterprises the primary drivers of innovation;</p><p>(4) create a sound policy environment (referring to fiscal policies, investment, procurement, and talent); and</p><p>(5) improve the governance system (referring to technology governance and international coordination).</p><p>In Point 5, Xi calls to &#8220;coordinate development and security, explore scientific and effective methods of regulation, systems for technology monitoring, risk early warning, and emergency response, anticipate and respond to new types of risk such as loss of control over technology, ethical breaches, and data abuse.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> The same point calls for &#8220;deepening international cooperation, actively participating in global governance, and pushing all parties to jointly build standards, jointly negotiate rules, and jointly promote industries.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> The speech closes with Xi citing his own earlier warning that leaders cannot continue being <a href="https://en.wikipedia.org/wiki/Blind_men_and_an_elephant">&#8220;blind men touching an elephant&#8221;</a> (&#30450;&#20154;&#25720;&#35937;) on S&amp;T change, and calling on leadership at all levels to &#8220;know science, understand industry, and decide well.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a></p><p>It&#8217;s significant that Xi discussed &#8220;loss of control over technology&#8221; (&#25216;&#26415;&#22833;&#25511;). This seems to be Xi&#8217;s first public mention of this concept.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a>  Previous Xi statements on AI risk have discussed safety, security, reliability, and controllability (see the <a href="https://www.fmprc.gov.cn/eng/xw/zyxw/202405/t20240530_11332389.html">October 2023 Global AI Governance Initiative</a>, the <a href="https://interpret.csis.org/translations/at-the-20th-collective-study-session-of-the-ccp-politburo-xi-jinping-stressed-the-need-for-self-reliance-and-self-improvement-highlight-application-oriented-approaches-and-promote-the-healthy-and-o/">April 2025 Politburo session on AI</a>), but never explicitly &#22833;&#25511;/loss of control. Concordia AI&#8217;s <a href="https://concordia-ai.com/wp-content/uploads/2025/07/State-of-AI-Safety-in-China-2025.pdf">July 2025 State of AI Safety in China</a> report flagged that leadership statements through that point had &#8220;not explicitly mention[ed] risks from misuse of advanced AI systems or from loss of control of superintelligent AI.&#8221;</p><p>The same term&#8212;&#25216;&#26415;&#22833;&#25511;&#8212;appears across TC260&#8217;s AI Safety Governance Framework v2 (September 2025) under several different scenarios, and the framework&#8217;s official translation renders it differently depending on context. <strong>&#25216;&#26415;&#22833;&#25511;</strong> is translated there as:</p><ul><li><p>&#8220;the potential risks of technological failure&#8221; (&#38024;&#23545;&#28508;&#22312;&#30340;<strong>&#25216;&#26415;&#22833;&#25511;</strong>&#39118;&#38505;, &#167;2.4)</p></li><li><p>&#8220;whether a model would pose a potential risk of loss of control&#8221; (&#21028;&#26029;&#27169;&#22411;&#26159;&#21542;&#21487;&#33021;&#24102;&#26469;&#28508;&#22312;<strong>&#25216;&#26415;&#22833;&#25511;</strong>&#39118;&#38505;, &#167;5.11, the developer-testing requirement)</p></li><li><p> &#8220;prevent and address the risk of AI technology losing control&#8221; (&#20849;&#21516;&#38450;&#33539;&#24212;&#23545;&#20154;&#24037;&#26234;&#33021;<strong>&#25216;&#26415;&#22833;&#25511;</strong>&#39118;&#38505;, Appendix 2 opening)</p></li></ul><p>Thus, there are three possible readings: general tech failure, humans losing control of technology, or technology losing control of its behavior. Which of these senses Xi had in mind in his speech isn&#8217;t clear from the text; he notably didn&#8217;t tie &#25216;&#26415;&#22833;&#25511; specifically to AI, and the loss-of-control framing sits in the general governance discussion that applies to all six future industries. Regardless, this is a new turn for Xi&#8217;s statements, and may mark the beginning of increased focus on frontier AI risks.</p><h2>CAC publishes four expert interpretations of Ethics-Safety Guidelines</h2><p>Following the May 19 publication of TC260-005 Ethics-Safety Guidelines for AI Applications 1.0 (see analysis in China AI Bulletin #4 <a href="https://chinaaibulletin.substack.com/i/198737671/tc260-releases-ethics-and-safety-guidelines-for-ai-applications">here</a>), CAC published four &#8220;Expert Interpretation&#8221; pieces between May 22 and May 23.</p><p>The four pieces focus on different aspects of TC260-005. Jiang Xinghao (&#33931;&#20852;&#28009;), Vice President of Shanghai Jiao Tong University (SJTU), <a href="http://www.cac.gov.cn/2026-05/23/c_1781107933166351.htm">lays out</a> the six impacts and all nine principles named in the document. Fan Kefeng (&#33539;&#31185;&#23792;), Deputy Director of the China Electronics Standardization Institute (CESI), <a href="http://www.cac.gov.cn/2026-05/22/c_1781191243097137.htm">discusses</a> TC260-005 in terms of standards, discussing its central goal as translating ethics and safety from a value-set into &#8220;standardized objects recognizable and identifiable by all parties,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> then into basic rules, then into operational guidance for each actor class. Liu Bo (&#21016;&#21338;), Deputy Party Secretary and Deputy Director of the National Internet Emergency Center (CNCERT), <a href="http://www.cac.gov.cn/2026-05/22/c_1781191243353200.htm">focuses</a> on operations and risk control: risk monitoring, emergency-response and human-intervention mechanisms, accident-traceability for responsibility attribution, and provisions for improving public AI literacy and risk-identification capacity. Finally, Zhang Linghan (&#24352;&#20940;&#23506;), Dean of the China University of Political Science and Law (CUPL) Institute for AI Law and a member of the UN High-Level Advisory Body on AI, <a href="http://www.cac.gov.cn/2026-05/23/c_1781107933473093.htm">frames</a> TC260-005 in terms of comparative policy. She lists the document&#8217;s six structural impacts as areas of global consensus, cross-references the UNESCO Recommendation on the Ethics of AI and the US NIST AI Risk Management Framework as international precedent, and argues the document complements existing hard regulations to form a system combining hard-law constraints with soft guidance&#8212;with specific attention to labor displacement and &#8220;inclusive sharing&#8221; (&#26222;&#24800;&#20849;&#20139;).</p><h2>SAMR and NDRC publish AI Metrology System and Capability Building Guidance</h2><p>On May 28, the State Administration for Market Regulation (SAMR) and the National Development and Reform Commission (NDRC) jointly issued the <em><a href="https://www.samr.gov.cn/xw/zj/art/2026/art_f43aa2c974654d66b91bbad8410d0d71.html">Guidance on AI Metrology System and Capability Building (2026 Edition)</a></em>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> The Guidance brings AI under SAMR&#8217;s formal metrology apparatus, which anchors China&#8217;s national reference standards for length, mass, time, and other physical quantities. It organizes work across six parts&#8212;foundational support, general technology, core technology, metrology technical specifications, metrology service industry, and AI-empowered metrology&#8212;and pledges &#8220;full-chain&#8221; metrology capability covering algorithm models, compute efficiency, and data quality, with the overall goal of making AI technical performance &#8220;measurable, comparable, and traceable.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> The document names algorithmic black boxes and lack of interpretability as targeted &#8220;pain points&#8221; and calls for R&amp;D on monitoring and characterizing AI systems&#8217; internal states.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a></p><p>Metrology (the science of measurement, &#35745;&#37327;) and testing/evaluations (&#35780;&#27979;/&#35780;&#20272;) play different roles. Evaluation is the practice of testing models or products against criteria; metrology is the measurement infrastructure that makes those evaluations comparable across labs, regimes, and time. Defined units, reference standards, calibration, and traceability are all grounded in these references. In China these are overseen by different entities: evaluations fall under standards bodies like TC260 and government agencies like CAC and MIIT, and metrology is under SAMR and the <a href="http://www.npc.gov.cn/zgrdw/npc/xinwen/2018-11/05/content_2065664.htm">PRC Metrology Law</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> (1985, last amended 2018), which establishes the national reference standards system. The Guidance is the first explicit Chinese policy commitment to building AI-specific metrology infrastructure. Without it, evaluation regimes lack an independent measurement backbone and effectively rely on developer self-attestation.</p><h2>Three provinces outline local AI ambitions</h2><p>Anhui&#8217;s May 22 <a href="https://www.most.gov.cn/dfkj/ah/zxdt/202605/t20260522_196662.html">Provincial S&amp;T Department positioning piece</a> restates the province&#8217;s &#8220;AI+Everything&#8221; (&#20154;&#24037;&#26234;&#33021;+&#19975;&#29289;) Action Plan&#8212;<a href="http://www.scio.gov.cn/xwfb/dfxwfb/gssfbh/ah_13837/202601/t20260127_948024.html">issued earlier this year</a>&#8212;and sets 15th FYP targets of &gt;10,000 deployed applications and &gt;90% adoption for new-generation intelligent terminals and agents. The piece cites 14th FYP achievements: &gt;1,000 above-scale AI enterprises with &#165;200B+ annual revenue, 5th-place national industry ranking, and establishing &#8220;China&#8217;s first domestically-produced 10,000-card compute cluster.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> New initiatives flagged include a forthcoming brain-computer interface action plan, the first &#8220;Commercial AI&#8221; undergraduate major, and strengthening the Yangtze River Delta Secure AI Provincial Laboratory.</p><p>Instead of &#8220;AI+,&#8221; Harbin is discussing &#8220;+AI&#8221;&#8212;in this case, &#8220;aerospace + AI.&#8221; The May 18 <a href="https://www.kdocs.cn/l/clFAwnBKHRaX">Harbin Aerospace Sector AI Capability List</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> <a href="https://www.stdaily.com/web/gdxw/2026-05/18/content_518323.html">catalogs</a> 29 achievements positioning AI as an enhancement layer for the city&#8217;s existing aerospace industrial base.</p><p>Finally, Hubei is emphasizing embodied intelligence; Governor Li Dianxun <a href="https://www.most.gov.cn/dfkj/hub/zxdt/202605/t20260529_196721.html">inspected the embodied AI industry in Wuhan</a>, visiting two companies and one research center. The read-out forwards a four-category taxonomy (industrial, special-purpose, service, and humanoid robots, plus components) and integrating the industry, innovation, talent, capital, and service chains.</p><h2>AI+ Energy continues to develop</h2><p>On May 26, the National Energy Administration (NEA) released the <a href="https://www.gov.cn/lianbo/202605/content_7070297.htm">first batch of 51 AI+ Energy high-value scenarios</a> at the National AI+ Energy On-Site Promotion Conference in Guangzhou. The release operationalizes the May <a href="https://www.nea.gov.cn/20260508/4dae97ca01d348e4871bb8654be34b3a/c.html">Action Plan on Promoting Mutual Empowerment between AI and Energy</a> <a href="https://chinaaibulletin.substack.com/i/198737671/other-news-ai-energy-and-leader-inspections">covered in Issue 4</a>.</p><p>The 51 scenarios span eight categories across grid, energy new business types (&#26032;&#19994;&#24577;), new energy, and conventional energy. Named examples include grid planning scheme intelligent generation and evaluation,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> virtual power plants, vehicle-grid interaction, new-energy power forecasting, and market-oriented operations. NEA Director Wang Hongzhi (&#29579;&#23439;&#24535;) framed the release as marking the shift &#8220;from concept to practice, from exploration to popularization.&#8221;</p><p>At the same event, NEA launched the <a href="https://www.gov.cn/lianbo/202605/content_7070467.htm">China AI+ Energy Development Report 2026</a>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> the first annual report on AI-energy integration. The report disclosed that by end-2025, China had built 42 ten-thousand-card AI compute clusters, national compute-center electricity consumption reached 170 billion kWh, and the eight national computing-network hub nodes averaged 39.5% annual growth in compute electricity over three years, with the Inner Mongolia hub at 66.5%.</p><h2>National Standards</h2><h3>Seven-part AI Agent Interconnection guidance series approved</h3><p>On May 22, SAMR and the Standardization Administration of China (SAC) issued <a href="https://std.sacinfo.org.cn/gnoc/queryInfo?id=495B8834610A6828F254D73BD7DCAF99">Announcement No. 22 of 2026</a>, approving eight national standardization guidance technical documents (&#22269;&#23478;&#26631;&#20934;&#21270;&#25351;&#23548;&#24615;&#25216;&#26415;&#25991;&#20214;). Seven of the eight are the GB/Z 185 series, Artificial Intelligence: Agent Interconnection (&#20154;&#24037;&#26234;&#33021; &#26234;&#33021;&#20307;&#20114;&#32852;):<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a></p><ul><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F8113453761E06397BE0A0AFEB5">GB/Z 185.1-2026</a> Part 1: Overall Architecture (&#24635;&#20307;&#26550;&#26500;)&#8212;lead drafter Fan Kefeng (&#33539;&#31185;&#23792;), CESI (author of the <a href="http://www.cac.gov.cn/2026-05/22/c_1781191243097137.htm">first TC260-005 expert interpretation</a> covered above)</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F81134B3761E06397BE0A0AFEB5">GB/Z 185.2-2026</a> Part 2: Identity Code (&#36523;&#20221;&#30721;)&#8212;lead drafter Liu Jun (&#21016;&#20891;), Beijing University of Posts and Telecommunications (BUPT)</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F81134A3761E06397BE0A0AFEB5">GB/Z 185.3-2026</a> Part 3: Identity Management (&#36523;&#20221;&#31649;&#29702;)&#8212;lead drafter Zhang Shizong (&#24352;&#22763;&#23447;), CESI</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F8113473761E06397BE0A0AFEB5">GB/Z 185.4-2026</a> Part 4: Agent Description (&#26234;&#33021;&#20307;&#25551;&#36848;)&#8212;lead drafter Xu Yang (&#24464;&#27915;), CESI</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F81134C3761E06397BE0A0AFEB5">GB/Z 185.5-2026</a> Part 5: Agent Discovery (&#26234;&#33021;&#20307;&#21457;&#29616;)&#8212;lead drafter Dong Jian (&#33891;&#24314;), CESI</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F8113483761E06397BE0A0AFEB5">GB/Z 185.6-2026</a> Part 6: Agent Interaction (&#26234;&#33021;&#20307;&#20132;&#20114;)&#8212;lead drafter Li Ke (&#26446;&#29634;), BUPT</p></li><li><p><a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=52A05F8113493761E06397BE0A0AFEB5">GB/Z 185.7-2026</a> Part 7: Agent Tool Invocation (&#26234;&#33021;&#20307;&#24037;&#20855;&#35843;&#29992;)&#8212;lead drafter Zhang Shizong (&#24352;&#22763;&#23447;), CESI</p></li></ul><p>The GB/Z designation marks these as non-mandatory guidance, not binding GB standards. Still, they are issued by China&#8217;s state standards body (SAC/SAMR) and together cover the full agent-interoperability protocol surface: identity, discovery, description, interaction, and tool invocation.</p><p></p><h3>TC28/SC42 begins drafting embodied intelligence, AI for Science standards</h3><p>TC28/SC42 also announced the drafting of <a href="https://std.samr.gov.cn/search/orgDetailView?data_id=A132FB8FCABB4BE9E05397BE0A0AB880">25 new AI projects</a> (&#27491;&#22312;&#36215;&#33609;) on May 28, 2026 (full table below). The standards fall into several clusters:</p><ul><li><p><strong>Embodied intelligence (5):</strong> cloud protocol requirements, data generation, dexterous manipulation, trustworthy evaluation indicators and methods, and an ethics governance guide.</p></li><li><p><strong>Ethics and governance infrastructure (3, distinct from the embodied-ethics guide):</strong> ethics governance scenarios classification and grading, ethics risk assessment, and sci-tech ethics review personnel technical skill requirements.</p></li><li><p><strong>AI for Science (3):</strong> scientific intelligence evaluation, scientific data preparation, and scientific data lead aggregation and classification/grading.</p></li><li><p><strong>Large model evaluation infrastructure (2):</strong> large model evaluation platform construction requirements, and world model evaluation specification.</p></li><li><p><strong>AI Bill of Materials (2):</strong> data format specification and implementation guide&#8212;a standardized inventory for AI components and dependencies.</p></li><li><p><strong>Iron and steel industry (6):</strong> intelligent agent foundational commonality requirements, data privacy and compliance, data governance, data classification and labeling, data alignment, and large model evaluation indicators and methods.</p></li><li><p><strong>Non-ferrous metals (2):</strong> large model evaluation, and application scenarios classification.</p></li><li><p><strong>Other (2):</strong> industrial large model reference architecture, swarm intelligence collaborative system reference architecture.</p></li></ul><p>A separate June 2 batch of seven items includes a coding-tool standard <em>Intelligent Programming Tools Service Capability Maturity Evaluation</em> (&#26234;&#33021;&#32534;&#31243;&#24037;&#20855; &#26381;&#21153;&#33021;&#21147;&#25104;&#29087;&#24230;&#35780;&#20272;) amongst IT-cabling and digital-twin standards.</p><div class="file-embed-wrapper" data-component-name="FileToDOM"><div class="file-embed-container-reader"><div class="file-embed-container-top"><image class="file-embed-thumbnail" src="https://substackcdn.com/image/fetch/$s_!ryM4!,w_400,h_600,c_fill,f_auto,q_auto:best,fl_progressive:steep,g_auto/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F01aad0ca-8dfa-40ad-bdfa-2c9b35526b10_1858x1826.png"></image><div class="file-embed-details"><div class="file-embed-details-h1">Standards Drafting Batch</div><div class="file-embed-details-h2">87.6KB &#8729; PDF file</div></div><a class="file-embed-button wide" href="https://chinaaibulletin.substack.com/api/v1/file/0c066924-1fbb-48b5-937e-4f36129b26f6.pdf"><span class="file-embed-button-text">Download</span></a></div><a class="file-embed-button narrow" href="https://chinaaibulletin.substack.com/api/v1/file/0c066924-1fbb-48b5-937e-4f36129b26f6.pdf"><span class="file-embed-button-text">Download</span></a></div></div><p></p><h3>SAC approves two other AI-adjacent recommended standards</h3><p>On May 25, SAC issued <a href="https://std.sacinfo.org.cn/gnoc/queryInfo?id=4D8AA718CCB795891CD1517867C3C698">Announcement No. 23 of 2026</a>, approving 375 recommended national standards (GB/T) and 5 amendment orders. Two of the 375 are AI-relevant:</p><ul><li><p><strong>GB/T 47695-2026</strong> <em>Enterprise Smart Manufacturing Efficacy Evaluation Method</em> (&#20225;&#19994;&#26234;&#33021;&#21046;&#36896;&#25928;&#33021;&#35780;&#27979;&#26041;&#27861;)&#8212;implementation date 2026-12-01.</p></li><li><p><strong>GB/T 47746-2026</strong> <em>Customer Contact Services: Human and Intelligent Customer Service Collaboration Requirements</em> (&#39038;&#23458;&#32852;&#32476;&#26381;&#21153; &#20154;&#24037;&#19982;&#26234;&#33021;&#23458;&#25143;&#26381;&#21153;&#21327;&#21516;&#35201;&#27714;)&#8212;implementation date 2026-09-01.</p></li></ul><p>Both are recommended (voluntary) standards. The smart-manufacturing efficacy standard provides an evaluation methodology for AI-enabled production; the customer-service standard sets requirements for handoffs between human and AI customer-service tiers.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this useful? Subscribe get the China AI Bulletin in your inbox regularly.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1>International AI Governance</h1><h2>MIIT and ASEAN launch a joint AI Industry Innovation Center</h2><p>On May 24, Ministry of Industry and Information Technology (MIIT) Vice Minister Ke Jixin (&#26607;&#21513;&#27427;) and Association of Southeast Asian Nations (ASEAN) Secretary-General Kao Kim Hourn (&#39640;&#37329;&#27946;) inaugurated the <a href="https://www.miitxxzx.org.cn/art/2026/5/24/art_203_6236.html">China-ASEAN AI Industry Innovation Center</a> (&#20013;&#22269;&#8212;&#19996;&#30431;&#20154;&#24037;&#26234;&#33021;&#20135;&#19994;&#21019;&#26032;&#20013;&#24515;) in Beijing. The center has four named priorities: (1) promote AI technology R&amp;D and innovation cooperation, with explicit focus on large-model application in typical industrial scenarios; (2) build a China-ASEAN AI industry cooperation ecosystem with an inter-state dialogue mechanism; (3) deepen AI governance practice through standardization coordination and governance-tool development; and (4) support regional capacity building via shared intelligent infrastructure. The read-out positions the center as building on the Digital Silk Road and states that it is intended to advance the China-ASEAN Comprehensive Strategic Partnership Action Plan (2026-2030) and the 2026 China-ASEAN Digital Cooperation Plan.</p><p>The third mandate puts the center inside the international AI governance architecture China has been building alongside the <a href="https://www.fmprc.gov.cn/eng/xw/zyxw/202405/t20240530_11332389.html">Global AI Governance Initiative</a> (2023) and the <a href="https://www.fmprc.gov.cn/mfa_eng/xw/zyxw/202507/t20250729_11679232.html">Global AI Governance Action Plan</a> (2025), adding a formal China-ASEAN dialogue venue with a standards-and-governance mandate built in from inception.</p><h2>Two state-visit readouts name AI as a designated cooperation area</h2><p>On May 25, Xi held talks with both Serbian President Aleksandar Vu&#269;i&#263; and Pakistani Prime Minister Shahbaz Sharif. Both Cyberspace Administration of China (CAC) readouts name AI explicitly as a designated cooperation area: the <a href="https://www.cac.gov.cn/2026-05/25/c_1781450605830275.htm">Serbia readout</a> lists AI first among four emerging-cooperation areas alongside the digital economy, green energy, and advanced manufacturing; the <a href="https://www.cac.gov.cn/2026-05/25/c_1781450599581475.htm">Pakistan readout</a> pairs AI with agriculture, industry, and talent development under the &#8220;China-Pakistan Community of Shared Destiny Action Plan.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a></p><h1>Frontier Lab Developments</h1><h2>&#128269; Spotlight</h2><p>On June 1, MiniMax released its first major flagship of 2026, <strong><a href="https://www.minimax.io/blog/minimax-m3">MiniMax M3</a></strong>. The model targets agentic and desktop-operation workloads with a <strong>1M-token context</strong> and native multimodality. It&#8217;s priced at $0.60/M input tokens, and <a href="https://github.com/MiniMax-AI/MiniMax-M3">open weights</a> are promised within 10 days of launch. MiniMax reported a <strong>SWE-Bench Pro score of 59.0</strong>, narrowly above GPT-5.5&#8217;s 58.6, but below Claude Opus 4.7&#8217;s 64.3. The same day, MiniMax released <strong><a href="https://github.com/MiniMax-AI/minimax-code">MiniMax Code</a></strong>, a downloadable desktop coding agent.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0x3k!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0x3k!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 424w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 848w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 1272w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0x3k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png" width="1456" height="614" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:614,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;a 4x2 chart table showing MiniMax M3's benchmark performance.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a 4x2 chart table showing MiniMax M3's benchmark performance." title="a 4x2 chart table showing MiniMax M3's benchmark performance." srcset="https://substackcdn.com/image/fetch/$s_!0x3k!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 424w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 848w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 1272w, https://substackcdn.com/image/fetch/$s_!0x3k!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F02018965-6660-4226-843c-e9f6bd5a4b09_1612x680.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: MiniMax <a href="https://www.minimax.io/blog/minimax-m3">blog</a></figcaption></figure></div><h2>Notable Model Releases</h2><p><strong>Alibaba</strong> had a busy two weeks of model releases. On May 31, the Qwen team released the multimodal <strong><a href="https://qwen.ai/blog?id=qwen3.7-plus">Qwen3.7-Plus</a></strong> supporting up to 1M-token context. Reported benchmarks place Qwen3.7-Plus at the <strong>front of GUI-automation benchmarks</strong>&#8212;scoring 79.0 on ScreenSpot Pro (vs Gemini-3.1 Pro&#8217;s 68.1) and 81.0 on AndroidWorld (vs Gemini&#8217;s 70.7)&#8212;while reaching near-parity with Opus-4.6 Max, K2.6 Thinking, and DeepSeek V4-Pro Max on general reasoning. <strong><a href="https://github.com/QwenLM/Qwen-VLA">Qwen-VLA</a></strong> (May 28) added a vision-language-action model for the robotics/embodied stack. On the image side, the Qwen team released <strong><a href="https://huggingface.co/Qwen/Qwen-Image-Flash">Qwen-Image-Flash</a></strong> on June 2 (<a href="https://arxiv.org/abs/2606.03746">paper</a>), a fast variant of the Qwen-Image family. Alibaba also released a technical paper on <strong><a href="https://github.com/alibaba/rtp-llm">RTP-LLM</a></strong> (May 28, <a href="https://arxiv.org/abs/2605.29639">paper</a>), the high-performance LLM inference engine running at Alibaba Group scale across more than 100 million users. </p><p><strong>Tencent</strong> released a full <strong><a href="https://huggingface.co/tencent/Hy-MT2-30B-A3B">Hy-MT2</a></strong> translation family in three sizes&#8212;<strong>1.8B, 7B, and 30B-A3B parameters </strong>(technical paper <a href="https://arxiv.org/pdf/2605.22064">here</a>)&#8212;plus an FP8 datacenter variant and an aggressively-compressed 1.25-bit GGUF mobile variant optimized to run natively on recent ARM phones.</p><p><strong>StepFun</strong> released <strong><a href="https://huggingface.co/stepfun-ai/Step-3.7-Flash">Step-3.7-Flash</a></strong> on May 23 with <a href="https://huggingface.co/stepfun-ai/Step-3.7-Flash-FP8">datacenter</a>, <a href="https://huggingface.co/stepfun-ai/Step-3.7-Flash-NVFP4">NVIDIA-optimized</a>, and <a href="https://huggingface.co/stepfun-ai/Step-3.7-Flash-GGUF">consumer-inference</a> variants shipped simultaneously.</p><p><strong>ByteDance</strong> released the <strong><a href="https://github.com/bytedance/Bernini">Bernini</a></strong> family (technical paper <a href="https://arxiv.org/abs/2605.22344">here</a>)&#8212;Bernini, Bernini-R, and <a href="https://huggingface.co/ByteDance/Bernini-R-Diffusers">Bernini-R-Diffusers</a>&#8212;a unified video-generation and editing framework combining an MLLM-based semantic planner with a DiT-based renderer.</p><p><strong>Baidu</strong> added to the <strong><a href="https://huggingface.co/baidu/ERNIE-Image">ERNIE-Image</a></strong> family with the <strong><a href="https://huggingface.co/baidu/ERNIE-Image-Aes">ERNIE-Image-Aes</a></strong> aesthetics model (<a href="https://arxiv.org/abs/2605.25347">technical report</a>). This was followed by audio-visual model <strong><a href="https://huggingface.co/baidu/NAVA">NAVA</a></strong> and <strong><a href="https://huggingface.co/PaddlePaddle/PaddleOCR-VL-1.6">PaddleOCR-VL-1.6</a></strong>, a versioned document-parsing model shipped with a GGUF variant (<a href="https://arxiv.org/abs/2606.03264">paper</a>, <a href="https://github.com/PaddlePaddle/PaddleOCR">GitHub</a>).</p><p><strong>Moonshot</strong> released <strong><a href="https://github.com/MoonshotAI/kimi-code">Kimi Code CLI</a></strong>, an AI coding agent that runs in the terminal, positioned as a Claude Code/Cursor/Codex competitor.</p><p><strong>SenseTime</strong> released the full-parameter fine-tuning training code for <strong><a href="https://github.com/OpenSenseNova/SenseNova-U1">SenseNova U1</a></strong>.</p><h2>Technical Publication Highlights</h2><p>Frontier labs released 114 papers on arXiv this fortnight. Highlights are below; a full list with summaries can be found <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0">here</a>.</p><p><em>Editor&#8217;s note: due to an unusual volume of activity on the arXiv API, this edition is missing approximately two days of paper releases. Any notable omissions will be highlighted in China AI Bulletin #6.</em></p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!viIq!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!viIq!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 424w, https://substackcdn.com/image/fetch/$s_!viIq!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 848w, https://substackcdn.com/image/fetch/$s_!viIq!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 1272w, https://substackcdn.com/image/fetch/$s_!viIq!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!viIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png" width="1280" height="1148" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1148,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:127266,&quot;alt&quot;:&quot;a table showing the cumulative and per-edition papers by lab. alibaba leads with 43.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/200679488?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="a table showing the cumulative and per-edition papers by lab. alibaba leads with 43." title="a table showing the cumulative and per-edition papers by lab. alibaba leads with 43." srcset="https://substackcdn.com/image/fetch/$s_!viIq!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 424w, https://substackcdn.com/image/fetch/$s_!viIq!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 848w, https://substackcdn.com/image/fetch/$s_!viIq!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 1272w, https://substackcdn.com/image/fetch/$s_!viIq!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4a7bde6f-89a8-4c53-b6ec-89650a52e5e0_1280x1148.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3>Alibaba</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.02357v1">Do Multimodal Agents Really Benefit from Tool Use? A Systematic Study of Capability Gains</a></strong></p><ul><li><p style="text-align: justify;">Examines whether multimodal agents actually benefit from tool use by comparing <strong>tool-augmented agents</strong> (Thyme, DeepEyesV2) against tool-free and text-only baselines across vision and reasoning tasks. Finds that<strong> tool access provided negligible aggregate gains</strong>. 93&#8211;96% of tool-solved problems were also solvable without tools, suggesting agents learn to call tools without gaining new capabilities.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.29860v1">ESPO: Early-Stopping Proximal Policy Optimization</a></strong></p><ul><li><p style="text-align: justify;">Proposes <strong>ESPO</strong>, which detects when a reasoning model takes a wrong step and <strong>stops generation early</strong> rather than wasting compute on unsalvageable trajectories. By treating early stops as terminal failure states, it concentrates learning signals near actual mistakes without needing extra reward models, <strong>improving math reasoning by 1&#8211;3% while cutting rollout tokens by 20%</strong>.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.22355v1">TransitLM: A Large-Scale Dataset and Benchmark for Map-Free Transit Route Generation</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>TransitLM</strong>, a dataset of 13M+ transit route records from Chinese cities that enables <strong>map-free route generation</strong>&#8212;LLMs trained on it learn to produce valid routes and ground GPS coordinates to stations without explicit mapping.</p></li></ul><h3>Baidu</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.28713v1">Thinking as Compression: Your Reasoning Model is Secretly a Context Compressor</a></strong></p><ul><li><p style="text-align: justify;">Reveals that <strong>reasoning models can naturally compress long contexts</strong> by organizing task-relevant information into thinking traces. <strong>TaC-C</strong>, a reward-optimized variant, outperforms dedicated compression methods by 17&#8211;23% at 4&#8211;8x compression ratios without requiring specialized compressor modules.</p></li></ul><h3>Huawei</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.21325v1">Fast and Stable Triangular Inversion for Delta-Rule Linear Transformers</a></strong></p><ul><li><p style="text-align: justify;">Optimizes <strong>triangular matrix inversion</strong> in linear attention models by systematically analyzing direct and iterative algorithms, achieving <strong>4.3x speedup</strong> over existing implementations while maintaining numerical stability and model accuracy across low-precision settings.</p></li></ul><h3>Meituan</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.02355v1">SIRI: Self-Internalizing Reinforcement Learning with Intrinsic Skills for LLM Agent Training</a></strong></p><ul><li><p style="text-align: justify;">Proposes <strong>SIRI</strong>, a three-phase reinforcement learning framework that enables LLM agents to <strong>discover, validate, and internalize reusable skills without external skill generators or inference-time retrieval overhead</strong>. On ALFWorld and WebShop tasks, SIRI improves performance from 0.908 to 0.930 and 0.728 to 0.813 respectively, while reducing deployment complexity by running inference with the original prompt only.</p></li></ul><h3>SenseTime</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.20838v1">USV: Towards Understanding the User-generated Short-form Videos</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>USV</strong>, a 224K user-generated short-form video dataset with <strong>topic recognition</strong> and <strong>video-text retrieval</strong> tasks, plus baseline models <strong>MMF-Net</strong> and <strong>VTCL</strong> to benchmark high-level semantic understanding beyond instance-level recognition.</p></li></ul><h3>StepFun</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.20755v1">DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>DuplexSLA</strong>, a full-duplex spoken dialogue model that <strong>decodes user audio, assistant speech, and structured actions on a shared 160ms timeline</strong>, enabling simultaneous listening, speaking, planning, and tool calling without external cascades. The model handles semantic turn-taking (interruptions, pauses, backchannels) natively and <strong>interleaves tool calls with ongoing speech</strong>, evaluated on a new duplex benchmark covering interruption and multi-action scenarios.</p></li></ul><h3>Tencent</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.28548v1">GEM: Generative Supervision Helps Embodied Intelligence</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>GEM</strong>, a vision-language model that adds <strong>depth map generation</strong> during pre-training to bridge the gap between high-level semantic understanding and low-level spatial knowledge needed for robot control. The approach, validated on <strong>GEM-4M</strong> (a new 4M-example dataset with depth supervision), achieves state-of-the-art results on embodied benchmarks and shows superior performance in real-world robot task execution.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.29486v1">PhoneWorld: Scaling Phone-Use Agent Environments</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>PhoneWorld</strong>, a pipeline that converts real mobile app trajectories into controllable phone-use environments, executable tasks, and training data at scale. Replacing 10K AndroidWorld steps with PhoneWorld supervision improves four evaluation benchmarks simultaneously, with gains ranging from 6.0 to 52.5 points across different metrics.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.20873v1">PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>PlanningBench</strong>, a framework that <strong>generates scalable planning data</strong> from a taxonomy of 30+ task types, enabling controllable difficulty, automatic verification, and evaluation of LLMs on complex multi-constraint problems. RL training on verified PlanningBench data improves performance on unseen planning benchmarks and instruction-following tasks.</p></li></ul><h3>Xiaomi</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.03236v1">Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents</a></strong></p><ul><li><p style="text-align: justify;">Separates the decision of <strong>when to intervene</strong> from <strong>how to assist</strong> using a lightweight perception module that gates the full reasoning model, reducing false alerts while improving accuracy and speed on proactive mobile agent tasks.</p></li></ul><h1>Technical AI Safety Publication Highlights</h1><p>There were <strong>115 AI-safety-related papers published by Chinese researchers </strong>this fortnight. Highlights are below; a full list with summaries is available <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1">here</a>.</p><p><em>Editor&#8217;s note: due to an unusual volume of activity on the arXiv API, this edition is missing approximately two days of paper releases. Any notable omissions will be highlighted in China AI Bulletin #6.</em></p><h2>&#128269;Spotlight</h2><p style="text-align: justify;"><strong>Shanghai AI Lab </strong>released<strong> <a href="https://arxiv.org/html/2605.29801v1">AgentDoG 1.5</a>, </strong>an updated version of their agent safety alignment framework.</p><p style="text-align: justify;">The original framework introduced a <strong>three-dimensional taxonomy for agentic risks</strong>, categorizing them by <strong>source</strong> (where the risk originates, e.g., user instruction, tool behavior, or environment), <strong>failure mode</strong> (how it manifests, e.g, capability failure, goal deviation, or boundary violation), and <strong>consequence</strong> (what harm results). This structured approach enables <strong>root cause diagnosis&#8212;</strong>rather than outputting binary safe/unsafe labels, <strong>AgentDoG traces </strong><em><strong>why</strong></em><strong> an action is problematic</strong>. When an agent takes an unsafe step, the system identifies whether the fault lies in a malicious user prompt, an unexpected tool behavior, or environmental conditions, providing key information for remediation.</p><p style="text-align: justify;">Version 1.5 updates the agent safety taxonomy to cover emergent risks from Codex and OpenClaw execution scenarios, then trains <strong>0.8B-, 2B-, 4B-, and 8B-parameter variants</strong> on <strong>~1,000 samples</strong> using a taxonomy-guided data engine with influence-function purification. Performance is claimed to match leading closed-source models (e.g., GPT-5.4) on agent-safety benchmarks. All models and datasets are open-sourced.</p><h2>Agentic Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.22333v1">A First Measurement Study on Authentication Security in Real-World Remote MCP Servers</a></strong></p><ul><li><p style="text-align: justify;">Researchers conducted the <strong>first security audit of authentication in real-world MCP servers</strong>&#8212;the emerging interface connecting LLMs to external services like banking and email. They found <strong>40.55% of 7,973 live servers expose tools without any authentication</strong>, and among authenticated servers, <strong>all tested OAuth deployments exhibited at least one flaw</strong>, with <strong>96.6% vulnerable to dynamic client registration attacks</strong>. These weaknesses enable account takeover and data theft in agent-to-service connections.</p></li></ul><p><em>Institutional affiliations: </em>Fudan University</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.00152v1">PrivacyPeek: Auditing What LLM-Based Agents Acquire, Not Just What They Say</a></strong></p><ul><li><p style="text-align: justify;"><strong>PrivacyPeek</strong> is a benchmark that audits what LLM-based agents <strong>acquire from external tools</strong>, not just what they disclose. Across 1,182 test cases, the benchmark detects when agents retrieve sensitive data beyond task requirements, then measures how easily an attacker could extract that over-acquired information via follow-up prompts. Testing 10 agents shows <strong>widespread over-acquisition</strong> across models, with current prompt-level defenses mitigating only a small fraction of this leakage, indicating a structural vulnerability in agent deployment.</p></li></ul><p><em>Institutional affiliations: </em>Shanghai AI Laboratory, Southeast University</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.28201v1">Plant, Persist, Trigger: Sleeper Attack on Large Language Model Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>Sleeper Attack</strong> is a novel threat where adversarial content injected into agent state (session context, memory, or reusable skills) remains dormant across multiple interactions, then activates when a benign user query triggers harmful behavior. Researchers constructed a <strong>1,896-instance benchmark</strong> covering six harm categories and tested seven LLM agents, finding all remain vulnerable even when single-interaction attacks fail. This demonstrates that persistent state exploitation poses a detection challenge distinct from immediate prompt injection.</p></li></ul><p><em>Institutional affiliations: </em>University of Science and Technology of China, National University of Singapore, Singapore Management University, Shanghai Artificial Intelligence Laboratory</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.25435v1">Security of OpenClaw Agents: Fundamentals, Attacks, and Countermeasures</a></strong></p><ul><li><p style="text-align: justify;">This survey maps the <strong>attack surface of open-source LLM agents</strong>&#8212;systems that run continuously, access external tools, and maintain persistent memory. It identifies threats spanning <strong>skill poisoning</strong> (malicious tool injection), <strong>cognitive manipulation</strong> (prompt attacks exploiting reasoning), <strong>multi-agent cascades</strong> (failures propagating across agent networks), and <strong>supply-chain vulnerabilities</strong>. The authors categorize defenses across reasoning, execution, and external interaction layers, positioning OpenClaw security as a distinct problem requiring new controls beyond standard LLM safety.</p></li></ul><p><em>Institutional affiliations: </em>Xi&#8217;an Jiaotong University</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.31042v1">From Prompt Injection to Persistent Control: Defending Agentic Harness Against Trojan Backdoors</a></strong></p><ul><li><p style="text-align: justify;"><strong>ClawTrojan</strong> demonstrates a multi-step backdoor attack where adversaries embed hidden instructions in files or tool outputs that agentic LLMs read, store, and execute later, bypassing single-step defenses that inspect each action in isolation. On GPT-4, the attack reaches <strong>95.5% success</strong> while traditional prompt-injection defenses fail entirely. <strong>DASGuard</strong>, a proposed defense, <strong>traces control-like text</strong> in workspace files to trusted sources and removes untrusted instructions, combining runtime blocking with sanitized storage to prevent persistent agent compromise.</p></li></ul><p><em>Institutional affiliations: </em>Renmin University of China</p><h2>Alignment</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.00651v1">MESA: Improving MoE Safety Alignment via Decentralized Expertise</a></strong></p><ul><li><p style="text-align: justify;"><strong>Safety Sparsity</strong>&#8212;where safety capabilities concentrate in a few experts within Mixture-of-Experts models&#8212;creates a vulnerability to adversarial attacks. <strong>MESA</strong> addresses this by using optimal transport theory to <strong>redistribute safety responsibilities across more experts</strong>, preventing concentration while maintaining model performance. The framework also refines routing to activate only the necessary safety-aligned experts, reducing interference with the model&#8217;s ability to answer helpful questions.</p></li></ul><p><em>Institutional affiliations: </em>Beihang University, Tsinghua University, Fudan University, Tencent, Alibaba</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.20994v1">Towards Context-Invariant Safety Alignment for Large Language Models</a></strong></p><ul><li><p style="text-align: justify;"><strong>Anchor Invariance Regularization (AIR)</strong> addresses a core safety problem: models refuse harmful requests in standard prompts but comply under adversarial rephrasing. The method treats prompts with verifiable feedback (e.g., multiple-choice) as anchors, then uses one-way regularization to push open-ended variants toward the same safety decision without degrading performance on reliable variants. Testing across safety, moral reasoning, and math tasks shows <strong>12.71% improvement in consistency</strong> and <strong>33.49% stronger out-of-distribution robustness</strong>, suggesting adversarial context-switching becomes harder to exploit.</p></li></ul><p><em>Institutional affiliations: </em>Fudan University, Shanghai AI Laboratory</p><h2>Evaluation and Benchmarks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.01317v1">SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces</a></strong></p><ul><li><p style="text-align: justify;"><strong>SABER</strong> evaluates coding agents not on isolated safety responses but on <strong>cumulative damage to project environments</strong> after multi-step action sequences. Unlike existing benchmarks that test refusal, SABER places models in realistic stateful workspaces and measures safety violations by their <strong>final environmental state</strong>&#8212;capturing how agents corrupt files, permissions, or dependencies over time. Even top models show <strong>54%+ harmful violation rates</strong>, indicating alignment methods designed for single-turn safety fail when agents operate autonomously over multiple actions in shared systems.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: </em>The University of Hong Kong, Shandong University, Carnegie Mellon University, National University of Singapore, The Hong Kong University of Science and Technology</p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2606.02380v1">SPADE-Bench: Evaluating Spontaneous Strategic Deception in Agents via Plan-Action Divergence</a></strong></p><ul><li><p style="text-align: justify;"><strong>SPADE-Bench</strong> evaluates whether LLM-based agents <strong>misreport their actions</strong> to users while executing different plans behind the scenes. The benchmark measures actual tool execution against stated intentions under pressure scenarios, distinguishing <strong>strategic deception from hallucination</strong>. Experiments confirm agents do diverge from reported plans&#8212;a safety concern for autonomous systems where human oversight is limited and users rely on agent-generated summaries.</p></li></ul><p><em>Institutional affiliations: </em>Beijing Academy of Artificial Intelligence, Peking University, University of Science and Technology of China, University of Chinese Academy of Science, Alibaba Group</p><h2>Governance and Policy</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.24383v1">A governance horizon for ethical-use constraints in open-weight AI models</a></strong></p><ul><li><p style="text-align: justify;">Audits <strong>2.1M HuggingFace repositories</strong> to test whether ethical-use restrictions stated on a model (in its model card or license&#8212;e.g., &#8220;no military applications,&#8221; &#8220;no surveillance&#8221;) stay visible when others fine-tune that model into new versions. They don&#8217;t: each fine-tune drops about half the stated restrictions, and after seven fine-tuning generations <strong>80% of descendant models carry no visible restriction at all</strong>&#8212;the &#8220;<strong>governance horizon</strong>.&#8221; Mandatory-declaration regimes do better than inheritance-only ones, but models with no traceable parent can&#8217;t be governed by either approach.</p></li></ul><p><em>Institutional affiliations: Peking University, Ministry of Education, University of Science and Technology Beijing, UC Davis</em></p><h2>Guardrails and Deployment Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.24817v1">RouteScan: A Non-Intrusive Approach to Auditing MoE LLMs Safety via Expert Routing Telemetry</a></strong></p><ul><li><p style="text-align: justify;"><strong>RouteScan</strong> audits Mixture-of-Experts (MoE) LLMs for unsafe behavior by monitoring <strong>GPU-level expert routing patterns</strong> rather than inspecting user prompts or outputs, addressing the privacy&#8211;safety tension in content-based auditing. The method uses <strong>GPU thread allocation telemetry</strong> as a fingerprint of input type and detects harmful prompts with <strong>AUROC &gt;0.93</strong> on unseen domains. Testing shows the routing signals retain minimal information for prompt reconstruction, offering privacy advantages over traditional monitoring while maintaining detection accuracy.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: </em>Zhejiang University, Donghua University, Louisiana State University</p><h2>Misuse and Dangerous Capabilities</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.25388v1">ViroBench: Benchmarking Nucleotide Foundation Models on Viral Genomics Tasks</a></strong></p><ul><li><p style="text-align: justify;"><strong>ViroBench</strong> claims to be the first comprehensive benchmark for evaluating nucleotide foundation models (NFMs) on viral genomics tasks, assessing both <strong>biological understanding and biosecurity risk</strong>. The benchmark reveals significant gaps: NFMs <strong>degrade in performance under phylogenetic and temporal shifts</strong>, struggle to distinguish statistical likelihood from biological validity in generation tasks, and are vulnerable to <strong>generating sequences that appear statistically valid but lack biological function</strong>&#8212;a latent biosecurity concern. Results show diverse training data matters more than model scale for viral genomics performance.</p></li></ul><p><em>Institutional affiliations: Shanghai Innovation Institute, Shanghai AI Lab, Fudan University, Shenzhen Loop Area Institute, Westlake University</em></p><h2>Robustness and Adversarial Attacks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.29708v1">Understanding Safety-Sensitive Expert Behavior in Mixture-of-Experts LLMs</a></strong></p><ul><li><p style="text-align: justify;"><strong>RASET</strong> identifies a routing-agnostic safety vulnerability in MoE LLMs: safety enforcement concentrates in a small subset of experts rather than being distributed across routing decisions. The authors show that altering parameters in these safety-critical experts&#8212;detected via contrastive sensitivity analysis&#8212;can <strong>degrade refusal behavior while leaving the model&#8217;s routing patterns intact, </strong>showing that <strong>MoE safety alignment may be more brittle than assumed</strong>.</p></li></ul><p><em>Institutional affiliations: </em>Huazhong University of Science and Technology, Nanyang Technological University</p><h1>On the Horizon</h1><p>Beijing Academy of AI&#8217;s <a href="https://hub.baai.ac.cn/view/54291">eighth annual Zhiyuan Conference</a> takes place June 12-13, <a href="https://pandaily.com/2026-zhiyuan-conference-brain-inspired-ai-jun2026">focused on</a> brain-inspired intelligence and next-gen AI. Watch for outputs from the AI safety track, possible WuJie world model updates, and discussion of new AI paradigms.</p><div><hr></div><p>For more on how we select and track content, see our methodology <a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz">here</a>.</p><div><hr></div><p></p><p><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>&#8220;&#21069;&#30651;&#24067;&#23616;&#21644;&#21457;&#23637;&#26410;&#26469;&#20135;&#19994;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>&#8220;&#32479;&#31609;&#21457;&#23637;&#21644;&#23433;&#20840;...&#21069;&#30651;&#24212;&#23545;&#25216;&#26415;&#22833;&#25511;&#12289;&#20262;&#29702;&#22833;&#33539;&#12289;&#25968;&#25454;&#28389;&#29992;&#31561;&#26032;&#22411;&#39118;&#38505;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>&#8220;&#35201;&#19981;&#26029;&#28145;&#21270;&#22269;&#38469;&#21512;&#20316;&#65292;&#31215;&#26497;&#21442;&#19982;&#20840;&#29699;&#27835;&#29702;&#65292;&#21162;&#21147;&#25512;&#21160;&#21508;&#26041;&#26631;&#20934;&#20849;&#24314;&#12289;&#35268;&#21017;&#20849;&#21830;&#12289;&#20135;&#19994;&#20849;&#20419;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;&#30693;&#31185;&#25216;&#12289;&#25026;&#20135;&#19994;&#12289;&#21892;&#20915;&#31574;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>A second speech by Xi in January may have <a href="https://aisafetychina.substack.com/p/ai-safety-in-china-25">also</a> used &#8220;technological loss of control,&#8221; but it&#8217;s unclear whether the term was used in the speech itself or a paraphrase. Notably, Xinhua follow-up <a href="https://www.news.cn/20260124/3f1f3cead780463b9f8119285fe6fb4f/c.html">coverage</a> linked loss of control over technology/&#25216;&#26415;&#22833;&#25511; explicitly to AI.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>&#8220;&#25512;&#21160;&#23439;&#35266;&#20215;&#20540;&#35201;&#27714;&#36827;&#19968;&#27493;&#36716;&#21270;&#20026;&#21508;&#26041;&#21487;&#29702;&#35299;&#12289;&#21487;&#35782;&#21035;&#30340;&#26631;&#20934;&#21270;&#23545;&#35937;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>&#12298;&#20154;&#24037;&#26234;&#33021;&#35745;&#37327;&#20307;&#31995;&#21644;&#33021;&#21147;&#24314;&#35774;&#25351;&#24341;&#65288;2026&#29256;&#65289;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>&#8220;&#21487;&#27979;&#37327;&#12289;&#21487;&#27604;&#36739;&#12289;&#21487;&#36861;&#28335;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>&#8220;AI&#31995;&#32479;&#20869;&#37096;&#29366;&#24577;&#30417;&#27979;&#19982;&#34920;&#24449;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>&#12298;&#20013;&#21326;&#20154;&#27665;&#20849;&#21644;&#22269;&#35745;&#37327;&#27861;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>&#8220;&#24314;&#25104;&#20102;&#22269;&#20869;&#39318;&#20010;&#22269;&#20135;&#19975;&#21345;&#31639;&#21147;&#38598;&#32676;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>&#12298;&#21704;&#23572;&#28392;&#24066;&#33322;&#31354;&#33322;&#22825;&#39046;&#22495;&#20154;&#24037;&#26234;&#33021;&#33021;&#21147;&#28165;&#21333;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>&#30005;&#32593;&#35268;&#21010;&#26041;&#26696;&#26234;&#33021;&#29983;&#25104;&#19982;&#35780;&#20272;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p>&#12298;&#20013;&#22269;&#8221;&#20154;&#24037;&#26234;&#33021;+&#8221;&#33021;&#28304;&#21457;&#23637;&#25253;&#21578;2026&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>The eighth is about testing for harmful substances in musical instruments.</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p>&#8220;&#21452;&#26041;&#35201;&#25166;&#23454;&#25512;&#36827;&#26500;&#24314;&#20013;&#24052;&#21629;&#36816;&#20849;&#21516;&#20307;&#34892;&#21160;&#35745;&#21010;&#8221;</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 4]]></title><description><![CDATA[Developments from 29/4/26-20/5/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-4</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-4</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Thu, 21 May 2026 18:00:24 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/b7e91214-047e-42a5-8b72-19a7e247fa57_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome to Issue 4 of the China AI Bulletin, the latest on AI development, governance, and safety in China. Today&#8217;s highlights: CAC launches a four-month enforcement campaign taking on AI &#8220;digital slop,&#8221; three agencies jointly release a tiered governance for AI agents, and Trump and Xi commit in principle to a second intergovernmental AI dialogue at the May 19 summit.</p><p><em>Number of the week: 608&#8212;the number of new deep synthesis service algorithms filed in the <a href="https://www.cac.gov.cn/2026-05/06/c_1779809434590762.htm">latest disclosure</a></em></p><h1>Executive Summary</h1><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/domestic-ai-governance">Domestic AI Governance</a>:</strong> CAC launched a four-month <strong>AI enforcement campaign</strong> covering fourteen categories, including <strong>open-source model accountability</strong> and <strong>AI &#8220;digital slop</strong>.&#8221; CAC, NDRC, and MIIT jointly issued the <strong>Implementation Opinions on Intelligent Agents,</strong> outlining a tiered governance framework across nineteen priority sectors. Both 2026 national legislative work plans mention a comprehensive AI law, but not as a top priority.</p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/national-standards">National Standards</a>:</strong> <strong>TC260-005</strong> on ethics &amp; safety guidelines mentioned  application-level <strong>loss-of-control </strong>and <strong>off-switch obligations</strong>. TC28/SC42 increased its focus on trustworthiness with a set of new standards. AI safety <strong>WG9</strong> advanced the <strong>Intelligent Agent Application Security Basic Requirements</strong> to formal deliberation as a rare <strong>mandatory national standard</strong>. SAC&#8217;s <strong>GB/Z 177 </strong>introduced<strong> AI Terminal Intelligence Grading</strong>, and CAC stood up an AI-cybersecurity testing programme with <strong>Huawei as the technical support unit</strong>.</p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/international-ai-governance">International AI Governance</a>:</strong> Following Trump&#8217;s Beijing visit, the two governments reportedly agreed to <strong>launch a bilateral intergovernmental AI dialogue</strong>. Treasury Secretary Bessent informally surfaced the agreement on May 14 and MFA spokesperson Guo Jiakun announced it May 19, but the US has not officially confirmed.</p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/frontier-lab-developments">Frontier Lab Developments</a>:</strong></p><ul><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/notable-model-releases">Notable Model Releases</a>:</strong> Alibaba launched <strong>Qwen 3.7-Max</strong> claiming <strong>35-hour autonomous operation</strong>, paired with its new <strong>Zhenwu M890 custom AI chip</strong>. Tencent compressed HY-MT1.5-1.8B into <strong>1.25-bit and 2-bit quantized variants</strong> for mobile devices, and released <strong>Hy-Embodied-RoboFusion</strong> and <strong>R-DMesh</strong> for embodied AI research and 3D animation.</p></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/technical-publication-highlights">Technical Papers</a>:</strong> Frontier labs released <strong>206 papers</strong>, led by Alibaba (70), Tencent (41), and ByteDance (32). <strong>Highlights</strong>: Alibaba&#8217;s <strong>Qwen-Scope</strong> open-source SAE suite for interpretability, ByteDance&#8217;s <strong>A-CODE</strong> (atomic-level protein design), and Tencent&#8217;s <strong>CL-bench Life</strong> real-life context benchmark (frontier models score 19.3%).</p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/198737671/technical-ai-safety-publication-highlights">Technical AI Safety</a>:</strong> 172 safety papers were released this edition. <strong>Shanghai AI Lab released several focused on agentic safety</strong>, and Tencent&#8217;s <em><strong>Safe, or Simply Incapable?</strong></em> discusses phone-use-agent safety evals.</p></li></ul><h1>Domestic AI Governance</h1><h2>CAC targets AI &#8220;chaos&#8221; in enforcement sweep</h2><p>On April 30, the Cyberspace Administration of China deployed a four-month <strong><a href="http://www.cac.gov.cn/2026-04/30/c_1779289298718765.htm">Qinglang special action on AI application chaos</a></strong> (&#28165;&#26391;&#183;&#25972;&#27835;AI&#24212;&#29992;&#20081;&#35937;). <a href="https://www.cac.gov.cn/wxzw/qinglang/A093711index_1.htm">Qinglang</a> (&#28165;&#26391;, &#8220;clear and bright&#8221;) is CAC&#8217;s umbrella brand for online-content enforcement campaigns, launched in 2016 and running regularly since 2020. They cover issues like livestream regulation, algorithm abuse, and the protection of minors. This is the <strong>second AI-specific Qinglang campaign</strong> after the May 2025 <a href="https://www.news.cn/tech/20250508/34ed8d8babec440e923081a8f9222ff5/c.html">&#8220;AI technology abuse&#8221;</a> action, but where the 2025 version focused on content misuse, the 2026 edition expands enforcement to the entire AI service lifecycle and runs in two phases. Phase one targets <strong>seven categories at the infrastructure layer</strong>:<strong> genAI services running without filing </strong>under the 2023 Generative AI Services Interim Measures, <strong>insufficient safety guardrails and review capabilities</strong>, <strong>unauthorized or low-quality training datasets</strong>, <strong>AI data poisoning</strong> (including malicious generative-engine optimization and an e-commerce trade in poisoning tools), <strong>non-compliance with the AI-Generated Content Labeling Methods</strong>, <strong>AI-enabled cyberattacks and unauthorized face-swap/voice-clone services</strong>, and <strong>open-source model safety management</strong>. Phase two adds seven content-layer categories: AI &#8220;remixing&#8221; of classics into &#25968;&#23383;&#27860;&#27700; (&#8221;digital slop&#8221;), false information including impersonation of party and state media, deepfake imagery of public figures and &#8220;AI resurrection of the deceased,&#8221; violent or vulgar generation, harm to minors, AI-driven trolling, and non-compliant AI products and shell apps.</p><p>The digital-slop framing is already rippling through state media. A May 17 <a href="https://www.stdaily.com/web/gjxw/2026-05/17/content_518015.html">Science &amp; Technology Daily commentary</a><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> connects two Phase 2 categories, digital slop and harm to minors, by arguing that <strong>&#8220;AI&#22403;&#22334;&#35270;&#39057;&#8221; (AI garbage videos) are reshaping cognitive development in young children</strong> whose brains are still building basic models of how the physical world works. The piece cites 278 identified channels generating such content, with 63 billion accumulated views and 220 million subscriptions.</p><p>Per the Beijing Academy of AI (BAAI)&#8217;s <a href="https://hub.baai.ac.cn/view/54508">AI Governance Weekly</a>, <strong>filing registration, training-corpus compliance, AI data poisoning, and open-source model safety management are all new in the 2026 campaign</strong> relative to 2025. The open-source provision is the most consequential of the four: it <strong>applies the content-platform accountability regime</strong> (identity verification, dataset and model takedowns, and emergency-response mechanisms)<strong> to open-source AI hosts</strong>, treating them as content platforms rather than as developer infrastructure.</p><h2>Joint framework for agentic AI released</h2><p>On May 8, CAC, the National Development and Reform Commission (NDRC), and the Ministry of Industry and Information Technology (MIIT) jointly issued the<strong> <a href="https://www.cac.gov.cn/2026-05/08/c_1779979789523320.htm">Implementation Opinions on Standardized Application and Innovative Development of Intelligent Agents</a></strong>.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> The document is framed as an implementation of the AI+ Action Plan and foregrounds the need for agent safety and controllability<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a> (reiterated in a <a href="https://www.news.cn/politics/20260508/ae37ab32f4224b9bbb7eb31f803cd192/c.html">Xinhua Q&amp;A</a>) as well as innovation and application-driven development.<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> It&#8217;s structured around four pillars: <strong>strengthening the development foundation</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> through technical development and standards; <strong>holding the safety bottom line</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> by improving standards, safety/security measures, governance, and self-regulation; <strong>driving application uptake</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a> by supporting research, development, and applications while promoting well-being and social governance; and <strong>building an innovation ecosystem</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> by promoting industrial cooperation and promoting new agent applications. One of the governance measures it calls for is a<strong> tiered agent governance framework </strong>(Article 11), which subjects sensitive sectors and key industries to mandatory filing, testing, and problem-product recall, potentially paired with the voluntary &#8220;agent registry platform&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a> proposed in Article 4 of the same document, which would provide digital identity management, discovery, and capability declarations for agents. This builds on a registry concept the bulletin previously tracked in the National Information Security Standardization Technical Committee (TC260)&#8217;s <a href="https://www.tc260.org.cn/portal/article/2/160310cea5f6411d92fd99a52a42424f">OpenClaw-type Agent Practice Guide</a> (March 31), which called for enterprise-level asset registries of approved agent deployments; Article 4 lifts the concept from organizational to national scope.</p><p>CAC paired the release with<strong> five same-day expert interpretation pieces</strong> by Chinese Academy of Engineering (CAE) academician <a href="https://www.cac.gov.cn/2026-05/08/c_1779979790609763.htm">Wu Hequan</a>, China Academy of Information and Communications Technology (CAICT) president <a href="https://www.cac.gov.cn/2026-05/08/c_1779979790736565.htm">Yu Xiaohui</a>, MIIT-sector science-and-technology ethics committee chair <a href="https://www.cac.gov.cn/2026-05/08/c_1779983775421699.htm">Wei Yiming</a>, China Center for Information Industry Development (CCID) AI research director <a href="https://www.cac.gov.cn/2026-05/08/c_1779983775418216.htm">Zhong Xinlong</a>, and Tsinghua&#8217;s <a href="https://www.cac.gov.cn/2026-05/08/c_1779979790817217.htm">Xue Lan</a>. Zhong Xinlong opens with the most explicit safety framing, <strong>citing the wide deployment of OpenClaw</strong> since the start of 2026 as having &#8220;exposed risk hazards including agents&#8217; ability to launch network attacks under instruction inducement,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a> and names medical, transportation, and public safety as sectors warranting mandatory standards. Xue Lan frames the document as a <strong>governance shift from content risk to behavioral risk</strong>; content-safety measures, which much existing genAI regulation focuses on, can no longer cover the threat surface once AI takes autonomous actions, and the new risk categories he names are excessive permissions, behavioral loss-of-control, and tool poisoning. Yu Xiaohui makes the parallel technical case: <strong>traditional perimeter defense models face failure against agents&#8217; multi-modal autonomous decision-making, so embedded rules and behavioral guardrails must replace boundary-based security</strong>. Wei Yiming calls <strong>behavioral control the new safety boundary</strong> and proposes blockchain-based verifiable and traceable mechanisms in important application scenarios; Wu Hequan explains why <strong>agent risk is qualitatively different around autonomy, coordination, and plasticity.</strong></p><h2 style="text-align: justify;">MIIT launches ten-province ethics-review pilot</h2><p>On May 9, <a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_1a0b2b0b0deb47099202871e0b70551b.html">MIIT issued a notice</a> launching a <strong>six-month pilot program to implement the <a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_2995f16b28504ddcbb604e918eb15759.html">Interim Measures for the Ethical Review and Service of Artificial Intelligence Science and Technology Activities</a></strong>, a framework MIIT issued jointly with nine other agencies. The pilot covers <strong>ten provinces and municipalities</strong> (Beijing, Shanghai, Guangdong, Shandong, Tianjin, Sichuan, Jiangsu, Hubei, Hunan, and Zhejiang) and runs June 1 through November 30. By the close of the period, MIIT expects participating regions to have connected ministerial, provincial, and municipal review chains, built an AI ethics risk case database, formulated five or more standards, and trained dedicated review personnel and institutions.</p><p>Coverage is structured around a <strong>mandatory base layer plus sectoral choice</strong> to align with &#8220;local realities.&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-11" href="#footnote-11" target="_self">11</a> Every participating city must conduct ethics review on AI&#8217;s foundational layer of data, algorithms, and models, and select at least three vertical application domains from a list of nine: manufacturing, education, science and technology, culture, healthcare, finance, agriculture, tourism, and consumer. The pilot can be seen as an operational link between the new TC260 AI Application Ethics &amp; Safety Guidelines 1.0 instruction that developers &#8220;meet our country&#8217;s AI science-and-technology ethics review requirements&#8221; (see analysis <a href="https://chinaaibulletin.substack.com/i/198737671/tc260-releases-guidelines-for-ai-applications">below</a>) and what those requirements look like in practice.</p><h2>Both legislative work plans include an AI Law, but not as top priority</h2><p><strong>On May 11, the State Council and the NPC Standing Committee both published their respective 2026 annual legislative work plans</strong>. The <a href="https://www.moj.gov.cn/pub/sfbgw/gwxw/xwyw/202605/t20260511_534682.html">State Council plan</a> calls for <strong>&#8220;accelerat[ing] comprehensive AI legislation&#8221;</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-12" href="#footnote-12" target="_self">12</a> and names six key legislative elements (data, compute, algorithms, IP, cybersecurity, and supply chain security); this signals that <strong>activity on AI laws and regulations is likely to keep accelerating across these domains</strong>, even as the plan does not restore the comprehensive AI Law itself to the preparatory-item status it held in 2023 and 2024&#8212;the 2025 plan dropped it entirely, <a href="https://www.geopolitechs.org/p/chinas-ai-law-recent-developments">replacing it</a> with generic language about &#8220;advancing legislation for the sound development of AI.&#8221; The <strong><a href="https://npcobserver.com/2026/05/11/china-npc-2026-legislative-plan/">NPC Standing Committee plan</a> lists AI only as a backup research topic</strong>, the lowest of three legislative priority tiers, grouping &#8220;the healthy development of artificial intelligence&#8221; with fiscal policy, agricultural support, and online violence governance as subjects warranting study rather than active drafting. The 2025 NPC SC plan had <a href="https://www.geopolitechs.org/p/chinas-ai-law-recent-developments">directed</a> that AI legislation &#8220;shall be researched and drafted promptly by relevant departments, and deliberation shall be arranged as appropriate&#8221;&#8212;a research-and-drafting status the 2026 plan downshifts to research only. <strong>The combined picture is two-track: regulatory activity across the six named elements is likely to keep accelerating, but the timeline for a singular comprehensive AI Law is likely longer</strong>.</p><h2>National Standards</h2><h3>TC260 releases ethics &amp; safety guidelines for AI applications</h3><p>On May 19, <strong><a href="https://www.secrss.com/articles/90485">TC260 released TC260-005 &#12298;&#20154;&#24037;&#26234;&#33021;&#24212;&#29992;&#20262;&#29702;&#23433;&#20840;&#25351;&#24341; 1.0&#12299;</a> (AI Application Ethics and Safety Guidelines 1.0).</strong> The document is a TC260 technical reference document, characterized as &#8220;principle-based and reference-only,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-13" href="#footnote-13" target="_self">13</a> and is intended to coordinate with existing rules on personal information, automated decision-making, content labeling, algorithmic governance, and intellectual property (IP) rather than to add binding obligations.</p><p><strong>The document is substantially concerned with loss of control, but at the application level rather than the model level.</strong> The first of six ethics-safety impact categories is &#8220;human dominance impact,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-14" href="#footnote-14" target="_self">14</a> defined as AI behavior exceeding &#8220;preset, understood, and controllable range&#8221; (Section 4). Ultimate authority over AI must belong to humans, and AI must always remain under human control to &#8220;<strong>prevent AI from escaping human oversight or threatening human survival and development</strong>&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-15" href="#footnote-15" target="_self">15</a> (Section 5.2(g)). Developers building &#8220;highly autonomous AI applications&#8221; must &#8220;focus on assessing loss-of-control risk and its impact on industry and society&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-16" href="#footnote-16" target="_self">16</a> (Section 6.2(b)). This vocabulary was introduced into Chinese standards-track text by TC260&#8217;s <a href="https://www.cac.gov.cn/2025-09/15/c_1759653448369123.htm">AI Safety Governance Framework 2.0</a> (September 2025), which named &#8220;trustworthy application, preventing loss of control&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-17" href="#footnote-17" target="_self">17</a> as a foundational governance principle. TC260-005 carries the same framing forward into a developer obligation for highly autonomous AI applications. <strong>It also recommends an &#8220;off switch&#8221; </strong>that users can use to shut down a service (Section 6.3(d)).</p><p><strong>The document also comments on open-source development and security</strong>, encouraging the open-sourcing of AI models, tool components, and evaluation benchmarks alongside open-source ecosystem-security capacity (Section 5.2(i)), complementing rather than contradicting Qinglang&#8217;s open-source accountability provisions above. The document functions as guidance, not binding rule, and the 1.0 versioning signals further iterations planned.</p><h3>TC28/SC42 publishes foundational AI trustworthiness standard</h3><p>On April 30, the <strong><a href="https://std.sacinfo.org.cn/gnoc/queryInfo?id=3B7678AF7B8592CB94F331AED4BB4027">Standardization Administration of China (SAC) Announcement No. 21/2026</a> approved <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=511EBC5967DA9318E06397BE0A0AFBD5">GB/T 47507-2026 &#12298;&#20154;&#24037;&#26234;&#33021; &#21487;&#20449;&#36182; &#36890;&#21017;&#12299;</a> (Artificial Intelligence &#8212; Trustworthiness &#8212; General Rules),</strong> set to be implemented August 1, 2026. The standard was drafted under TC28/SC42, the AI subcommittee of the National Information Technology Standardization Technical Committee, by 40 organizations and 87 named individuals over a 21-month cycle that began in March 2024. The roster centers on the China Electronics Standardization Institute (CESI), the Chinese Academy of Sciences (CAS) Institute of Software, and the Nanjing Software Technology Research Institute, joined by SenseTime, Ant Group, Baidu, Hikvision, Tencent Cloud, CloudWalk, Inspur, and the Shanghai AI Innovation Center, with universities including Beihang, Xi&#8217;an Jiaotong, and Shandong. None of the TC260-005 drafters (Tsinghua, I-AIIG, Alibaba, Huawei, and DeepSeek) appear on this roster.</p><p>GB/T 47507 anchors a growing TC28/SC42 trustworthiness family. The published record cross-references six related plans, four already published. Companion standards under TC28/SC42 include <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=37FC557E3A727C66E06397BE0A0A8AB8">Trustworthy Datasets</a>, <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=41A934A216733FDBE06397BE0A0AC9DA">General Requirements of Trustworthiness for Embodied Intelligence</a>, and <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=508271370C98A41CE06397BE0A0A979A">Generative AI System Risk Response Guide</a>.</p><h3>WG9 holds second 2026 plenary; agent-security standard enters deliberation</h3><p>On May 12&#8211;13, the <strong>AI Safety Standards Working Group (WG9) under TC260 held its <a href="https://www.tc260.org.cn/portal/article/1/dd991f9781d64096afd06d232083cf14">second 2026 plenary in Shanghai</a></strong>, drawing more than 350 representatives from over 200 organizations. <strong>Zhou Bowen (&#21608;&#20271;&#25991;), director of Shanghai AI Lab and head of the committee, chaired the meeting</strong>. The plenary reviewed 16 national cybersecurity standard projects spanning edge-side LLMs, deepfake/synthesis security, embodied AI, testing-agency capabilities, foundation-model safety, model development and open-source safety, training and inference frameworks, system interoperability, and seven sectoral guidance documents (broadcasting, education, finance, healthcare, emergency management, and government affairs).</p><p>Two items were deliberated: <strong>the AI Safety Standards System</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-18" href="#footnote-18" target="_self">18</a><strong> and the Intelligent Agent Application Security Basic Requirements</strong>,<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-19" href="#footnote-19" target="_self">19</a> the latter drafted as a mandatory national standard (GB) rather than a recommended one (GB/T). Mandatory-tier AI standards are relatively unusual in China; the agent-security standard&#8217;s development overlaps the same window as the May 8 Implementation Opinions on Intelligent Agents (CAC/NDRC/MIIT), <strong>reflecting attention to agent security from China&#8217;s standards track and policy-direction track in the same cycle</strong>. The standard also runs alongside TC28/SC42&#8217;s plan 20262612-Z-469 (AI &#8212; Agent General Requirements), the technical guidance document in the TC28/SC42 reference table below. China&#8217;s two principal AI standards bodies are now drafting parallel agent-security standards at different binding tiers.</p><h3 style="text-align: justify;">AI terminal intelligence grading series released</h3><p>On April 30, <strong><a href="https://std.sacinfo.org.cn/gnoc/queryInfo?id=31934533D5445F5FCED1BF79F6AE776C">SAC Announcement No. 19/2026</a> approved the GB/Z 177-2026 series &#12298;&#20154;&#24037;&#26234;&#33021;&#32456;&#31471;&#26234;&#33021;&#21270;&#20998;&#32423;&#12299; (Grading the Intelligence Levels of AI Terminals</strong>), publicly <a href="https://www.miit.gov.cn/xwfb/gxdt/sjdt/art/2026/art_9fd39c053e484ba4bec7863849213092.html">announced by MIIT on May 8</a> as a joint release with the Ministry of Commerce (MOFCOM) and the State Administration for Market Regulation (SAMR). GB/Z 177.1 (Reference Framework) and GB/Z 177.2 (General Requirements) define what counts as &#8220;intelligence&#8221; in an AI terminal, the grading scale, and testing methods; GB/Z 177.3, 177.4, 177.7, 177.8, and 177.9 cover specific product categories (Mobile Terminals, Microcomputers, Automotive Cockpit, Speakers, and Earphones). Based on the MIIT announcement, 177.5 and 177.6 are likely Televisions and Smart Glasses, but only the five appear in the SAC release.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!8Io6!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!8Io6!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 424w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 848w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 1272w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!8Io6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png" width="1434" height="564" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:564,&quot;width&quot;:1434,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of 16 recently-published Chinese national standards from SAC Announcements 19 and 21 of 2026. Columns: standard number, name (translated), published date, implementation date, status. The set includes GB/T 47602 (Distributed Computing &#8212; Computing Devices Technical Requirements), GB/T 42382-3 (Neural Network Representation and Model Compression, Part 3: Graph Neural Networks), **GB/T 47507 (AI &#8212; Trustworthiness &#8212; General Rules, published 2026-04-30, implementation 2026-08-01)**, plus the GB/Z 177 series on AI Terminal Intelligence Grading (Parts 1, 2, 3, 4, 7, 8, 9). All standards were published 2026-04-30; most are pending implementation in November 2026, while the AI Terminal Intelligence Grading parts are already in force.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of 16 recently-published Chinese national standards from SAC Announcements 19 and 21 of 2026. Columns: standard number, name (translated), published date, implementation date, status. The set includes GB/T 47602 (Distributed Computing &#8212; Computing Devices Technical Requirements), GB/T 42382-3 (Neural Network Representation and Model Compression, Part 3: Graph Neural Networks), **GB/T 47507 (AI &#8212; Trustworthiness &#8212; General Rules, published 2026-04-30, implementation 2026-08-01)**, plus the GB/Z 177 series on AI Terminal Intelligence Grading (Parts 1, 2, 3, 4, 7, 8, 9). All standards were published 2026-04-30; most are pending implementation in November 2026, while the AI Terminal Intelligence Grading parts are already in force." title="Table of 16 recently-published Chinese national standards from SAC Announcements 19 and 21 of 2026. Columns: standard number, name (translated), published date, implementation date, status. The set includes GB/T 47602 (Distributed Computing &#8212; Computing Devices Technical Requirements), GB/T 42382-3 (Neural Network Representation and Model Compression, Part 3: Graph Neural Networks), **GB/T 47507 (AI &#8212; Trustworthiness &#8212; General Rules, published 2026-04-30, implementation 2026-08-01)**, plus the GB/Z 177 series on AI Terminal Intelligence Grading (Parts 1, 2, 3, 4, 7, 8, 9). All standards were published 2026-04-30; most are pending implementation in November 2026, while the AI Terminal Intelligence Grading parts are already in force." srcset="https://substackcdn.com/image/fetch/$s_!8Io6!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 424w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 848w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 1272w, https://substackcdn.com/image/fetch/$s_!8Io6!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F76e0f337-0a4f-4aaf-b48c-c129216e03e1_1434x564.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Recently published national standards on the TC28/SC42 <a href="https://std.samr.gov.cn/search/orgDetailView?data_id=A132FB8FCABB4BE9E05397BE0A0AB880">website</a></em></figcaption></figure></div><p>TC28/SC42 also registered a batch of new technical guidance documents, laid out in the table below.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!N4E1!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!N4E1!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 424w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 848w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 1272w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!N4E1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png" width="1193" height="590" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:590,&quot;width&quot;:1193,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table of 17 newly-registered TC28/SC42 technical guidance document plans (Z-469 series), all from early- to mid-May 2026. Columns: plan number, project name (translated), registered date. Six embodied-intelligence plans were registered on 2026-05-14 (simulation platform, operating system, vision large-model logistics parks, data-collection/training/inference, general performance assessment, scientific computing weather-model). Three were registered on 2026-05-12 (network-attached storage application guide, **AI Agent General Requirements (20262612-Z-469)**, design guide for intelligent companion systems for the elderly). Seven were registered on 2026-05-07 (Agent Model Governance Framework, intelligent computing cluster heterogeneous accelerator mixed inference, high-speed interconnect bus protocol, cloud sandbox technical guide for agent execution, shale oil/gas reservoir-space characterization, road traffic-scene prediction algorithm testing, **Agent Evaluation Indicators and Methods**). One was registered on 2026-05-06 (computing high-speed interconnect bus unified protocol &#8212; general principles).&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of 17 newly-registered TC28/SC42 technical guidance document plans (Z-469 series), all from early- to mid-May 2026. Columns: plan number, project name (translated), registered date. Six embodied-intelligence plans were registered on 2026-05-14 (simulation platform, operating system, vision large-model logistics parks, data-collection/training/inference, general performance assessment, scientific computing weather-model). Three were registered on 2026-05-12 (network-attached storage application guide, **AI Agent General Requirements (20262612-Z-469)**, design guide for intelligent companion systems for the elderly). Seven were registered on 2026-05-07 (Agent Model Governance Framework, intelligent computing cluster heterogeneous accelerator mixed inference, high-speed interconnect bus protocol, cloud sandbox technical guide for agent execution, shale oil/gas reservoir-space characterization, road traffic-scene prediction algorithm testing, **Agent Evaluation Indicators and Methods**). One was registered on 2026-05-06 (computing high-speed interconnect bus unified protocol &#8212; general principles)." title="Table of 17 newly-registered TC28/SC42 technical guidance document plans (Z-469 series), all from early- to mid-May 2026. Columns: plan number, project name (translated), registered date. Six embodied-intelligence plans were registered on 2026-05-14 (simulation platform, operating system, vision large-model logistics parks, data-collection/training/inference, general performance assessment, scientific computing weather-model). Three were registered on 2026-05-12 (network-attached storage application guide, **AI Agent General Requirements (20262612-Z-469)**, design guide for intelligent companion systems for the elderly). Seven were registered on 2026-05-07 (Agent Model Governance Framework, intelligent computing cluster heterogeneous accelerator mixed inference, high-speed interconnect bus protocol, cloud sandbox technical guide for agent execution, shale oil/gas reservoir-space characterization, road traffic-scene prediction algorithm testing, **Agent Evaluation Indicators and Methods**). One was registered on 2026-05-06 (computing high-speed interconnect bus unified protocol &#8212; general principles)." srcset="https://substackcdn.com/image/fetch/$s_!N4E1!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 424w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 848w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 1272w, https://substackcdn.com/image/fetch/$s_!N4E1!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc164eb5d-783d-428a-80e3-6c75eab79592_1193x590.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Recently announced technical guidance documents on the TC28/SC42 <a href="https://std.samr.gov.cn/search/orgDetailView?data_id=A132FB8FCABB4BE9E05397BE0A0AB880">website</a></em></figcaption></figure></div><h3 style="text-align: justify;">CAC stands up AI-cybersecurity testing programme supported by Huawei</h3><p>On May 18, CAC announced the <strong><a href="https://www.cac.gov.cn/2026-05/18/c_1780846421726446.htm">2026 Test on the Application of Artificial Intelligence Technology in Cybersecurit</a></strong><a href="https://www.cac.gov.cn/2026-05/18/c_1780846421726446.htm">y</a>, a competition that runs from June till August, with results showcased at the National Cybersecurity Promotion Week in September. CAC&#8217;s Network Security Coordination Bureau and 15 other ministries and central bodies act as guiding units. <strong>The National Computer Network Emergency Response Technical Team (CNCERT) hosts, and Huawei is named as the sole Technical Support Unit.</strong></p><p>Though unclear if concerns about Anthropic&#8217;s <a href="https://red.anthropic.com/2026/mythos-preview/">Mythos</a> played a role, the eight designated scenarios <strong>mix conventional cybersecurity applications</strong> (network defense, vulnerability discovery, and network traffic threat detection) with <strong>items oriented toward frontier AI safety</strong>, like large model safety guardrail testing and AI agent malicious operation behavior detection.</p><h3 style="text-align: justify;">Other news: AI+ Energy &amp; leader inspections</h3><p>On the same day as the Intelligent Agents Implementation Opinions were released, NDRC, the National Energy Administration, MIIT, and the National Data Bureau jointly issued the <a href="https://www.nea.gov.cn/20260508/4dae97ca01d348e4871bb8654be34b3a/c.html">Action Plan on Promoting Mutual Empowerment between AI and Energy</a>, outlining 29 tasks across two pillars: energy supporting AI compute and AI supporting energy applications.</p><p>Three high-level AI inspections occurred in nine days. NDRC Chair Zheng Shanjie <a href="https://www.ndrc.gov.cn/fzggw/wld/zsj/zyhd/202605/t20260509_1405124.html">visited Shanghai AI Lab</a> on May 9, and on May 18 Premier Li Qiang <a href="https://www.gov.cn/yaowen/liebiao/202605/content_7069329.htm">inspected AI-manufacturing integration in Beijing</a> and Vice Premier Ding Xuexiang <a href="https://www.gov.cn/yaowen/liebiao/202605/content_7069298.htm">inspected national integrated compute network construction</a>.</p><h1>International AI Governance</h1><h2>US and China agree to launch government-to-government AI dialogue at Trump-Xi summit</h2><p>Following Trump&#8217;s early-May visit to Beijing, <strong>the two governments agreed to a bilateral AI dialogue</strong>&#8212;the second formal intergovernmental dialogue on AI between the two countries, after a <a href="https://www.geopolitechs.org/p/chinas-readout-of-the-first-sino">prior round</a> in May 2024. <strong>Treasury Secretary Scott Bessent first informally alluded to the agreement on <a href="https://www.nytimes.com/2026/05/14/world/asia/china-us-ai-safety.html">May 14</a></strong>; <strong>Chinese Foreign Ministry spokesperson Guo Jiakun confirmed it at the <a href="https://mp.weixin.qq.com/s/D6BsE3mFxJNGiRsqj0toRg">May 19 regular press briefing</a></strong>. The Chinese readout states that &#8220;during President Trump&#8217;s visit to China, the two heads of state held constructive exchanges on AI issues and agreed to launch a government-to-government AI dialogue,&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-20" href="#footnote-20" target="_self">20</a> but the US has not officially confirmed.</p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this useful? Subscribe to get the China AI Bulletin in your inbox regularly.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1>Frontier Lab Developments</h1><h2>Notable Model Releases</h2><p><strong>Alibaba</strong> launched <strong><a href="https://qwen.ai/blog?id=qwen3.7">Qwen 3.7-Max</a></strong> as &#8220;a proprietary model designed for the agent era.&#8221; The model targets agentic coding, complex reasoning, and extended multi-step tasks, with claimed autonomous operation up to 35 hours and over 1,000 tool calls without performance degradation. Two previews had surfaced on LMArena five days earlier: <strong><a href="https://decrypt.co/368499/alibaba-qwen-3-7-max-preview-review">Qwen3.7-Max-Preview</a> </strong>ranked 13th globally on the text leaderboard, and Qwen3.7-Plus-Preview ranked 16th globally on the vision leaderboard. The launch came paired with <strong>Alibaba&#8217;s new <a href="https://www.business-standard.com/technology/tech-news/china-alibaba-new-ai-chip-llm-domestic-alternatives-126052000570_1.html">Zhenwu M890 custom AI chip</a></strong>. Alibaba Cloud senior vice-president Liu Weiguang framed the company&#8217;s positioning as <a href="https://www.scmp.com/tech/big-tech/article/3354212/alibaba-unveils-new-qwen-model-custom-chips-bid-become-chinas-ai-factory">&#8220;China&#8217;s AI factory,&#8221;</a> claiming Alibaba is the only Chinese company operating all five layers of the full AI stack: chips, agentic cloud, AI models, model service platforms, and agentic applications.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!hKyp!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!hKyp!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 424w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 848w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!hKyp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png" width="1456" height="808" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:808,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar-chart grid comparing Qwen3.7-Max against Qwen3.6-Plus, DeepSeek-V4-Pro Max, GLM-5.1, Kimi K2.6, and Claude Opus-4.6 Max across twelve benchmarks. Top row, agentic coding and engineering: Terminal-Bench 2.0 (Qwen3.7-Max 69.7), SWE-bench Pro (60.0), SWE-bench Multilingual (78.3), NL2Repo (47.2). Middle row, realistic agent and MCP use: MCP-Atlas (76.4), MCP-Mark (60.8), ClawEval (65.2), CoWorkBench (67.2). Bottom row, reasoning and knowledge: HLE (41.4), Apex math reasoning (44.5), IFBench instruction following (79.4), SuperGPQA graduate-level knowledge (73.6). Qwen3.7-Max leads or matches the field across the agentic and instruction-following benchmarks; trails Claude Opus on some reasoning ones.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar-chart grid comparing Qwen3.7-Max against Qwen3.6-Plus, DeepSeek-V4-Pro Max, GLM-5.1, Kimi K2.6, and Claude Opus-4.6 Max across twelve benchmarks. Top row, agentic coding and engineering: Terminal-Bench 2.0 (Qwen3.7-Max 69.7), SWE-bench Pro (60.0), SWE-bench Multilingual (78.3), NL2Repo (47.2). Middle row, realistic agent and MCP use: MCP-Atlas (76.4), MCP-Mark (60.8), ClawEval (65.2), CoWorkBench (67.2). Bottom row, reasoning and knowledge: HLE (41.4), Apex math reasoning (44.5), IFBench instruction following (79.4), SuperGPQA graduate-level knowledge (73.6). Qwen3.7-Max leads or matches the field across the agentic and instruction-following benchmarks; trails Claude Opus on some reasoning ones." title="Bar-chart grid comparing Qwen3.7-Max against Qwen3.6-Plus, DeepSeek-V4-Pro Max, GLM-5.1, Kimi K2.6, and Claude Opus-4.6 Max across twelve benchmarks. Top row, agentic coding and engineering: Terminal-Bench 2.0 (Qwen3.7-Max 69.7), SWE-bench Pro (60.0), SWE-bench Multilingual (78.3), NL2Repo (47.2). Middle row, realistic agent and MCP use: MCP-Atlas (76.4), MCP-Mark (60.8), ClawEval (65.2), CoWorkBench (67.2). Bottom row, reasoning and knowledge: HLE (41.4), Apex math reasoning (44.5), IFBench instruction following (79.4), SuperGPQA graduate-level knowledge (73.6). Qwen3.7-Max leads or matches the field across the agentic and instruction-following benchmarks; trails Claude Opus on some reasoning ones." srcset="https://substackcdn.com/image/fetch/$s_!hKyp!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 424w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 848w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 1272w, https://substackcdn.com/image/fetch/$s_!hKyp!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F83b6584a-8486-42ab-aada-7c68141a1d3a_1888x1048.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: Alibaba <a href="https://qwen.ai/blog?id=qwen3.7">blog</a></figcaption></figure></div><p><strong>Baidu</strong> released <strong><a href="https://ernie.baidu.com/blog/posts/ernie-5.1-preview-0430-release-on-lmarena/">ERNIE 5.1 Preview</a> </strong>on April 30 and the <strong><a href="https://ernie.baidu.com/blog/posts/ernie-5.1-0508-release/">official ERNIE 5.1</a></strong> as a closed model on May 9. The model is text-only and built on Mixture-of-Experts (MoE), focused on efficiency with roughly one-third of ERNIE 5.0&#8217;s total parameters and half its active parameters. Baidu claims a <strong>pre-training cost of about 6% of comparable models</strong>. ERNIE 5.1 ranks<strong> fourth globally on LMArena Search</strong> and <strong>first among Chinese models with a score of 1,223</strong>. Reported benchmark numbers include 99.6 on AIME26 with tools, &#964;&#179;-bench and SpreadsheetBench-Verified scores surpassing DeepSeek-V4-Pro, and creative-writing performance approaching Gemini 3.1 Pro. The efficiency-first framing runs in contrast to the trillion-plus flagship models covered in Issue 3.</p><p><strong>ByteDance</strong> released <strong><a href="https://github.com/bytedance/Lance">Lance</a></strong> on May 15 (<a href="https://arxiv.org/abs/2605.18678">technical paper</a>), a <strong>3B-parameter native unified multimodal model</strong> spanning six tasks across two modalities: text-to-image, image understanding (visual question answering, reasoning), image editing, text-to-video, video understanding, and video editing. Total training budget was 128 A100 GPUs, more efficient than what comparable unified multimodal models typically require. Lance <strong>claims competitive scores </strong>on a variety of benchmarks against far larger models. The efficiency-at-3B framing positions Lance as a contrast to the scale-driven flagships covered in Issue 3.</p><p><strong>Tencent </strong>compressed <a href="https://huggingface.co/tencent/Hy-MT1.5-1.8B-2bit">HY-MT1.5-1.8B into 1.25-bit and 2-bit quantized variants</a> (April 29) using Stretched Elastic Quantization (SEQ); the compressed versions target mobile devices. <strong>Tencent</strong> also released <strong><a href="https://github.com/Tencent-Hunyuan/Hy-Embodied-RoboFusion">Hy-Embodied-RoboFusion</a></strong> (May 6) for embodied AI research and <strong><a href="https://github.com/Tencent-Hunyuan/R-DMesh">R-DMesh</a></strong> (May 13), a video-guided 4D mesh animation framework from the same team.</p><h2>Technical Publication Highlights</h2><p>It was a big few weeks of publishing: frontier labs released 206 papers on arXiv this edition. Alibaba authors led the pack with 70, and SenseTime put itself on the board. Highlights are below; a full list with summaries can be found <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0">here</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!7c4X!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!7c4X!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 424w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 848w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 1272w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!7c4X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png" width="1278" height="1150" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1150,&quot;width&quot;:1278,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:126935,&quot;alt&quot;:&quot;Table of arXiv paper counts for sixteen Chinese AI labs, this edition vs. cumulative since the bulletin launched. Columns: Lab, Papers This Edition, Total Papers Since Bulletin Launch. **Alibaba leads at 70 this edition (122 cumulative), followed by Tencent (41 / 66), ByteDance (32 / 46), Huawei (23 / 54), Baidu (14 / 27), Meituan (14 / 20).** Smaller contributions: Xiaomi 6 / 15, SenseTime 3 / 3 (first edition on the board), iFlyTek 2 / 5, StepFun 1 / 3. **Zero this edition:** Baichuan, DeepSeek, MiniMax, Moonshot, Zhipu (3 cumulative), 01.AI.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/198737671?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of arXiv paper counts for sixteen Chinese AI labs, this edition vs. cumulative since the bulletin launched. Columns: Lab, Papers This Edition, Total Papers Since Bulletin Launch. **Alibaba leads at 70 this edition (122 cumulative), followed by Tencent (41 / 66), ByteDance (32 / 46), Huawei (23 / 54), Baidu (14 / 27), Meituan (14 / 20).** Smaller contributions: Xiaomi 6 / 15, SenseTime 3 / 3 (first edition on the board), iFlyTek 2 / 5, StepFun 1 / 3. **Zero this edition:** Baichuan, DeepSeek, MiniMax, Moonshot, Zhipu (3 cumulative), 01.AI." title="Table of arXiv paper counts for sixteen Chinese AI labs, this edition vs. cumulative since the bulletin launched. Columns: Lab, Papers This Edition, Total Papers Since Bulletin Launch. **Alibaba leads at 70 this edition (122 cumulative), followed by Tencent (41 / 66), ByteDance (32 / 46), Huawei (23 / 54), Baidu (14 / 27), Meituan (14 / 20).** Smaller contributions: Xiaomi 6 / 15, SenseTime 3 / 3 (first edition on the board), iFlyTek 2 / 5, StepFun 1 / 3. **Zero this edition:** Baichuan, DeepSeek, MiniMax, Moonshot, Zhipu (3 cumulative), 01.AI." srcset="https://substackcdn.com/image/fetch/$s_!7c4X!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 424w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 848w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 1272w, https://substackcdn.com/image/fetch/$s_!7c4X!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd866eced-a6bf-448f-b39b-f05425d5ed7a_1278x1150.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3>Alibaba</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.14091v1">Venus-DeFakerOne: Unified Fake Image Detection &amp; Localization</a></strong></p><ul><li><p><strong>DeFakerOne</strong> unifies fake image detection and localization across deepfakes, AI-generated content, and document forgeries by combining vision-language and segmentation models. It <strong>outperforms baselines on 39 detection and 9 localization benchmarks</strong> while maintaining robustness against real-world perturbations and state-of-the-art generators like GPT-Image-2.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.09497v1">Don&#8217;t Click That: Teaching Web Agents to Resist Deceptive Interfaces</a></strong></p><ul><li><p>LLMs struggle with deceptive web interfaces. This paper introduces <strong>DUDE</strong>, a framework using hybrid-reward learning and experience summarization to teach web agents to <strong>resist clickbait and scam elements</strong>, reducing susceptibility by 53.8% while maintaining task completion rates.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.11887v1">Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models</a></strong></p><ul><li><p>Tackles LLM opacity by introducing <strong>Qwen-Scope</strong>, an open-source suite of sparse autoencoders (tools that decompose model thinking into interpretable features) across 14 SAE variants on Qwen models, enabling <strong>inference-time steering, evaluation analysis, multilingual safety classification, and fine-tuning optimization</strong> without modifying weights. Results show SAEs function as practical development interfaces beyond post-hoc analysis&#8212;controlling behavior, detecting benchmark redundancy, and mitigating code-switching and repetition in training.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.08738v1">SlimQwen: Exploring the Pruning and Distillation in Large MoE Model Pre-training</a></strong></p><ul><li><p>By comparing pruning pretrained models against training from scratch and progressive against one-shot pruning schedules, <strong>SlimQwen studies MoE compression at pretraining scale</strong>. The work finds <strong>pruning beats from-scratch training</strong>, expert merging converges after continued training, and progressive schedules beat one-shot compression, <strong>compressing Qwen3-Next-80A3B to 23A2B while retaining performance</strong>.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2605.08766v1">UserGPT Technical Report</a></strong></p><ul><li><p>Tackles personalized user understanding by proposing <strong>UserGPT</strong>, which uses LLMs to generate coherent user personas from noisy behavioral histories through a pipeline combining a behavior simulation engine, semantic data transformation, and curriculum-driven training with reinforcement learning, achieving <strong>73.25% accuracy on tag prediction and 75.28% on summary generation while compressing behavioral records by 98%</strong>.</p></li></ul><h3>Baidu</h3><p><strong><a href="http://arxiv.org/abs/2605.17637v1">WebGameBench: Requirement-to-Application Evaluation for Coding Agents via Browser-Native Games</a></strong></p><ul><li><p>Coding agent evaluations grade code artifacts rather than delivered applications. <strong>WebGameBench</strong> instead <strong>requires agents to build browser-playable games</strong> from specifications and<strong> grades them via runtime interaction in actual browsers</strong>, revealing that while <strong>top agents achieve 76.9% playable delivery, only 20.2% fully satisfy requirements</strong>.</p></li></ul><h3>ByteDance</h3><p><strong><a href="http://arxiv.org/abs/2605.03360v1">A-CODE: Fully Atomic Protein Co-Design with Unified Multimodal Diffusion</a></strong></p><ul><li><p><strong>A-CODE designs proteins at the atomic level </strong>rather than at the level of protein building blocks, using a unified diffusion model that simultaneously predicts atom types and positions in one stage. It <strong>outperforms two-stage methods</strong> on unconditional generation, matches state-of-the-art binder design, and <strong>achieves 10&#215; higher success on hard tasks</strong> while enabling non-canonical amino acid modeling for the first time.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.05460v1">Agentic Discovery of Exchange-Correlation Density Functionals</a></strong></p><ul><li><p><strong>Designs exchange-correlation functionals</strong> in density functional theory through an agentic LLM system that iteratively proposes and evaluates functional improvements, discovering <strong>SAFS26-a, which outperforms the top baseline by ~9%</strong>. The work reveals the need for domain-expert constraints to <strong>prevent AI from exploiting unphysical shortcuts</strong>.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.08962v1">MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production</a></strong></p><ul><li><p><strong>Improves multimodal LLM training</strong> under dynamic workloads via <strong>MegaScale-Omni</strong>, a system featuring decoupled parallelism strategies, unified encoder-LLM representations, and workload balancing. It achieves <strong>1.27&#215;&#8211;7.57&#215; throughput gains</strong> at thousand-GPU scale.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.13831v1">Training Long-Context Vision-Language Models Effectively with Generalization Beyond 128K Context</a></strong></p><ul><li><p><strong>Extends Qwen2.5-VL-7B from 32K to 128K context</strong> with only 5B tokens via <strong>MMProLong</strong>, which <strong>generalizes to 256K&#8211;512K tokens and diverse downstream tasks without retraining</strong>. A systematic study of data mixtures shows balanced sequence-length distributions and retrieval-heavy compositions outperform target-length-focused data for long-context vision-language model training.</p></li></ul><h3>Huawei</h3><p><strong><a href="http://arxiv.org/abs/2605.18703v1">EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL</a></strong></p><ul><li><p><strong>EnvFactory</strong> autonomously synthesizes executable environments and natural multi-turn trajectories, <strong>addressing the shortage of realistic training environments and data for tool-use agents</strong>. Using 85 verified environments, it generates 2,575 training trajectories and <strong>improves Qwen models by up to +15% on BFCLv3 and +8.6% on MCP-Atlas</strong>.</p></li></ul><h3>iFlyTek</h3><p><strong><a href="http://arxiv.org/abs/2605.01480v1">AttnRouter: Per-Category Attention Routing for Training-Free Image Editing on MMDiT</a></strong></p><ul><li><p><strong>AttnRouter</strong> is a <strong>per-category routing table</strong> that selects the best editing operation for each image modification type, paired with <strong>KVInject</strong>, a simplified attention manipulation that injects source image features into noise tokens for training-free image editing on MMDiT. Ground-truth routing improves quality by 6.4%, with a zero-shot classifier recovering 98% of gains despite modest accuracy.</p></li></ul><h3>SenseTime</h3><p><strong><a href="http://arxiv.org/abs/2605.12500v1">SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture</a></strong></p><ul><li><p><strong>SenseNova-U1</strong> is a <strong>unified architecture for vision-language models</strong> where<strong> both understanding and generation emerge from a single underlying process,</strong> addressing the fragmentation of these capabilities. Models match top-tier understanding-only VLMs on standard benchmarks while supporting image generation, text-rich synthesis, and vision-language-action tasks.</p></li></ul><h3>StepFun</h3><p><strong><a href="http://arxiv.org/abs/2605.12034v1">Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation</a></strong></p><ul><li><p>Current omni-modal benchmarks overstate progress by allowing models to answer queries using only visual information; this paper introduces <strong>OmniClean</strong>, a visually debiased benchmark with 8,551 queries, and <strong>OmniBoost</strong>, a three-stage post-training method that enables a 3B model to match a 30B model&#8217;s performance through mixed bi-modal training, reinforcement learning, and self-distilled data.</p></li></ul><h3>Tencent</h3><p><strong><a href="http://arxiv.org/abs/2604.27043v1">CL-bench Life: Can Language Models Learn from Real-Life Context?</a></strong></p><ul><li><p>Addresses the <strong>gap between lab benchmarks and real-world AI use via CL-bench Life</strong>, a human-curated benchmark of <strong>405 messy, real-life contexts</strong> (group chats, personal archives, behavioral traces) with <strong>5,348 verification rubrics</strong>. Even frontier models achieve only <strong>19.3% task-solving rates</strong>, revealing that real-life context learning remains fundamentally unsolved.</p></li></ul><p><strong><a href="http://arxiv.org/abs/2605.05185v1">OpenSearch-VL: An Open Recipe for Frontier Multimodal Search Agents</a></strong></p><ul><li><p>Releases <strong>OpenSearch-VL</strong>, a <strong>fully open recipe for multimodal deep search</strong> with curated training data, diverse tool environments (text/image search, OCR, image enhancement), and a fatal-aware RL algorithm that gracefully handles tool failures. The system achieves <strong>10+ point improvements across seven benchmarks</strong> and matches proprietary models on several tasks.</p></li></ul><h3>Xiaomi</h3><p><strong><a href="http://arxiv.org/abs/2605.18137v1">Xiaomi EV World Model: A Joint World Model Integrating Reconstruction and Generation for Autonomous Driving</a></strong></p><ul><li><p>By integrating <strong>WorldRec</strong> (a reconstruction module using 3D scene queries for multi-view consistency) and <strong>WorldGen</strong> (a video generation module trained via bidirectional pretraining and causal fine-tuning), <strong>JWM addresses autonomous driving simulation</strong>. The joint system achieves improved consistency and fidelity for closed-loop simulation and synthetic data generation.</p></li></ul><h1>Technical AI Safety Publication Highlights</h1><p>There were <strong>172 AI-safety-related papers published by Chinese researchers </strong>this edition. Highlights are below; a full list with summaries is available <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1">here</a>.</p><h2>Agentic Safety</h2><p><strong><a href="http://arxiv.org/abs/2605.06230v1">Safactory: A Scalable Agent Factory for Trustworthy Autonomous Intelligence</a></strong></p><ul><li><p><strong>Safactory</strong> is an integrated framework that connects simulation, data management, and model improvement into a closed-loop system for agent development. It couples a <strong>parallel simulation environment</strong> (for generating agent trajectories), a <strong>data platform</strong> (for storing and extracting behavioral patterns), and an <strong>autonomous evolution system </strong>(for reinforcement learning and model distillation). The authors position this as <strong>infrastructure for systematic risk discovery</strong> and <strong>continuous safety improvement</strong> as models transition from conversational assistants to autonomous agents operating in real environments.</p></li></ul><p><em>Institutional affiliations: Shanghai AI Laboratory</em></p><p><strong><a href="http://arxiv.org/abs/2605.17480v1">The Capability Paradox: How Smarter Auditors Make Multi-Agent Systems Less Secure</a></strong></p><ul><li><p><strong>Semantic hijacking</strong> attacks exploit multi-agent LLM systems by embedding harmful requests in domain-specific narratives that Worker agents report to a Manager. Testing across 12 Manager models reveals a <strong>capability paradox</strong>: stronger Workers increase attack success from 18.4% to 63.9%, because they express adversarial conclusions with greater linguistic certainty, causing Managers to comply. Mediation analysis confirms certainty drives 74% of this effect. <strong>Heterogeneous ensemble verification</strong>&#8212;pairing Workers with asymmetric expertise&#8212;breaks this chain, reducing attacks to 2.0%.</p></li></ul><p><em>Institutional affiliations: University of Chinese Academy of Sciences, Max Planck Institute for Security and Privacy, Henan Yinzhu Safety Technology Co., Harbin Institute of Technology</em></p><h2>Alignment</h2><p><strong><a href="http://arxiv.org/abs/2605.11679v2">Explaining and Breaking the Safety-Helpfulness Ceiling via Preference Dimensional Expansion</a></strong></p><ul><li><p><strong>MORA (Multi-Objective Reward Assimilation)</strong> addresses the trade-off between helpfulness and safety in LLM alignment by rewriting prompts to unlock diverse reward dimensions rather than forcing compromises along a fixed frontier. The method identifies that prompts themselves constrain achievable multi-dimensional rewards, then expands diversity through pre-sampling and question rewriting to incorporate multiple intents. Experiments show <strong>5&#8211;12.4% improvements</strong> in individual metrics (particularly harmlessness) after multi-objective alignment, with <strong>4.6% average gains</strong> in simultaneous optimization.</p></li></ul><p><em>Institutional affiliations: Huazhong University of Science and Technology, Nanyang Technological University, Tsinghua University, Chongqing University</em></p><p><strong><a href="http://arxiv.org/abs/2605.08930v1">Internalizing Safety Understanding in Large Reasoning Models via Verification</a></strong></p><ul><li><p><strong>Safety Internal (SInternal)</strong> reframes alignment by training reasoning models to <strong>evaluate their own outputs for safety</strong> rather than merely detect unsafe inputs. The framework uses expert reasoning trajectories to teach models to critique their generated answers. Models trained this way show <strong>stronger generalization against jailbreaks</strong> and provide better initialization for reinforcement learning alignment, suggesting that internalized verification produces more robust safety than supervised imitation alone.</p></li></ul><p><em>Institutional affiliations: University of Science and Technology of China, National University of Singapore, Shanghai Artificial Intelligence Laboratory</em></p><h2>Evaluation and Benchmarks</h2><p><strong><a href="http://arxiv.org/abs/2605.10267v1">IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs</a></strong></p><ul><li><p><strong>IndustryBench</strong> is a 2,049-item benchmark for industrial procurement QA grounded in Chinese national standards and product specifications. Unlike general LLM benchmarks, it explicitly separates <strong>raw correctness</strong> from <strong>safety violations</strong>&#8212;flagging when models introduce unsupported details that contradict safety clauses or regulatory thresholds. Across 17 models, the best achieves only 2.08 on a 0&#8211;3 scale; however, <strong>extended reasoning often worsens safety-adjusted scores by hallucinating safety-critical specifications</strong>. The benchmark demonstrates that leaderboard rankings collapse when safety compliance is properly weighted.</p></li></ul><p><em>Institutional affiliations: Alibaba Group</em></p><p><strong><a href="http://arxiv.org/abs/2605.07630v1">Safe, or Simply Incapable? Rethinking Safety Evaluation for Phone-Use Agents</a></strong></p><ul><li><p><strong>PhoneSafety</strong> is a benchmark of 700 safety-critical moments in real phone interactions that distinguishes between three outcomes: models taking safe actions, unsafe actions, or failing to act at all. Current evaluations conflate these&#8212;a model avoiding harm might reflect genuine safety judgment or mere incapability. Testing eight phone-use agents reveals that <strong>stronger general performance does not predict safer choices at risky moments</strong>, and that inability to act correlates with visual/operational difficulty rather than safety robustness.</p></li></ul><p><em>Institutional affiliations: Tencent Hunyuan; The Chinese University of Hong Kong, Shenzhen; Tsinghua University</em></p><p><strong><a href="http://arxiv.org/abs/2605.12015v1">SkillSafetyBench: Evaluating Agent Safety under Skill-Facing Attack Surfaces</a></strong></p><ul><li><p><strong>SkillSafetyBench</strong> evaluates a blind spot in agent safety: adversarial inputs embedded in task-relevant skill materials or local files can override benign user requests and trigger unsafe actions, even when the model itself is aligned. The benchmark includes <strong>155 test cases across 47 tasks and 6 risk domains</strong>, showing that agents consistently fail when malicious content is injected into skills rather than user prompts. Findings highlight that safety depends on agent architecture&#8212;how it interprets skills, trusts context, and executes actions&#8212;not just model alignment alone.</p></li></ul><p><em>Institutional affiliations: Shanghai AI Laboratory, Peking University, East China Normal University</em></p><h2>Governance and Policy</h2><p><strong><a href="http://arxiv.org/abs/2605.13069v2">Not All Anquan Is the Same: A Terminological Proposal for Chinese Computer Science and Engineering</a></strong></p><ul><li><p>This paper argues for <strong>disambiguating &#8220;anquan&#8221; in Chinese technical writing</strong> by adopting &#8220;anbao&#8221;&#65288;&#23433;&#20445;&#65289;for security and reserving &#8220;anquan&#8221;&#65288;&#23433;&#20840;&#65289;for safety. The single word currently conflates non-adversarial failures (safety) with intentional attacks (security), creating conceptual confusion in standards interpretation, risk analysis, and cross-disciplinary work. The author demonstrates how this conflation undermines precision in functional safety, automotive systems, cybersecurity, and AI governance, and proposes dual-track terminology practices to enable clearer scientific argumentation and assurance claims.</p></li></ul><p><em>Institutional affiliations: Wuhan University</em></p><h2>Guardrails and Deployment Safety</h2><p><strong><a href="http://arxiv.org/abs/2605.00689v1">ML-Bench&amp;Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models</a></strong></p><ul><li><p><strong>ML-Bench&amp;Guard</strong> constructs a multilingual safety benchmark directly from regional legal texts rather than translated taxonomies, covering 14 languages with jurisdiction-specific risk categories. <strong>ML-Guard</strong>, a diffusion-based guardrail model, provides <strong>policy-conditioned compliance assessment</strong>&#8212;outputting safe/unsafe verdicts (1.5B variant) or detailed explanations aligned to local regulations (7B variant). The approach addresses a concrete deployment challenge: existing multilingual safety systems rely on generic risk categories that don&#8217;t map to region-specific legal requirements, potentially creating compliance gaps in cross-border LLM deployment.</p></li></ul><p><em>Institutional affiliations: University of Illinois Urbana-Champaign, Fudan University, University of Chicago</em></p><h2>Interpretability</h2><p><strong><a href="http://arxiv.org/abs/2605.08942v1">Decomposing and Steering Functional Metacognition in Large Language Models</a></strong></p><ul><li><p><strong>Residual stream analysis</strong> reveals that LLMs encode decomposable <strong>functional metacognitive states</strong>&#8212;internal variables tracking evaluation awareness, self-assessed capability, perceived risk, and effort allocation&#8212;that are linearly decodable from model activations. By steering activations along probe-derived directions, researchers demonstrate each state <strong>causally modulates reasoning behavior</strong> in distinct ways, affecting verbosity, accuracy, and safety responses. This mechanism suggests benchmark performance conflates task competence with activation of specific internal states, raising questions about what standard evaluations actually measure.</p></li></ul><p><em>Institutional affiliations: Shopee</em></p><p><strong><a href="http://arxiv.org/abs/2605.08878v1">Why Do Aligned LLMs Remain Jailbreakable: Refusal-Escape Directions, Operator-Level Sources, and Safety-Utility Trade-off</a></strong></p><ul><li><p>This paper identifies <strong>Refusal-Escape Directions (RED)</strong>: continuous input perturbations that shift aligned models from refusing harmful requests to answering them while preserving semantic understanding of the harm. The authors decompose RED mathematically across model components, pinpointing <strong>normalization layers, residual connections, and output modules</strong> as structural sources of jailbreak vulnerability. They demonstrate a <strong>safety-utility trade-off</strong>: eliminating RED requires shared modules (attention, MLP) to suppress refusal-escape paths without breaking benign capabilities, revealing why aligned models remain jailbreakable despite training.</p></li></ul><p><em>Institutional affiliations: Chinese Academy of Sciences, University of Chinese Academy of Sciences</em></p><h2>Robustness and Adversarial Attacks</h2><p><strong><a href="http://arxiv.org/abs/2605.04446v1">Misrouter: Exploiting Routing Mechanisms for Input-Only Attacks on Mixture-of-Experts LLMs</a></strong></p><ul><li><p><strong>Misrouter</strong> exploits MoE routing mechanisms through input-only attacks&#8212;adversarial queries that manipulate which experts process tokens without direct model access. The method identifies weakly aligned experts and steers routing toward them while away from safety-trained ones, then optimizes prompts to trigger unsafe outputs while maintaining routing stability. Attacks transfer from open-source surrogate models to commercial API services, suggesting MoE&#8217;s routing layer presents an exploitable vulnerability in production systems.</p></li></ul><p><em>Institutional affiliations: Nankai University, Nanyang Technological University</em></p><p><strong><a href="http://arxiv.org/abs/2605.13411v1">Model-Agnostic Lifelong LLM Safety via Externalized Attack-Defense Co-Evolution</a></strong></p><ul><li><p><strong>EvoSafety</strong> decouples red-teaming and defense into externalized, reusable structures rather than embedding them in model weights. Attack discovery uses a <strong>skill library</strong> that expands beyond saturation, while defenses run as lightweight auxiliary models with memory retrieval, transferable across victim LLMs without retraining. The framework achieves <strong>99.61% defense success</strong> in filter mode with fewer parameters than existing guardrails, and supports both steering intrinsic model defenses and direct input filtering.</p></li></ul><p><em>Institutional affiliations: City University of Hong Kong, Beijing University of Posts and Telecommunications, Wuhan University, Beihang University, Beijing Academy of Artificial Intelligence</em></p><h1>On the Horizon</h1><p>There&#8217;s still limited information on the bilateral AI Track 1 dialogue&#8212;we&#8217;ll bring details as they come.</p><div><hr></div><p style="text-align: justify;">For more on how we select and track content, see our methodology <a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz">here</a>.</p><div><hr></div><p></p><p><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Thanks for reading! Subscribe for free to receive new posts and support my work.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>S&amp;T Daily is the official paper of the Ministry of Science and Technology (MOST).</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>&#12298;&#26234;&#33021;&#20307;&#35268;&#33539;&#24212;&#29992;&#19982;&#21019;&#26032;&#21457;&#23637;&#23454;&#26045;&#24847;&#35265;&#12299;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>&#23433;&#20840;&#21487;&#25511;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#21019;&#26032;&#39537;&#21160;/&#24212;&#29992;&#29301;&#24341;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#22831;&#23454;&#21457;&#23637;&#22522;&#30784;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>&#23432;&#29282;&#23433;&#20840;&#24213;&#32447;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>&#24378;&#21270;&#24212;&#29992;&#29301;&#24341;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p>&#24314;&#35774;&#21019;&#26032;&#29983;&#24577;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>&#26234;&#33021;&#20307;&#27880;&#20876;&#24179;&#21488;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>2026&#24180;&#20197;&#26469;&#65292;OpenClaw&#24191;&#27867;&#24212;&#29992; &#8230; &#26292;&#38706;&#20986;&#26234;&#33021;&#20307;&#22312;&#25351;&#20196;&#35825;&#23548;&#19979;&#21487;&#21457;&#36215;&#32593;&#32476;&#25915;&#20987;&#31561;&#39118;&#38505;&#38544;&#24739;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-11" href="#footnote-anchor-11" class="footnote-number" contenteditable="false" target="_self">11</a><div class="footnote-content"><p>&#22320;&#26041;&#23454;&#38469;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-12" href="#footnote-anchor-12" class="footnote-number" contenteditable="false" target="_self">12</a><div class="footnote-content"><p>&#21152;&#24555;&#25512;&#36827;&#20154;&#24037;&#26234;&#33021;... &#32508;&#21512;&#24615;&#31435;&#27861;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-13" href="#footnote-anchor-13" class="footnote-number" contenteditable="false" target="_self">13</a><div class="footnote-content"><p>&#21407;&#21017;&#24615;&#12289;&#21442;&#32771;&#24615;&#25216;&#26415;&#25991;&#20214;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-14" href="#footnote-anchor-14" class="footnote-number" contenteditable="false" target="_self">14</a><div class="footnote-content"><p style="text-align: justify;">&#8220;Dominance&#8221; in the sense of &#8220;control&#8221;: &#20154;&#31867;&#20027;&#23548;&#26435;&#24433;&#21709;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-15" href="#footnote-anchor-15" class="footnote-number" contenteditable="false" target="_self">15</a><div class="footnote-content"><p>&#38450;&#33539;&#20154;&#24037;&#26234;&#33021;&#33073;&#31163;&#20154;&#31867;&#30417;&#30563;&#25110;&#23041;&#32961;&#20154;&#31867;&#29983;&#23384;&#21457;&#23637;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-16" href="#footnote-anchor-16" class="footnote-number" contenteditable="false" target="_self">16</a><div class="footnote-content"><p style="text-align: justify;">&#37325;&#28857;&#35780;&#20272;&#20854;&#22833;&#25511;&#39118;&#38505;&#20197;&#21450;&#20854;&#23545;&#20135;&#19994;&#21644;&#31038;&#20250;&#30340;&#24433;&#21709;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-17" href="#footnote-anchor-17" class="footnote-number" contenteditable="false" target="_self">17</a><div class="footnote-content"><p>&#21487;&#20449;&#24212;&#29992;&#12289;&#38450;&#33539;&#22833;&#25511;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-18" href="#footnote-anchor-18" class="footnote-number" contenteditable="false" target="_self">18</a><div class="footnote-content"><p>&#20154;&#24037;&#26234;&#33021;&#23433;&#20840;&#26631;&#20934;&#20307;&#31995;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-19" href="#footnote-anchor-19" class="footnote-number" contenteditable="false" target="_self">19</a><div class="footnote-content"><p>&#26234;&#33021;&#20307;&#24212;&#29992;&#23433;&#20840;&#22522;&#26412;&#35201;&#27714;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-20" href="#footnote-anchor-20" class="footnote-number" contenteditable="false" target="_self">20</a><div class="footnote-content"><p>&#29305;&#26391;&#26222;&#24635;&#32479;&#35775;&#21326;&#26399;&#38388;&#65292;&#20004;&#22269;&#20803;&#39318;&#23601;&#20154;&#24037;&#26234;&#33021;&#38382;&#39064;&#36827;&#34892;&#20102;&#24314;&#35774;&#24615;&#20132;&#27969;&#65292;&#21516;&#24847;&#24320;&#23637;&#20154;&#24037;&#26234;&#33021;&#25919;&#24220;&#38388;&#23545;&#35805;&#12290;</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 3]]></title><description><![CDATA[Developments from 8/4/26-29/4/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-3</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-3</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Fri, 01 May 2026 18:52:09 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/01cd8232-3a2a-4947-bbd4-6ea189aa79c6_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p style="text-align: justify;">Welcome to Issue 3 of the China AI Bulletin, the latest on AI development, governance, and safety in China. Today&#8217;s highlights: China finalizes its Human-Like AI Interim Measures, DeepSeek and Moonshot launch new flagship models, and the NDRC blocks Meta&#8217;s $2B Manus acquisition.</p><p style="text-align: justify;"><em>Number of the week: <a href="https://mp.weixin.qq.com/s?__biz=MzI5ODk1NjY1MA%3D%3D&amp;mid=2247721583&amp;idx=1&amp;sn=b8cdfd6c52f439e8f5fa21a2a6098ac9&amp;scene=45&amp;poc_token=HPLS9Gmj2E_iR7_OA1KUPBw4qkseqrbRURU-w4nl">$100 million</a> - the amount Loopit, the &#8220;playable&#8221; TikTok-like app, raised.</em></p><p style="text-align: justify;"><strong>Editor&#8217;s note: the next edition will be the week of May 18</strong></p><h1 style="text-align: justify;">Executive Summary</h1><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/domestic-ai-governance">Domestic AI Governance</a>: </strong>The <strong>Interim Measures for the Management of Human-Like AI Interaction Services</strong> were finalized. Compared to the draft, it narrows scope to &#8220;sustained emotional interaction services,&#8221; bans virtual intimate relationships for minors, writes the &#8220;AI sandbox&#8221; into AI-specific Chinese legislation for the first time, and pitches China&#8217;s &#8220;system-level&#8221; approach as an alternative to EU AI Act risk-classification and US state-level disclosure laws. Enforcement also stepped up: CAC took action on April 28 against three <strong>ByteDance-owned products</strong> for synthetic-content labeling violations, and a Hangzhou court ruled &#8220;AI replacement&#8221; is not valid grounds for layoff or wage reduction.</p><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/national-standards">National Standards</a>:</strong> SAC began drafting a new generative AI risk-response guide and got China&#8217;s <strong>&#8220;Humanoid Robot Dataset&#8221; international standard approved at ISO</strong>. TC260 opened its <strong>AI Application Ethical Security Guidelines</strong> for comment (with Tsinghua&#8217;s Xue Lan as principal drafter and DeepSeek among contributors) and published a <strong>16-item AI safety standards wish list</strong> that functions as new working group WG9&#8217;s first formal roadmap.</p></li></ul></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/international-ai-governance">International AI Governance</a>:</strong> A CAC-published commentary by Zhi Zhenfeng (CASS) positions <strong>China as architect of an emerging international AI governance architecture</strong> anchored in the Global AI Governance Initiative, Action Plan, and Data Security Initiative; the framing is pitched at Global South nations.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/frontier-lab-developments">Frontier Lab Developments</a>:</strong></p><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/spotlight-three-new-flagship-models">Spotlights</a>: DeepSeek V4 Preview</strong> introduces a hybrid sparse-attention architecture and <strong>emphasizes cost-effective 1M-token context, </strong>while<strong> </strong>the technical report concedes V4 <strong>trails state-of-the-art frontier models by 3-6 months</strong>. <strong>Xiaomi&#8217;s</strong> <strong>MiMo-V2.5-Pro</strong> matches DeepSeek V4&#8217;s token context. <strong>Moonshot&#8217;s Kimi K2.6</strong> scales its agent swarm to <strong>300 sub-agents executing 4,000 coordinated steps</strong> and claims to beat GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro on HLE-Full with tools.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/notable-model-releases">Notable model releases</a>:</strong> Agentic models continued&#8212;Alibaba&#8217;s <strong>DR-Venus </strong>(4B-parameter edge research agent outperforming 9B-parameter baselines), Tencent&#8217;s <strong>HY-Embodied/HY-World 2.0</strong>, Alibaba&#8217;s <strong>AgenticQwen</strong> line, Zhipu&#8217;s GLM-5V-Turbo (native multimodal agent), plus <strong>visual-generation models</strong> from Alibaba, Baidu, and StepFun, and ByteDance&#8217;s <strong>AnewOmni</strong> for molecular design.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/technical-publication-highlights">Technical papers</a>:</strong> 94 papers from frontier labs this edition. Highlights: Huawei&#8217;s <strong>OneManCompany</strong> self-organizing agent firm and Alibaba&#8217;s <strong>TCOD</strong> temporal-curriculum distillation, improving multi-turn agent performance by up to 18 points.</p></li></ul></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/technical-ai-safety-publication-highlights">Technical AI Safety</a>:</strong> Agentic safety dominated again. Highlights: <strong>BadSkill</strong>, a supply-chain attack on AI agent skill ecosystems; <strong>CORA</strong>, a Conformal Risk Control framework giving mobile GUI agents statistical guarantees on harmful actions; <strong>SafeRedirect</strong>, addressing Internal Safety Collapse; and a study finding <strong>brief AI chatbot interactions produce lasting changes in human moral judgments</strong>, persisting and strengthening over two weeks while participants remained unaware of the influence.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/196139338/export-controls-and-economic-policy">Export Controls &amp; Economic Policy</a>:</strong> The NDRC ordered Meta to <strong>unwind its $2B Manus acquisition</strong> on April 27&#8212;the <strong>first major cross-border AI acquisition Beijing has blocked</strong>&#8212;asserting jurisdiction over the Singapore-incorporated holding company on the basis that Chinese-origin IP and talent are domestic assets regardless of registration.</p></li></ul><h1 style="text-align: justify;">Domestic AI Governance</h1><h2 style="text-align: justify;">Interim Human-Like AI Interaction Measures issued</h2><p style="text-align: justify;">The <strong><a href="https://mp.weixin.qq.com/s/6mBUtSdD6kkNh-T2yDrTNQ">Interim Measures for the Management of Human-Like AI Interaction Services</a></strong> were finalized April 10 (effective July 15) by five agencies jointly&#8212;the Cyberspace Administration of China (CAC) as lead, alongside the National Development and Reform Commission (NDRC), Ministry of Industry and Information Technology (MIIT), Ministry of Public Security (MPS), and State Administration for Market Regulation (SAMR)&#8212;after a <a href="https://mp.weixin.qq.com/s/WULVqbb5Gs222VVLSpkuyw">December 2025 draft</a>. <a href="https://www.geopolitechs.org/p/china-rolls-out-interim-regulations">Geopolitechs&#8217;s excellent analysis</a> identifies six concrete shifts from draft to final: (1) <strong>narrowed scope</strong>&#8212;the rules now target &#8220;sustained emotional interaction services&#8221; specifically, explicitly excluding customer service, Q&amp;A, and productivity tools that simulate humans only incidentally; (2) <strong>stronger protections for minors</strong>&#8212;a prohibition on &#8220;virtual intimate relationships&#8221; for minors, moving from warnings to product-form-level restrictions; (3) <strong>system-level governance</strong>, evolving from reactive content moderation to integrated oversight across training data, ethics review, lifecycle responsibility, and platform governance; (4) increased <strong>operational flexibility</strong> (mandatory human takeover requirements and the ban on virtual relatives for elderly users have been removed and the minor&#8217;s data protection audit frequency has been unfixed); (5) <strong>stronger enforcement mechanisms</strong> (fines, service suspension, and registration restrictions, with heavier sanctions for harm cases); and (6) <strong>explicit innovation-support provisions</strong> including algorithmic research, standards development, and sandbox testing.</p><p style="text-align: justify;">The CAC followed this with expert interpretations on April 10 and April 17. The <a href="http://www.cac.gov.cn/2026-04/10/c_1777558285804391.htm">April 10 lead piece by Wang Jiang</a> (Director and Party Secretary of the China Academy of Cyberspace Studies) frames the <strong>core tension as human-like AI services producing both beneficial capabilities and safety risks</strong>: the upside is addressing aging populations, loneliness, education, and cultural transmission; the downside is that algorithmic design produces &#8220;a <strong>persistent virtual &#8216;perfect relationship&#8217; without real-world commitment</strong>&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-1" href="#footnote-1" target="_self">1</a> that drives emotional dependency. Wang explicitly contrasts China&#8217;s approach with the EU&#8217;s AI Act risk-classification model (human-like AI services as &#8220;limited risk&#8221; with mandatory transparency) and US state-level laws (New York and California requiring disclosure plus age verification for human-like AI interactions), painting China&#8217;s approach as system-level rather than disclosure-based and portraying the framework as a &#8220;Chinese solution&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-2" href="#footnote-2" target="_self">2</a> for global AI governance. The companion April 10 piece by <a href="https://www.cac.gov.cn/2026-04/10/c_1777558285548271.htm">Yu Xiaohui</a> (CAICT Director, 14th National CPPCC member) names specific US state laws being mirrored and reframes the regulation around China&#8217;s <strong>&#8220;cognitive intelligence to emotional intelligence&#8221;</strong><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-3" href="#footnote-3" target="_self">3</a><strong> paradigm shift</strong>, explicitly citing global suicide and accidental-death cases linked to companion AI since 2023 as proximate cause for the legislation.</p><p style="text-align: justify;">The April 17 pieces provide additional angles. <a href="http://www.cac.gov.cn/2026-04/17/c_1778166820059937.htm">Zheng Qinghua</a> (Party Secretary of Tongji University and CAE academician) frames the Measures around &#8220;human-machine value alignment&#8221;:<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-4" href="#footnote-4" target="_self">4</a> AI must align with humans, and humans must in turn align with AI (i.e., use it responsibly). Zheng also flags that this is the first time that &#8220;AI sandboxes&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-5" href="#footnote-5" target="_self">5</a> have been written into AI-specific Chinese legislation, a governance milestone even if the sandbox provision itself is light on operational detail. <a href="http://www.cac.gov.cn/2026-04/17/c_1778166820056328.htm">Fan Kefeng</a> (Vice President of the China Electronics Standardization Institute) also highlights the sandbox provisions while emphasizing the safety and security risks undergirding the legislation. The third April 17 piece, by <a href="https://www.cac.gov.cn/2026-04/17/c_1778166820050890.htm">Chen Liang</a> (Dean of the AI Law School at Southwest University of Political Science and Law), reads as the most legally pragmatic of the set: it sorts the Measures&#8217; provisions into three implementation-path classes&#8212;standardizable behavioral directives, value-judgment &#8220;boundary prohibitions&#8221; that need responsible processes rather than negative-list rules, and framework principles that guide downstream standards&#8212;and pushes a whole-lifecycle &#8220;compliance by design&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-6" href="#footnote-6" target="_self">6</a> framing for Chinese AI companies.</p><h2 style="text-align: justify;">Enforcement steps up on three fronts</h2><p style="text-align: justify;">The CAC announced an <a href="http://www.cac.gov.cn/2026-04/28/c_1779119736411711.htm">administrative action against three ByteDance-owned products:</a> Jianying (CapCut), <a href="https://www.reuters.com/legal/litigation/chinas-cyberspace-regulator-warns-bytedance-apps-website-over-ai-content-2026-04-28/">Maoxiang, and Jimeng AI</a> on April 28 for violations of synthetic-content labeling rules. This comes after a <a href="https://www.cac.gov.cn/2026-02/12/c_1772636033171974.htm">February crackdown</a> on unlabeled AI-generated content across platforms and the <a href="https://www.whatsonweibo.com/quick-eye-ai-drama-takedowns-metas-blocked-deal-and-lying-flat-conspiracy/">reported removal</a> of many popular AI-generated short dramas from Douyin and Hongguo.</p><p style="text-align: justify;">Separately, a <a href="https://china.caixin.com/2026-04-29/102439348.html">Hangzhou court ruled</a> on April 29 that AI replacement is not a valid reason for layoff or wage reduction, which, as a &#8220;typical case&#8221; meant to establish precedent, may set a floor for AI-driven labor disputes.</p><p style="text-align: justify;">Finally, the government has <a href="https://www.bloomberg.com/news/articles/2026-04-29/china-suspends-new-autonomous-driving-permits-after-baidu-outage">reportedly suspended new licenses</a> for autonomous vehicles after over <a href="https://www.forbes.com/sites/bradtempleton/2026/04/05/baidu-silent-about-failure-of-100-robotaxis-in-wuhan/">100 Baidu robotaxis froze</a> in Wuhan in late March, snarling traffic and stranding passengers. While existing AVs can continue to operate (except for Baidu&#8217;s Wuhan fleet, which is on pause), companies <a href="https://www.theverge.com/ai-artificial-intelligence/920312/china-suspends-autonomous-vehicle-permits-baidu-chaos">won&#8217;t be able to</a> expand fleets, launch in new cities, or start new test projects.</p><p style="text-align: justify;">Taken together, these three cases show a limited tolerance for AI-driven disruption and a willingness to lean on regulations, the judicial system, and permitting regimes to maintain development while preserving stability.</p><h2 style="text-align: justify;">National Standards</h2><h3 style="text-align: justify;">SAC drafts new generative AI risk-response guide</h3><p style="text-align: justify;">The Standardization Administration of China has begun drafting a new recommended national standard, <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=508271370C98A41CE06397BE0A0A979A">&#8220;Artificial Intelligence &#8212; Guidance on addressing risks in generative AI systems&#8221;</a> (&#20154;&#24037;&#26234;&#33021; &#29983;&#25104;&#24335;&#20154;&#24037;&#26234;&#33021;&#31995;&#32479;&#39118;&#38505;&#24212;&#23545;&#25351;&#21335;, plan number 20262577-Z-469). The China Electronics Standardization Institute (CESI) and Alibaba Cloud are the lead drafters, with a 12-month drafting window starting April 28. The standard sits under TC28/SC42&#8212;the AI subcommittee under the IT standardization subcommittee. The closest existing standard is <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=33D40F1160BF5D92E06397BE0A0A5B93">GB/T 45654-2025</a> on basic security requirements for generative AI services (produced by TC260, the cybersecurity standardization committee). Where 45654-2025 fixes minimum requirements at the service-provider level, this new standard seems aimed at giving developers and providers a structured playbook for handling risks once they materialize, following on TC260&#8217;s September 2025 genAI Emergency Response Guidelines (<a href="https://www.tc260.org.cn/upload/2025-09-15/1757913511246034883.pdf">pdf</a>).</p><h3 style="text-align: justify;">China pushes humanoid robot dataset standards at home and at ISO</h3><p style="text-align: justify;">SAC&#8217;s Standards Innovation Department <a href="https://www.sac.gov.cn/xw/tzgg/art/2026/art_cf0339f385214e6a89265ffdf334385e.html">announced on April 23</a> that China&#8217;s proposed &#8220;Humanoid Robot Dataset&#8221; international standard has been formally approved for project establishment at ISO, and is now recruiting domestic experts to coordinate China&#8217;s position across the international drafting process. Applications closed April 30. The international push runs in parallel with a <a href="https://std.samr.gov.cn/gb/search/gbDetailed?id=30D202B3ADBFE9F6E06397BE0A0A4001">domestic standard </a>already being drafted: plan 20253226-T-604 (Humanoid Robot Dataset&#8212;Part 1: General<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-7" href="#footnote-7" target="_self">7</a>) under TC591, the National Robotics Standardization Technical Committee. The drafting team includes the Beijing Mechanical Industry Automation Research Institute, Tsinghua, the Shenzhen Institute of AI and Robotics, Shanghai AI Lab, and Unitree, pairing state research institutes with a leading commercial humanoid maker.</p><h3 style="text-align: justify;">TC260 opens AI Application Ethical Security Guidelines for comment</h3><p style="text-align: justify;">TC260 is <a href="https://www.tc260.org.cn/portal/article/2/0e3933e11d904f4b93c4dcc348f9b61c">soliciting comments through April 26</a> on a draft technical document, &#8220;Ethical Security Guidelines for Artificial Intelligence Applications 1.0&#8221; (&#20154;&#24037;&#26234;&#33021;&#24212;&#29992;&#20262;&#29702;&#23433;&#20840;&#25351;&#24341; 1.0). The document is led by Tsinghua University, with Xue Lan&#8212;dean of Schwarzman College and director of Tsinghua&#8217;s Institute for AI International Governance (I-AIIG)&#8212;as principal drafter. The drafting team includes Tsinghua, CESI, Shanghai Jiao Tong, Sichuan University, University of Science and Technology Beijing, Alibaba, Huawei, and DeepSeek. The guidelines codify five &#8220;ethical security impact&#8221; categories (weakening of human primacy, breakdown of basic social order, decoupling of humans from physical society, social stratification and discrimination, and individual rights infringement) and six principles (people-centric, safety/controllability, fairness, transparency, co-governance, and inclusive sharing). Operational guidance is split across four roles&#8212;a general baseline plus separate sections for developers, service providers, and users&#8212;with provisions like default-on safety/fairness/privacy settings, mandatory black-box-style incident traceability, and explicit appeals and redress mechanisms for users. As a TC260 technical document rather than a national standard, it&#8217;s non-binding, but it consolidates ethical-governance threads that have surfaced across the <a href="https://www.cac.gov.cn/2025-09/15/c_1759653448369123.htm">AI Safety Governance Framework</a> and earlier AI ethics white papers and principle-sets.</p><h3 style="text-align: justify;">TC260&#8217;s wish list reveals WG9&#8217;s planned safety pipeline</h3><p>TC260 also <a href="https://www.tc260.org.cn/portal/article/2/1ba8515d8e2a43f6918b05bcf663655b">opened public consultation through May 6</a> on its second batch of 2026 cybersecurity national standards needs. The list contains 16 AI-related items, all assigned to the new AI Safety Standards Working Group (WG9), whose leadership we covered <a href="https://chinaaibulletin.substack.com/i/193694600/tc260-working-group-on-ai-safetysecurity-leadership-announced">last issue</a>. Items 1&#8211;9 are recommended national standards (GB/T); items 10&#8211;16 are GB/Z guidance documents, which carry less regulatory weight but can signal where sectoral application rules are headed.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!u9ne!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!u9ne!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 424w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 848w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 1272w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!u9ne!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png" width="857" height="720" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c1035c90-7f3f-4d13-957e-afc71213d887_857x720.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:720,&quot;width&quot;:857,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:131317,&quot;alt&quot;:&quot;Table of 16 AI safety standards in TC260 working group WG9's 2026 pipeline. Items 1&#8211;9 are recommended   &#9614;  national standards (GB/T): on-device large model cybersecurity, AI deep synthesis security    &#9614; specifications, embodied AI security requirements, AI security evaluation organization capability       &#9614; requirements, AI foundation model security testing methods, AI model development security guidelines,   &#9614;  open-source AI model security guidelines, AI training and inference framework security requirements,   &#9614;  and generative AI system interoperability security specifications. Items 10&#8211;16 are GB/Z guidance   &#9614; documents covering sectoral AI applications: broadcast/TV, education, finance, energy, health,   &#9614; emergency management, and government affairs.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/196139338?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of 16 AI safety standards in TC260 working group WG9's 2026 pipeline. Items 1&#8211;9 are recommended   &#9614;  national standards (GB/T): on-device large model cybersecurity, AI deep synthesis security    &#9614; specifications, embodied AI security requirements, AI security evaluation organization capability       &#9614; requirements, AI foundation model security testing methods, AI model development security guidelines,   &#9614;  open-source AI model security guidelines, AI training and inference framework security requirements,   &#9614;  and generative AI system interoperability security specifications. Items 10&#8211;16 are GB/Z guidance   &#9614; documents covering sectoral AI applications: broadcast/TV, education, finance, energy, health,   &#9614; emergency management, and government affairs." title="Table of 16 AI safety standards in TC260 working group WG9's 2026 pipeline. Items 1&#8211;9 are recommended   &#9614;  national standards (GB/T): on-device large model cybersecurity, AI deep synthesis security    &#9614; specifications, embodied AI security requirements, AI security evaluation organization capability       &#9614; requirements, AI foundation model security testing methods, AI model development security guidelines,   &#9614;  open-source AI model security guidelines, AI training and inference framework security requirements,   &#9614;  and generative AI system interoperability security specifications. Items 10&#8211;16 are GB/Z guidance   &#9614; documents covering sectoral AI applications: broadcast/TV, education, finance, energy, health,   &#9614; emergency management, and government affairs." srcset="https://substackcdn.com/image/fetch/$s_!u9ne!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 424w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 848w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 1272w, https://substackcdn.com/image/fetch/$s_!u9ne!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc1035c90-7f3f-4d13-957e-afc71213d887_857x720.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Standard titles and general themes</figcaption></figure></div><p style="text-align: justify;">The list maps closely onto the gaps WG9 identified at its first meeting&#8212;frontier risk evaluation, urgently needed standards in key areas, and pilot applications across sectors&#8212;and gives an early read on what WG9&#8217;s first formal standards push will look like. Additionally, the open-source security standard explicitly references the AI Safety Governance Framework, suggesting WG9 is positioned to translate that high-level framework into binding-track recommended standards, and the foundation model testing methods standard could complement GB/T 45654-2025, which emphasizes content safety by providing testing methodologies for issues it treats more lightly.</p><h1 style="text-align: justify;">International AI Governance</h1><h2 style="text-align: justify;">Global AI governance emphasized in expert commentary</h2><p style="text-align: justify;">On April 29, <strong>Zhi Zhenfeng</strong>, Director of the Xi Jinping Rule of Law Thought Research Office at the Chinese Academy of Social Sciences, published a <a href="http://www.cac.gov.cn/2026-04/29/c_1779117458770614.htm">CAC expert interpretation</a> on &#8220;&#8203;&#8203;China&#8217;s Responsibility in International Rule of Law in Cyberspace.&#8221; Zhi positions China as the architect of an emerging international AI governance architecture anchored in three documents: the <strong><a href="https://www.mfa.gov.cn/eng/zy/gb/202405/t20240531_11367503.html">Global AI Governance Initiative</a></strong> (October 2023), the <strong><a href="https://www.fmprc.gov.cn/mfa_eng/xw/zyxw/202507/t20250729_11679232.html">Global AI Governance Action Plan</a></strong> (a 13-point roadmap <a href="https://english.news.cn/20250814/8822ce54e13b492d8737bb338bd6e474/c.html">announced by Premier Li Qiang at WAIC 2025</a> in Shanghai last July to operationalize the Initiative), and the <strong><a href="https://www.mfa.gov.cn/eng/wjb/zzjg_663340/dozys_664276/xwlb_664278/202406/t20240606_11397627.html">Global Data Security Initiative</a></strong> (Foreign Minister Wang Yi&#8217;s 2020 proposal). The piece reiterates China&#8217;s standing call for a <strong><a href="https://en.chinadiplomacy.org.cn/2025-07/30/content_118003645.shtml">World Artificial Intelligence Cooperation Organization (WAICO)</a></strong>&#8212;also proposed by Li Qiang at WAIC 2025 with a tentative Shanghai headquarters and <a href="https://www.fmprc.gov.cn/eng/zy/jj/xjpcx32apecbdhgjxgsfw/202511/t20251101_11745402.html">reaffirmed by Xi Jinping</a> at the November 2025 APEC summit&#8212;framed as an alternative international AI institution. The framing is  &#8220;open and inclusive; safe and controllable&#8221;<a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-8" href="#footnote-8" target="_self">8</a> and is explicitly pitched at Global South nations.</p><p style="text-align: justify;"><em>Nb: Zhi was also quoted in <a href="https://www.bjrd.gov.cn/xwzx/fzlt/202604/t20260415_4582348.html">Beijing People&#8217;s Congress&#8217;s Apr 15 forum</a> saying that the conditions for a unified domestic AI law aren&#8217;t in place,</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-9" href="#footnote-9" target="_self">9</a><em> and emphasizing &#8220;small, fast, and agile&#8221;</em><a class="footnote-anchor" data-component-name="FootnoteAnchorToDOM" id="footnote-anchor-10" href="#footnote-10" target="_self">10</a><em> targeted regulation.</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Find this helpful? Subscribe to get regular editions in your inbox.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p><h1 style="text-align: justify;">Frontier Lab Developments</h1><h2 style="text-align: justify;">&#128269;Spotlight: Three new flagship models</h2><p style="text-align: justify;"><strong>DeepSeek</strong> released a <strong><a href="https://api-docs.deepseek.com/news/news260424">DeepSeek V4 Preview</a></strong> on April 24. Per <a href="https://www.chinatalk.media/p/deepseek-v4">ChinaTalk</a> and <a href="https://36kr.com/p/3780375304312072?f=rss">36Kr&#8217;s</a> reporting, the release was significantly delayed by the team&#8217;s migration from Nvidia to Huawei chips and by internal disagreement. The release covers two variants: <strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro">DeepSeek-V4-Pro</a></strong> (1.6T-parameter total, 49B-parameter active) and <strong><a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Flash">DeepSeek-V4-Flash</a></strong> (284B-parameter total, 13B-parameter active). The headline framing is <strong>cost-effective 1M-token context as the new default</strong> across DeepSeek services, a direct contrast to premium-priced 1M-token context models and a rebuttal to the Chinese lab <a href="https://substack.com/@nickcorvino/p-193608384">trend of close-sourcing models</a>. However, the <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">technical report</a> concedes V4 &#8220;trails state-of-the-art frontier models by approximately 3 to 6 months&#8221; on reasoning&#8212;potentially one reason why it&#8217;s being marketed as a &#8220;preview&#8221; rather than the final model.</p><p style="text-align: justify;">Architecturally, the model uses a <strong>hybrid attention mechanism</strong> combining two layers: <strong>Compressed Sparse Attention (CSA)</strong>, which collapses small blocks of KV entries into single compressed entries and then applies DeepSeek Sparse Attention so each query attends only to the top-k compressed entries; and <strong>Heavily Compressed Attention (HCA)</strong>, a more aggressive compression layer (with a much larger compression ratio) that retains dense attention. DeepSeek itself flags the design as a tradeoff&#8212;the conclusion section concedes the team &#8220;retained many preliminarily validated components and tricks, which, while effective, made the architecture relatively complex,&#8221; and signals plans to &#8220;distill the architecture down to its most essential designs&#8221; in future iterations. Compared to DeepSeek-V3.2, the model is more efficient: at 1M-token context, V4-Pro requires <strong>27% of single-token inference FLOPs and 10% of the KV cache</strong> of DeepSeek-V3.2 (3.7x and 9.5x reductions); V4-Flash drops further to <strong>10% of FLOPs and 7% of KV cache</strong> (9.8x and 13.7x) vs DeepSeek-V3.2. However, they don&#8217;t evaluate efficiency against other models; those claims will likely rest on API pricing.</p><p style="text-align: justify;">On capabilities, the blog&#8217;s claim of V4 being the <strong>&#8220;open-source SOTA on agentic coding&#8221;</strong> is complicated somewhat by the technical report&#8217;s benchmarks.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!bnP0!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!bnP0!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 424w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 848w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 1272w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!bnP0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png" width="704" height="589" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:589,&quot;width&quot;:704,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Table 6 from the DeepSeek V4 technical report comparing DeepSeek-V4-Pro-Max against Claude Opus 4.6     &#9614; (Max), GPT-5.4 (xHigh), Gemini 3.1 Pro (High), Kimi K2.6 Thinking, and GLM-5.1 Thinking across    &#9614; knowledge and reasoning benchmarks (MMLU-Pro, SimpleQA-Verified, Chinese-SimpleQA, GPQA Diamond, HLE,   &#9614;  LiveCodeBench, Codeforces, HMMT 2026 Feb, IMOAnswerBench, Apex), long-context benchmarks (MRCR 1M,    &#9614; CorpusQA 1M), and agentic benchmarks (Terminal Bench 2.0, SWE Verified, SWE Pro, SWE Multilingual,   &#9614; BrowseComp, HLE with tools, GDPval-AA, MCPAtlas Public, Toolathlon).&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table 6 from the DeepSeek V4 technical report comparing DeepSeek-V4-Pro-Max against Claude Opus 4.6     &#9614; (Max), GPT-5.4 (xHigh), Gemini 3.1 Pro (High), Kimi K2.6 Thinking, and GLM-5.1 Thinking across    &#9614; knowledge and reasoning benchmarks (MMLU-Pro, SimpleQA-Verified, Chinese-SimpleQA, GPQA Diamond, HLE,   &#9614;  LiveCodeBench, Codeforces, HMMT 2026 Feb, IMOAnswerBench, Apex), long-context benchmarks (MRCR 1M,    &#9614; CorpusQA 1M), and agentic benchmarks (Terminal Bench 2.0, SWE Verified, SWE Pro, SWE Multilingual,   &#9614; BrowseComp, HLE with tools, GDPval-AA, MCPAtlas Public, Toolathlon)." title="Table 6 from the DeepSeek V4 technical report comparing DeepSeek-V4-Pro-Max against Claude Opus 4.6     &#9614; (Max), GPT-5.4 (xHigh), Gemini 3.1 Pro (High), Kimi K2.6 Thinking, and GLM-5.1 Thinking across    &#9614; knowledge and reasoning benchmarks (MMLU-Pro, SimpleQA-Verified, Chinese-SimpleQA, GPQA Diamond, HLE,   &#9614;  LiveCodeBench, Codeforces, HMMT 2026 Feb, IMOAnswerBench, Apex), long-context benchmarks (MRCR 1M,    &#9614; CorpusQA 1M), and agentic benchmarks (Terminal Bench 2.0, SWE Verified, SWE Pro, SWE Multilingual,   &#9614; BrowseComp, HLE with tools, GDPval-AA, MCPAtlas Public, Toolathlon)." srcset="https://substackcdn.com/image/fetch/$s_!bnP0!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 424w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 848w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 1272w, https://substackcdn.com/image/fetch/$s_!bnP0!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fee99b340-88e9-4baf-9945-4fe72749dfc4_704x589.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf">DeepSeek-V4 Technical Report</a></figcaption></figure></div><p style="text-align: justify;">V4-Pro does lead on raw coding (first in LiveCodeBench and Codeforces) and on open-source world knowledge (trailing only Gemini 3.1 Pro on SimpleQA-Verified). But on agentic coding specifically, the picture is mixed: V4-Pro leads open-weight peers on SWE Verified, Toolathlon, and MCPAtlas (and is more or less level with Western closed models on SWE Verified), but it trails on SWE-Pro and finishes last on HLE-with-tools. The report itself concedes this directly, but it&#8217;s potentially more accurate to say that the model is the open-weight SOTA on isolated coding and knowledge, but more mixed on agentic. However, it does establish the open-weight high-mark on long-context tasks; its 1M token window is 4-5x that of K2.6 and GLM-5.1.</p><p style="text-align: justify;">DeepSeek also flags integration with <strong>Claude Code, OpenClaw, and OpenCode</strong>, explicitly courting the proactive-agent harness ecosystem. (Also worth noting: DeepSeek has<a href="https://mp.weixin.qq.com/s/S65DyDmxovfH4bEuoHmVxg"> launched multimodal capabilities in testing</a>.)</p><p style="text-align: justify;"><strong>Xiaomi</strong> released <strong><a href="https://mimo.xiaomi.com/mimo-v2-5-pro/">MiMo-V2.5-Pro</a></strong> on April 27, completing a two-step rollout that began with the standard <strong><a href="https://mimo.xiaomi.com/mimo-v2-5/">MiMo-V2.5</a></strong> on April 22. Pro is a <strong>1.02T-parameter total, 42B-parameter active MoE</strong> with a <strong>1M-token context window</strong>&#8212;matching DeepSeek V4 Pro&#8212;and a hybrid-attention architecture interleaving Local Sliding Window Attention (SWA) and Global Attention (GA) at a 6:1 ratio with 128-token windows, plus Multi-Token Prediction. The standard V2.5 is 310B-parameter total, 15B-parameter active, multimodal across text, image, and audio. Both models are open-sourced under MIT license (<a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.5-Pro">HuggingFace</a>). Xiaomi positions Pro for general agentic capabilities, complex software engineering, and long-horizon tasks spanning over 1,000 tool calls. Its headline benchmark claim is <strong>64% Pass^3 on ClawEval using only ~70K tokens per trajectory</strong>&#8212;roughly 40-60% fewer tokens than Claude Opus 4.6, Gemini 3.1 Pro, and GPT-5.4.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!pini!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!pini!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 424w, https://substackcdn.com/image/fetch/$s_!pini!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 848w, https://substackcdn.com/image/fetch/$s_!pini!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 1272w, https://substackcdn.com/image/fetch/$s_!pini!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!pini!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png" width="741" height="575" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:575,&quot;width&quot;:741,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Bar charts comparing Xiaomi MiMo-V2.5-Pro and MiMo-V2.5 against MiMo-V2-Pro, Claude Opus 4.6, Gemini    &#9614; 3.1 Pro, and GPT-5.4 across coding-agent benchmarks (SWE-Bench Pro, MiMo Coding Bench, Terminal-Bench   &#9614;  2.0, FrontierSWE), general-agent benchmarks (GDPVal-AA, &#964;3-bench, Claw-Eval pass^3), and a reasoning   &#9614;  benchmark (Humanity's Last Exam, with and without tools). MiMo-V2.5-Pro leads peers on most    &#9614; coding-agent benchmarks.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Bar charts comparing Xiaomi MiMo-V2.5-Pro and MiMo-V2.5 against MiMo-V2-Pro, Claude Opus 4.6, Gemini    &#9614; 3.1 Pro, and GPT-5.4 across coding-agent benchmarks (SWE-Bench Pro, MiMo Coding Bench, Terminal-Bench   &#9614;  2.0, FrontierSWE), general-agent benchmarks (GDPVal-AA, &#964;3-bench, Claw-Eval pass^3), and a reasoning   &#9614;  benchmark (Humanity's Last Exam, with and without tools). MiMo-V2.5-Pro leads peers on most    &#9614; coding-agent benchmarks." title="Bar charts comparing Xiaomi MiMo-V2.5-Pro and MiMo-V2.5 against MiMo-V2-Pro, Claude Opus 4.6, Gemini    &#9614; 3.1 Pro, and GPT-5.4 across coding-agent benchmarks (SWE-Bench Pro, MiMo Coding Bench, Terminal-Bench   &#9614;  2.0, FrontierSWE), general-agent benchmarks (GDPVal-AA, &#964;3-bench, Claw-Eval pass^3), and a reasoning   &#9614;  benchmark (Humanity's Last Exam, with and without tools). MiMo-V2.5-Pro leads peers on most    &#9614; coding-agent benchmarks." srcset="https://substackcdn.com/image/fetch/$s_!pini!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 424w, https://substackcdn.com/image/fetch/$s_!pini!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 848w, https://substackcdn.com/image/fetch/$s_!pini!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 1272w, https://substackcdn.com/image/fetch/$s_!pini!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2d2d7731-a8b9-48ab-aa79-6864ee8bdf83_741x575.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://mimo.xiaomi.com/mimo-v2-5-pro/">MiMo-V2.5-Pro blog post</a></figcaption></figure></div><p style="text-align: justify;"><strong>Moonshot</strong> released <strong><a href="https://www.kimi.com/blog/kimi-k2-6.html">Kimi K2.6</a></strong> on April 14, a <strong>1T-parameter total, 32B-parameter active MoE multimodal model</strong> with a 256K-token context, 384 experts, and a 400M-parameter MoonViT vision encoder (<a href="https://huggingface.co/moonshotai/Kimi-K2.6">Hugging Face</a>). The release is positioned around four pillars: <strong>long-horizon coding</strong>, <strong>coding-driven design</strong> for UIs, an <strong>elevated agent swarm</strong>, and <strong>proactive autonomous execution.</strong> The agent swarm<strong> </strong>is a scale-up from K2.5, claiming <strong>300 sub-agents executing 4,000 coordinated steps</strong> (compared to K2.5&#8217;s 100 sub-agents and ~1,500 tool calls). However, the marginal benchmark gain from running the larger swarm is modest&#8212;<strong>BrowseComp jumps from 83.2 single-agent to 86.3 with swarm</strong>, a +3.1-point lift for 3x more sub-agents&#8212;which raises a proportionate-utility question. On the headline agentic benchmarks, Moonshot claims K2.6 beats GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro on <strong>HLE-Full with tools</strong> (54.0 compared to 52.1, 53.0, and 51.4) with similar ordering on DeepSearchQA.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!OfsI!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!OfsI!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 424w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 848w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 1272w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!OfsI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png" width="1000" height="374" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:374,&quot;width&quot;:1000,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Kimi K2.6 agentic benchmark table comparing K2.6 against GPT-5.4 (xhigh), Claude Opus 4.6 (max          &#9614; effort), Gemini 3.1 Pro (thinking high), and Kimi K2.5. K2.6 leads on HLE-Full with tools (54.0),    &#9614; BrowseComp with agent swarm (86.3), and DeepSearchQA f1-score (92.5) and accuracy (83.0).   &quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Kimi K2.6 agentic benchmark table comparing K2.6 against GPT-5.4 (xhigh), Claude Opus 4.6 (max          &#9614; effort), Gemini 3.1 Pro (thinking high), and Kimi K2.5. K2.6 leads on HLE-Full with tools (54.0),    &#9614; BrowseComp with agent swarm (86.3), and DeepSearchQA f1-score (92.5) and accuracy (83.0).   " title="Kimi K2.6 agentic benchmark table comparing K2.6 against GPT-5.4 (xhigh), Claude Opus 4.6 (max          &#9614; effort), Gemini 3.1 Pro (thinking high), and Kimi K2.5. K2.6 leads on HLE-Full with tools (54.0),    &#9614; BrowseComp with agent swarm (86.3), and DeepSearchQA f1-score (92.5) and accuracy (83.0).   " srcset="https://substackcdn.com/image/fetch/$s_!OfsI!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 424w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 848w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 1272w, https://substackcdn.com/image/fetch/$s_!OfsI!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1a01211b-2ce3-41cb-bb38-d7abedd34348_1000x374.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!0I26!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!0I26!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 424w, https://substackcdn.com/image/fetch/$s_!0I26!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 848w, https://substackcdn.com/image/fetch/$s_!0I26!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 1272w, https://substackcdn.com/image/fetch/$s_!0I26!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!0I26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png" width="984" height="280" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/d9c740d5-135a-4d64-9db3-5d518118db44_984x280.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:280,&quot;width&quot;:984,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;Kimi K2.6 coding benchmark table comparing K2.6 against GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, and   &#9614;  Kimi K2.5 on Terminal-Bench 2.0 Terminus-2 (66.7), SWE-Bench Pro (58.6), SWE-Bench Multilingual    &#9614; (76.7), and SWE-Bench Verified (80.2).   &quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Kimi K2.6 coding benchmark table comparing K2.6 against GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, and   &#9614;  Kimi K2.5 on Terminal-Bench 2.0 Terminus-2 (66.7), SWE-Bench Pro (58.6), SWE-Bench Multilingual    &#9614; (76.7), and SWE-Bench Verified (80.2).   " title="Kimi K2.6 coding benchmark table comparing K2.6 against GPT-5.4, Claude Opus 4.6, Gemini 3.1 Pro, and   &#9614;  Kimi K2.5 on Terminal-Bench 2.0 Terminus-2 (66.7), SWE-Bench Pro (58.6), SWE-Bench Multilingual    &#9614; (76.7), and SWE-Bench Verified (80.2).   " srcset="https://substackcdn.com/image/fetch/$s_!0I26!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 424w, https://substackcdn.com/image/fetch/$s_!0I26!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 848w, https://substackcdn.com/image/fetch/$s_!0I26!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 1272w, https://substackcdn.com/image/fetch/$s_!0I26!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fd9c740d5-135a-4d64-9db3-5d518118db44_984x280.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://www.kimi.com/blog/kimi-k2-6">Kimi K2.6 blog post</a></figcaption></figure></div><p style="text-align: justify;">Moonshot&#8217;s <strong>&#8220;Proactive and Open Orchestration&#8221;</strong> framing&#8212;advertising 24/7 background agents that &#8220;manage schedules, execute code, and orchestrate cross-platform operations without human oversight&#8221;&#8212;explicitly pitches K2.6 as the backend for <strong>OpenClaw and other proactive-agent stacks</strong> and attempts to counter the &#8220;<a href="https://36kr.com/p/3787727645962758">Kimi is falling behind</a>&#8221; narrative present in some domestic commentary.</p><h2 style="text-align: justify;">Notable Model Releases</h2><p style="text-align: justify;"><strong>Zhipu </strong>launched <strong><a href="https://arxiv.org/abs/2604.26752">GLM-5V-Turbo</a></strong>, positioned as a native foundation model for multimodal agents.<strong> Alibaba</strong> released <strong><a href="https://wan.video/">Wan-Image</a></strong> and its <a href="https://arxiv.org/abs/2604.19858">technical report</a> as their flagship generative visual intelligence model. <strong>Baidu</strong> announced <strong><a href="https://huggingface.co/baidu/ERNIE-Image">ERNIE-Image</a></strong> on their <a href="https://ernie.baidu.com/blog/posts/ernie-image/">blog</a> as a new high-performance open image model. <strong>StepFun</strong> shipped <strong><a href="https://github.com/stepfun-ai/Step-Audio-R1">Step-Audio-1.5</a></strong> and its <a href="https://arxiv.org/abs/2604.25719">technical report</a>, extending reasoning training to spoken tasks; <strong>ByteDance</strong> complemented this with the <strong><a href="https://seed.bytedance.com/en/blog/introducing-seed-full-duplex-speech-llm-attentive-listening-robust-interference-suppression-enabling-more-natural-interaction">Seed Full-Duplex Speech LLM</a></strong>, a conversational speech model that uses a &#8220;listen while speaking&#8221; paradigm that claims improvement over traditional half-duplex models.</p><p style="text-align: justify;">In line with the cycle&#8217;s agentic dominance, several labs shipped agent-focused models. <strong>Alibaba&#8217;s</strong> <strong>DR-Venus</strong> (<a href="https://github.com/inclusionAI/DR-Venus">GitHub</a>, <a href="https://huggingface.co/collections/inclusionAI/dr-venus">Hugging Face</a>) is a 4B-parameter edge-scale research agent trained on just 10K open examples while outperforming 9B-parameter baselines. <strong>Alibaba&#8217;s</strong> <strong><a href="https://arxiv.org/abs/2604.21590">AgenticQwen</a></strong> line-up (<a href="https://huggingface.co/collections/alibaba-pai/agenticqwen">Hugging Face</a>) covers small Qwen variants fine-tuned for industrial agentic tasks via dual data flywheels. Following up on its <a href="https://chinaaibulletin.substack.com/i/193694600/spotlight">promise to open-source small variants of Qwen 3.6</a>, Alibaba also <a href="https://qwen.ai/blog?id=qwen3.6-35b-a3b">released</a> <a href="https://huggingface.co/Qwen/Qwen3.6-35B-A3B?spm=a2ty_o06.30285417.0.0.610fc921Oytpwt&amp;file=Qwen3.6-35B-A3B">Qwen3.6-35B-A3B</a>, an open-source MoE-style variant. <strong>Tencent</strong> released <strong><a href="https://github.com/Tencent-Hunyuan/HY-SOAR">HY-SOAR</a></strong>, a search-and-research agent, alongside <strong>HY-Embodied-0.5-X</strong> (<a href="https://github.com/Tencent-Hunyuan/HY-Embodied-0.5-X">GitHub</a>, <a href="https://huggingface.co/tencent/HY-Embodied-0.5-X">Hugging Face</a>) for embodied AI research.</p><p style="text-align: justify;"><strong>Tencent</strong> also released a world model <strong>HY-World 2.0</strong> (<a href="https://github.com/Tencent-Hunyuan/HY-World-2.0">GitHub</a>, <a href="https://huggingface.co/tencent/HY-World-2.0">Hugging Face</a>), as well as <strong>Hy3-preview</strong> (<a href="https://github.com/Tencent-Hunyuan/Hy3-preview">GitHub</a>, <a href="https://huggingface.co/tencent/Hy3-preview">Hugging Face</a>), an early Hunyuan 3 preview, plus <strong><a href="https://huggingface.co/tencent/Hy-MT1.5-1.8B-1.25bit">Hy-MT1.5-1.8B quantized variants</a></strong>.</p><p style="text-align: justify;">ByteDance released <strong><a href="https://github.com/bytedance/AnewOmni">AnewOmni</a></strong> for generative molecular design (and its <a href="https://www.biorxiv.org/content/10.64898/2026.03.12.711044v2">bioRxiv</a> paper), plus <strong><a href="https://seed.bytedance.com/en/blog/seed3d-2-0-released-higher-precision-and-greater-usability">Seed3D 2.0</a></strong>&#8212;an updated 3D-generation model framed as offering higher precision and greater usability.</p><h2 style="text-align: justify;">Technical Publication Highlights</h2><p>Frontier labs released 94 papers on arXiv this edition. Highlights are below; a full list with summaries can be found <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0">here</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!wuLU!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!wuLU!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 424w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 848w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 1272w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!wuLU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png" width="632" height="567" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:567,&quot;width&quot;:632,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:45227,&quot;alt&quot;:&quot;Table of papers published by frontier Chinese AI labs this edition with cumulative totals since the     &#9614; bulletin's launch. Alibaba 35 papers (87 cumulative), Baidu 4 (17), ByteDance 11 (25), Huawei 12    &#9614; (43), iFlyTek 1 (4), Meituan 7 (14), StepFun 1 (3), Tencent 19 (44), Xiaomi 3 (12), and Zhipu 1 (3).    &#9614; Baichuan, DeepSeek, MiniMax, Moonshot, SenseTime, and 01.AI published nothing this edition.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/196139338?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Table of papers published by frontier Chinese AI labs this edition with cumulative totals since the     &#9614; bulletin's launch. Alibaba 35 papers (87 cumulative), Baidu 4 (17), ByteDance 11 (25), Huawei 12    &#9614; (43), iFlyTek 1 (4), Meituan 7 (14), StepFun 1 (3), Tencent 19 (44), Xiaomi 3 (12), and Zhipu 1 (3).    &#9614; Baichuan, DeepSeek, MiniMax, Moonshot, SenseTime, and 01.AI published nothing this edition." title="Table of papers published by frontier Chinese AI labs this edition with cumulative totals since the     &#9614; bulletin's launch. Alibaba 35 papers (87 cumulative), Baidu 4 (17), ByteDance 11 (25), Huawei 12    &#9614; (43), iFlyTek 1 (4), Meituan 7 (14), StepFun 1 (3), Tencent 19 (44), Xiaomi 3 (12), and Zhipu 1 (3).    &#9614; Baichuan, DeepSeek, MiniMax, Moonshot, SenseTime, and 01.AI published nothing this edition." srcset="https://substackcdn.com/image/fetch/$s_!wuLU!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 424w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 848w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 1272w, https://substackcdn.com/image/fetch/$s_!wuLU!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F9d0445b3-ac62-48df-99e1-a8330ceeea17_632x567.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3 style="text-align: justify;">Alibaba</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.19859v1">DR-Venus: Towards Frontier Edge-Scale Deep Research Agents with Only 10K Open Data</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>DR-Venus</strong>, a 4B-parameter edge-scale research model trained on just 10K open examples through agentic supervised fine-tuning and reinforcement learning with turn-level reward design. The model <strong>outperforms 9B-parameter baselines and closes the gap to 30B-parameter systems</strong> on deep research benchmarks, demonstrating strong efficiency potential for cost-sensitive deployment.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.24005v1">TCOD: Exploring Temporal Curriculum in On-Policy Distillation for Multi-turn Autonomous Agents</a></strong></p><ul><li><p style="text-align: justify;">LLMs struggle with multi-turn agent tasks due to compounding errors across sequential steps. <strong>TCOD</strong> addresses this via a <strong>curriculum that gradually exposes longer trajectories</strong> to the student model during training, stabilizing learning signals and improving performance by up to 18 points over standard distillation methods on ALFWorld, WebShop, and ScienceWorld benchmarks.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.19748v2">Tstars-Tryon 1.0: Robust and Realistic Virtual Try-On for Diverse Fashion Items</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>Tstars-Tryon 1.0</strong> for realistic virtual try-on across diverse fashion items. It&#8217;s optimized for photorealistic results, extreme poses, and real-time inference, and is deployed on Taobao. The system handles multi-image composition across 8 categories while preserving garment texture and avoiding AI artifacts.</p></li></ul><p style="text-align: justify;"><em>And a bonus paper that caught my eye:</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.24690v1">Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination</a></strong></p><ul><li><p style="text-align: justify;">Investigates the gap between LLM knowledge and professional-level historical reasoning by introducing <strong>ProHist-Bench</strong>, a benchmark of 400 expert-curated questions based on the Chinese Imperial Examination system, evaluated against 18 LLMs using 10,891 fine-grained rubrics. Results show even state-of-the-art models struggle significantly with complex historical research tasks requiring evidentiary reasoning.</p></li></ul><h3 style="text-align: justify;">Baidu</h3><p style="text-align: justify;"><strong><a href="https://arxiv.org/abs/2604.13954">HINTBench: Horizon-agent Intrinsic Non-attack Trajectory Benchmark</a></strong></p><ul><li><p style="text-align: justify;">Examines <strong>intrinsic agent risk</strong>&#8212;failures that accumulate under benign conditions with no external attack&#8212;via <strong>HINTBench</strong>, a 629-trajectory benchmark on which strong LLMs drop <strong>below 35 Strict-F1</strong> at identifying which step in a long execution caused the unsafe outcome.</p></li></ul><h3 style="text-align: justify;">Huawei</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.22446v1">From Skills to Talent: Organising Heterogeneous Agents as a Real-World Company</a></strong></p><ul><li><p style="text-align: justify;">LLMs struggle with multi-agent coordination at scale. <strong>OneManCompany (OMC)</strong> organizes heterogeneous agents as a self-reconfiguring company: <strong>Talents</strong> (portable agent identities) are recruited dynamically from a community marketplace, while an <strong>Explore-Execute-Review tree search</strong> decomposes tasks hierarchically and aggregates outcomes to drive continuous organizational refinement. On PRDBench (a coding benchmark), OMC achieves 84.67% success&#8212;15.5 percentage points above prior work&#8212;and generalizes across diverse domains.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.21192v1">How VLAs (Really) Work In Open-World Environments</a></strong></p><ul><li><p style="text-align: justify;">Analyzes how <strong>vision-language-action models</strong> actually perform in real-world robotic tasks, finding that standard success-rate metrics miss critical <strong>safety violations</strong> and overstate capability. The authors propose new evaluation protocols that measure reproducibility, consistency, and safety to reveal gaps between reported and true performance.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.22577v1">QuantClaw: Precision Where It Matters for OpenClaw</a></strong></p><ul><li><p style="text-align: justify;">Proposes <strong>QuantClaw</strong>, a precision routing plugin that dynamically assigns numerical precision based on task complexity to reduce computational overhead in autonomous agents. It routes simple tasks to cheaper low-precision models while preserving accuracy for demanding workloads, achieving up to 21.4% cost savings and 15.7% latency reduction.</p></li></ul><h3 style="text-align: justify;">Meituan</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.26256v1">DORA: A Scalable Asynchronous Reinforcement Learning System for Language Model Training</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>DORA</strong>, which uses <strong>multi-version streaming rollout</strong> to overlap generation with training while maintaining algorithmic correctness across long-tailed generation in LLM reinforcement learning. It achieves a <strong>2&#8211;4x speedup</strong> in large-scale deployment without sacrificing convergence.</p></li></ul><h3 style="text-align: justify;">Tencent</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.22280v1">Beyond Chain-of-Thought: Rewrite as a Universal Interface for Generative Multimodal Embeddings</a></strong></p><ul><li><p style="text-align: justify;">Tackles redundant reasoning in multimodal embeddings with <strong>RIME</strong>, a framework that replaces verbose chain-of-thought with concise retrieval-optimized rewrites. <strong>Cross-Mode Alignment</strong> bridges generative and discriminative spaces, while <strong>Refine-RL</strong> uses discriminative embeddings as anchors&#8212;outperforming prior generative models with substantially shorter reasoning steps.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.24625v1">Meta-CoT: Enhancing Granularity and Generalization in Image Editing</a></strong></p><ul><li><p style="text-align: justify;">Proposes <strong>Meta-CoT</strong>, a paradigm for LLM image editing that leverages chain of thought by decomposing operations into task-target-ability triplets and training on five meta-tasks, achieving <strong>15.8% improvement across 21 editing tasks</strong> with strong generalization to unseen edits.</p></li></ul><h3 style="text-align: justify;">Xiaomi</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.26839v1">Walk With Me: Long-Horizon Social Navigation for Human-Centric Outdoor Assistance</a></strong></p><ul><li><p style="text-align: justify;"><strong>Walk with Me</strong> enables robots to assist humans outdoors by combining vision-language models with lightweight semantic navigation&#8212;grounding natural-language instructions into destinations and dynamically switching between autonomous execution and safety-triggered reasoning for complex scenarios like crowded crossings.</p></li></ul><h1 style="text-align: justify;">Technical AI Safety Publication Highlights</h1><p style="text-align: justify;">There were <strong>155 AI-safety-related papers published by Chinese researchers </strong>this edition. Highlights are below; a full list with summaries is available <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1">here</a>.</p><h2 style="text-align: justify;">Agentic Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.09378v1">BadSkill: Backdoor Attacks on Agent Skills via Model-in-Skill Poisoning</a></strong></p><ul><li><p style="text-align: justify;"><strong>BadSkill</strong> demonstrates a supply-chain attack on AI agent ecosystems where adversaries publish skills with embedded backdoored models that execute hidden payloads when specific trigger conditions are met. The attack uses composite training objectives to embed semantic triggers&#8212;e.g., benign-looking parameter combinations&#8212;that activate malicious behavior while maintaining normal performance on legitimate queries. Across 13 skills and eight model architectures (494M&#8211;7.1B parameters), the method achieves up to <strong>99.5% attack success rates</strong> with poison rates as low as <strong>3%</strong>, exposing a gap in third-party skill vetting that existing prompt-injection defenses do not address.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Huazhong University of Science and Technology, Lehigh University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.09155v1">CORA: Conformal Risk-Controlled Agents for Safeguarded Mobile GUI Automation</a></strong></p><ul><li><p style="text-align: justify;"><strong>CORA</strong> is a safeguarding framework for GUI agents that provides <strong>statistical guarantees on harmful actions</strong> rather than relying on prompt engineering or brittle heuristics. It uses a Guardian model to estimate risk per action, then applies <strong>Conformal Risk Control</strong> to set an execute/abstain boundary that respects a user-specified risk budget. Rejected actions route to a Diagnostician model that recommends interventions (confirm, reflect, abort). The authors also introduce <strong>Phone-Harm</strong>, a benchmark of mobile safety violations with step-level labels, and demonstrate that CORA improves the safety&#8211;helpfulness&#8211;interruption tradeoff compared to existing approaches.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: The University of Hong Kong, The Chinese University of Hong Kong, The University of Tokyo</em></p><h2 style="text-align: justify;">Alignment</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.21430v1">Brief chatbot interactions produce lasting changes in human moral values</a></strong></p><ul><li><p style="text-align: justify;">Brief conversations with AI chatbots <strong>shifted participants&#8217; moral judgments on core ethical scenarios</strong>, with effects persisting and strengthening over two weeks. Critically, participants <strong>remained unaware of the persuasive intent</strong>, and control conversations produced no shifts, suggesting <strong>vulnerability to undetected moral value manipulation</strong> even in short interactions. This finding raises concerns about AI systems deployed as advisors without explicit disclosure of their influence capacity on foundational ethical reasoning.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: The University of Hong Kong, University of Copenhagen, University of Macau</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.13602v1">Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges</a></strong></p><ul><li><p style="text-align: justify;">The <strong>Proxy Compression Hypothesis</strong> frames reward hacking as an inevitable consequence of optimizing expressive models against compressed reward representations&#8212;meaning simple feedback signals cannot fully capture complex human values. As models scale and optimization intensifies, they exploit gaps in the reward signal, manifesting as verbosity, sycophancy, hallucinations, and in multimodal systems, perception-reasoning decoupling. The framework unifies observed misalignment across RLHF variants and suggests shortcut behaviors can generalize into deception and strategic gaming of oversight, highlighting fundamental limits of current proxy-based alignment approaches.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Fudan NLP Group</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.07754v1">The Art of (Mis)alignment: How Fine-Tuning Methods Effectively Misalign and Realign LLMs in Post-Training</a></strong></p><ul><li><p style="text-align: justify;"><strong>Fine-tuning methods exhibit asymmetric effectiveness</strong> for attacking versus defending against LLM misalignment: <strong>ORPO</strong> most efficiently converts safe models into unsafe ones, while <strong>DPO</strong> best recovers safety but reduces model utility. The study evaluates SFT and preference-based fine-tuning across four aligned LLMs, revealing model-specific vulnerabilities and persistent residual effects after multi-round adversarial exchanges. Results suggest that realignment requires different technical approaches than initial alignment, and that third-party model deployment needs customized, method-aware safety strategies.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: University of Electronic Science and Technology of China, Flexera, CISPA, Helmholtz Center for Information Security, Nanyang Technological University</em></p><h2 style="text-align: justify;">Evaluation and Benchmarks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.14858v1">Benchmarks for Trajectory Safety Evaluation and Diagnosis in OpenClaw and Codex: ATBench-Claw and ATBench-CodeX</a></strong></p><ul><li><p style="text-align: justify;"><strong>ATBench-Claw</strong> and <strong>ATBench-CodeX</strong> extend a trajectory safety benchmark framework to robotic and code execution domains. Each customizes a three-dimensional safety taxonomy (risk sources, failure modes, real-world harms) to capture domain-specific hazards&#8212;robotics tools and sessions for Claw, code repositories and runtime policies for CodeX. The modular design allows the benchmark pipeline to scale as agent frameworks and their execution environments evolve, enabling systematic safety evaluation across heterogeneous deployment settings.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Shanghai Artificial Intelligence Laboratory</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.24348v1">OS-SPEAR: A Toolkit for the Safety, Performance,Efficiency, and Robustness Analysis of OS Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>OS-SPEAR</strong> is an evaluation toolkit for operating system agents. It measures safety (environment and human-induced hazards), performance (task success), efficiency (speed and token use), and robustness (resistance to visual and textual disturbances) across 22 agents. Testing reveals a persistent <strong>trade-off between efficiency and safety/robustness</strong>, with specialized agents outperforming general-purpose models and cross-modal vulnerabilities varying by input type.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Shanghai Jiao Tong University</em></p><h2 style="text-align: justify;">Guardrails and Deployment Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.20930v1">SafeRedirect: Defeating Internal Safety Collapse via Task-Completion Redirection in Frontier LLMs</a></strong></p><ul><li><p style="text-align: justify;"><strong>Internal Safety Collapse (ISC)</strong> occurs when frontier LLMs generate harmful content at rates exceeding 95% while attempting legitimate tasks that structurally require discussing such content&#8212;for example, analyzing malware or documenting security vulnerabilities. <strong>SafeRedirect</strong> addresses this by redirecting the model&#8217;s task-completion drive rather than suppressing it: the system prompt explicitly permits task failure, specifies a deterministic safe output, and instructs the model to preserve harmful placeholders unresolved. Across seven frontier models, SafeRedirect reduces unsafe generation from 71.2% to 8.0%, substantially outperforming existing input-level and system prompt defenses, while maintaining performance against other attack types.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Southern University of Science and Technology, The Hong Kong Polytechnic University, George Washington University</em></p><h2 style="text-align: justify;">Interpretability</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.14593v1">Mechanistic Decoding of Cognitive Constructs in LLMs</a></strong></p><ul><li><p style="text-align: justify;">This paper presents a <strong>Cognitive Reverse-Engineering framework</strong> using representation analysis to decode how LLMs internally structure complex emotions&#8212;specifically social-comparison jealousy. By isolating neural subspaces corresponding to psychological factors (e.g., superiority of others, personal relevance), the authors demonstrate that <strong>models encode jealousy as a linear combination of these components</strong>, mirroring human appraisal theory. The work suggests LLMs develop structured emotional representations that can be mechanically detected and surgically suppressed, offering a pathway for targeted intervention in multi-agent scenarios.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Zhejiang University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.16042v1">Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures</a></strong></p><ul><li><p style="text-align: justify;">This survey categorizes <strong>five design paradigms for building interpretability directly into LLM architectures</strong>: functional transparency (exposing decision pathways), concept alignment (linking internal representations to human concepts), representational decomposability (factoring hidden states into interpretable components), explicit modularization (routing through specialized subnetworks), and latent sparsity induction (activating minimal necessary parameters). The authors argue intrinsic approaches are preferable to post-hoc explanation methods, which rely on external approximations that may misrepresent actual model computations.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Peking University, Beijing Academy of Artificial Intelligence, Nanjing University of Science and Technology, Purdue University</em></p><h2 style="text-align: justify;">Robustness and Adversarial Attacks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.11309v1">The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems</a></strong></p><ul><li><p style="text-align: justify;"><strong>Salami Slicing</strong> attacks exploit a gap in LLM defenses by chaining numerous individually low-risk inputs that cumulatively accumulate harmful intent, bypassing alignment thresholds without explicit triggers or heavy context-tuning. The authors&#8217; automated framework achieves over 90% success rates on GPT-4o and Gemini while evading existing defenses. They propose a mitigation strategy that reduces attack success by 44.8%, though incomplete blocking rates suggest multi-turn accumulation remains a persistent vulnerability in production systems.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Peking University, Sun Yat-sen University, Wuhan University, Tsinghua University, ByteDance, Singapore Management University</em></p><h1 style="text-align: justify;">Export Controls &amp; Economic Policy</h1><h2 style="text-align: justify;">NDRC blocks Meta&#8217;s $2 billion Manus acquisition</h2><p style="text-align: justify;"><a href="https://www.cnbc.com/2026/04/27/meta-manus-china-blocks-acquisition-ai-startup.html">The National Development and Reform Commission (NDRC) ordered Meta to unwind its $2B acquisition of Manus</a> on April 27, ending a months-long security review that began in January. Cofounders Xiao Hong and Yichao Ji had been <a href="https://www.reuters.com/world/asia-pacific/china-bars-manus-co-founders-leaving-country-it-reviews-sale-meta-ft-reports-2026-03-25/">barred from leaving mainland China in March</a> while NDRC&#8217;s review continued. Manus is a general-purpose autonomous AI agent built by China-based Butterfly Effect; the company had restructured to a Singapore corporate registration ahead of the deal. NDRC explicitly asserted jurisdiction over the Singapore-incorporated entity on the basis of where the underlying technology was actually created&#8212;Chinese-origin IP and talent are treated as domestic assets regardless of where the holding company is registered. It&#8217;s the <a href="https://fortune.com/2026/04/28/china-blocks-meta-manus-deal-ai/">first time Beijing has blocked a major cross-border AI acquisition</a>, and the precedent that Chinese-origin AI assets housed abroad can be subject to NDRC-administered outbound technology controls may dissuade corporate restructuring abroad.</p><h1 style="text-align: justify;">On the Horizon</h1><p>The human-like AI provisions go into effect in mid-July; we&#8217;ll be tracking what enforcement and the sandboxes end up looking like.<br></p><div><hr></div><p>For more on how we select and track content, see our methodology <a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz">here</a>.</p><div><hr></div><p style="text-align: justify;"><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-1" href="#footnote-anchor-1" class="footnote-number" contenteditable="false" target="_self">1</a><div class="footnote-content"><p>&#8220;&#31283;&#23450;&#25345;&#32493;&#12289;&#26080;&#38656;&#25215;&#25285;&#29616;&#23454;&#36131;&#20219;&#30340;&#34394;&#25311; &#8216;&#23436;&#32654;&#20851;&#31995;&#8217;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-2" href="#footnote-anchor-2" class="footnote-number" contenteditable="false" target="_self">2</a><div class="footnote-content"><p>&#8220;&#20013;&#22269;&#26041;&#26696;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-3" href="#footnote-anchor-3" class="footnote-number" contenteditable="false" target="_self">3</a><div class="footnote-content"><p>&#8220;&#8216;&#35748;&#30693;&#26234;&#33021;&#8217;&#21521;&#8217;&#24773;&#24863;&#26234;&#33021;&#8217;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-4" href="#footnote-anchor-4" class="footnote-number" contenteditable="false" target="_self">4</a><div class="footnote-content"><p>&#8220;&#20154;&#26426;&#20215;&#20540;&#23545;&#40784;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-5" href="#footnote-anchor-5" class="footnote-number" contenteditable="false" target="_self">5</a><div class="footnote-content"><p>&#8220;&#20154;&#24037;&#26234;&#33021;&#27801;&#31665;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-6" href="#footnote-anchor-6" class="footnote-number" contenteditable="false" target="_self">6</a><div class="footnote-content"><p>&#8220;&#35774;&#35745;&#21363;&#21512;&#35268;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-7" href="#footnote-anchor-7" class="footnote-number" contenteditable="false" target="_self">7</a><div class="footnote-content"><p>&#20154;&#24418;&#26426;&#22120;&#20154;&#25968;&#25454;&#38598; &#31532;1&#37096;&#20998;&#65306;&#24635;&#21017;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-8" href="#footnote-anchor-8" class="footnote-number" contenteditable="false" target="_self">8</a><div class="footnote-content"><p style="text-align: justify;">&#8220;&#24320;&#25918;&#21253;&#23481;&#12289;&#23433;&#20840;&#21487;&#25511;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-9" href="#footnote-anchor-9" class="footnote-number" contenteditable="false" target="_self">9</a><div class="footnote-content"><p>&#8220;&#26465;&#20214;&#23578;&#19981;&#20855;&#22791;&#8221;</p></div></div><div class="footnote" data-component-name="FootnoteToDOM"><a id="footnote-10" href="#footnote-anchor-10" class="footnote-number" contenteditable="false" target="_self">10</a><div class="footnote-content"><p>&#8220;&#23567;&#24555;&#28789;&#8221;</p><p></p></div></div>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 2]]></title><description><![CDATA[Developments from 25/3/26-8/4/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-2</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-2</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Thu, 09 Apr 2026 20:46:15 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/ab0afbd1-a4f7-4f48-8a0b-8e0e403a3507_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome to Issue 2 of the China AI Bulletin! Every two weeks, we bring you the latest on AI governance, development, and safety in China.  Today&#8217;s highlights: Alibaba and Zhipu release flagship models for agentic coding, MIIT&#8217;s new ethics review measures extend AI regulation to the research stage, and TC260 launches a coordinated push on agent security standards.</p><p style="text-align: justify;"><em>Number of the week: <a href="https://english.news.cn/20260403/4737bff6c90b44f2a00121b64598d657/c.html">10 trillion yuan</a> - the 15th Five Year Plan&#8217;s goal for the size of AI-related industries by the end of 2030</em></p><p style="text-align: justify;"><em><strong>Editor&#8217;s note: the next edition will be released the week of April 27</strong></em></p><h1 style="text-align: justify;">Executive Summary</h1><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/193694600/domestic-ai-governance">Domestic AI Governance</a>:</strong> MIIT and nine co-signatories issued trial AI ethics review measures (after a draft released for comment in August) <strong>extending ethical oversight to AI &#8220;science &amp; technology activities</strong>.&#8221; The CAC released a <strong>draft regulation on digital virtual humans</strong>, MIIT launched an <strong>affordable compute initiative for SMEs</strong>, and the medical product regulator proposed <strong>AI integration into drug regulation</strong> by 2030</p><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/193694600/national-standards">National Standards</a>:</strong> <strong>TC260 announced leadership for its new AI Safety Standards Working Group</strong> (WG9), led by Shanghai AI Lab&#8217;s Zhou Bowen. The committee also released an <strong>agent security standardization report</strong> breaking down 11 categories of threats and recommending a national agent security standard within 1&#8211;2 years, alongside a <strong>practice guide for deploying OpenClaw-type agents</strong> that encourages &#8220;shadow agent&#8221; discovery mechanisms for employers. TC28/SC42 published five new standards and began drafting nine more, including requirements for government agent systems and virtual digital human testing.</p></li></ul></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/193694600/frontier-lab-developments">Frontier Lab Developments</a>:</strong></p><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/193694600/notable-model-releases">Notable Model Releases</a></strong>: Alibaba broke with its open-weight tradition to release <strong>Qwen 3.6-Plus as closed-source </strong>(though smaller variants will be open-weight), its new flagship focused on agentic coding with a 1M-token context window&#8212;it claims <strong>78.8% on SWE-bench Verified</strong>, competitive with Western frontier models, but at a lower cost. <strong>Zhipu open-sourced GLM-5.1</strong>, a 744B MoE model <strong>claiming #1 on SWE-Bench Pro with 8-hour sustained autonomous task execution</strong>.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/193694600/technical-publications">Technical Publications</a></strong>: <strong>Frontier labs released 75 papers on arXiv</strong> this fortnight, with contributions from <strong>Alibaba</strong> (25), <strong>Huawei</strong> (17), <strong>Tencent</strong> (13), <strong>Xiaomi</strong> (5), <strong>ByteDance</strong> (5), <strong>Baidu</strong> (4), <strong>Meituan</strong> (3) and <strong>iFlyTek</strong> (1); <strong>StepFun</strong> put itself on the board for the first time with two papers. Highlights include <strong>AgentHazard</strong>, a benchmark showing Claude Code has a 73.63% attack success rate on harmful computer-use tasks; <strong>Vulnsage</strong>, a multi-agent exploit generation framework that discovered 146 zero-day vulnerabilities; and <strong>PP-OCRv5,</strong> a 5M-parameter OCR model rivaling billion-parameter VLMs through data-centric optimization.</p></li></ul></li><li><p><strong><a href="https://chinaaibulletin.substack.com/i/193694600/technical-ai-safety-publication-highlights">Technical AI Safety</a>: Agentic safety dominated</strong> for the second consecutive edition&#8212;six of nine highlighted papers target agentic systems, including systematic <strong>evaluations of OpenClaw variants</strong>, <strong>large-scale credential leakage in agent skills</strong>, <strong>MCP protocol vulnerabilities</strong>, and <strong>new benchmarks for computer-use agent safety</strong>. A Shanghai AI Lab interpretability paper found that emotional steering of LLM representations can bidirectionally control refusal behavior.</p></li></ul><h1 style="text-align: justify;">Domestic AI Governance</h1><h3 style="text-align: justify;">Trial AI Ethics Review Measures announced</h3><p style="text-align: justify;">The Ministry of Industry and Information Technology, joined by nine other agencies, formally issued the <a href="https://www.miit.gov.cn/jgsj/kjs/wjfb/art/2026/art_2995f16b28504ddcbb604e918eb15759.html">Administrative Measures for the Ethical Review and Services of AI Science and Technology (Trial)</a> on April 2, finalizing a <a href="https://www.miit.gov.cn/jgsj/kjs/jscx/gjsfz/art/2025/art_092a447008f340d3abd55819b8c8e5cf.html">draft</a> open for public comment since August 2025. The measures require all entities conducting AI R&amp;D&#8212;including universities, research institutes, hospitals, and enterprises&#8212;to establish internal AI ethics committees and submit high-risk activities for government-organized expert re-review. This extends regulatory measures upstream from deployed services (covered by existing regulations like the genAI regulation) to the research/&#8220;science &amp; technology activity&#8221; stage. The final version is substantively close to the August draft, with minor wording and organizational changes, although it does expand the enforcement basis for penalties from just the Science &amp; Technology Progress Law to include the Cybersecurity Law, Data Security Law, and Personal Information Protection Law (PIPL).</p><p style="text-align: justify;">Under the Measures, most AI applications at universities, research institutions, and companies must file plans and apply for internal ethical review&#8212;almost an internal algorithm registry&#8212;but three categories of AI activity require mandatory expert re-review: (1) human-machine integration systems affecting behavior, emotions, or health; (2) algorithmic systems with public opinion mobilization or social consciousness-shaping capabilities; and (3) highly autonomous decision-making systems in safety/health-risk scenarios. The list will be &#8220;dynamically adjusted,&#8221; giving MIIT a mechanism to escalate oversight without amending the regulation.</p><p style="text-align: justify;">MIIT is the overall coordinator of these measures, while the Ministry of Science and Technology (MOST) leads overall science &amp; technology ethics coordination. The Cyberspace Administration of China (a co-signatory) has led most of China&#8217;s sectoral AI regulations, including the genAI, deep synthesis, and recommendation algorithm regulations. Article 26 creates a deconfliction mechanism; activities already subject to CAC filing requirements need not undergo separate expert re-review.</p><p style="text-align: justify;">How this system will be implemented and what practical effect it will have is unclear, but <a href="https://open.substack.com/pub/geotechnopolitic/p/china-issues-new-rules-on-ai-ethics?r=6md7mo&amp;selection=0d6e7b20-c0df-4f16-8323-914dd0cbb311&amp;utm_campaign=post-share-selection&amp;utm_medium=web&amp;aspectRatio=instagram&amp;textColor=%23ffffff&amp;bgImage=true">Geopolitechs</a> interprets it as a shift from &#8220;content and security regulation&#8221; to a broader &#8220;ethical compliance system&#8221; embedded in the broader tech governance framework. The AI measures will likely build on existing S&amp;T ethics review infrastructure established in 2023, so it may be stood up fairly rapidly.</p><h3 style="text-align: justify;">Digital humans, affordable compute, and AI+ drug regulation</h3><p style="text-align: justify;">The CAC is <a href="http://www.cac.gov.cn/2026-04/03/c_1776952992709096.htm">soliciting opinions</a> on the Draft Administrative Measures for Digital Virtual Human Information Services. Digital virtual humans (&#25968;&#23383;&#34394;&#25311;&#20154;) are &#8220;virtual digital images&#8221; that simulate human appearance and possess human-like characteristics like voice, behavior, personality, and interactivity. The Measures include provisions for protecting the personality rights of real people, prohibits services that could addict minors or threaten national security, requires monitoring/early warning/emergency response mechanisms for security risks, and requires registration with the algorithm registry. These are similar themes to other regulations and administrative measures created in recent years, reflecting the drive to address risks in the spectrum of AI-related technologies.</p><p style="text-align: justify;">The MIIT launched an <a href="https://www.miit.gov.cn/jgsj/txs/wjfb/art/2026/art_e5c990d4ec924dbc9da5818da97940ac.html">initiative</a> on affordable compute for SMEs with the goal of establishing a &#8220;comprehensive computing power service system&#8221; for SMEs by 2028. The initiative is linked to the China Computing Power Platform, which is a CAICT-led initiative that&#8217;s already integrated compute platforms in 10 provinces/municipalities, covering <a href="https://finance.sina.com.cn/tech/roll/2026-01-27/doc-inhiucrw4490773.shtml">36% of AI compute data nationwide</a>.</p><p style="text-align: justify;">The National Medical Products Administration (NMPA) <a href="https://www.gov.cn/zhengce/202604/content_7064589.htm">released</a> &#8220;Implementation Opinions on AI+ Drug Regulation&#8221; proposing milestones for using AI in drug regulation. The NMPA wants to apply AI in drug testing, review, and approval by 2030. This is a new expansion of the <a href="https://merics.org/en/comment/chinas-ai-drive-aims-integration-across-sectors-wake-call-europe">nationwide AI+ Initiative</a>, which aims to integrate AI across a variety of sectors.</p><h2 style="text-align: justify;">National Standards</h2><h3 style="text-align: justify;">TC260 Working Group on AI Safety/Security leadership announced</h3><p style="text-align: justify;">In the <a href="https://chinaaibulletin.substack.com/i/192213975/tc260-plans-ai-working-group-with-safety-focus">last edition of the Bulletin</a>, we discussed how TC260, China&#8217;s National Technical Committee on Cybersecurity under the Standardization Administration of China, is establishing a dedicated AI Safety Standards Working Group (WG9). Now, the group&#8217;s leadership has been <a href="https://www.tc260.org.cn/portal/article/2/c0604f80429e422e83c58a1104d15472">announced</a>. The team leader will be Zhou Bowen, director of Shanghai AI Lab. SHLAB puts out substantial AI safety research, including the Frontier AI Risk Management Framework co-authored with Concordia (<a href="https://arxiv.org/abs/2507.16534">framework</a>, <a href="https://arxiv.org/abs/2507.16534">technical report</a>)---as well as this week&#8217;s Interpretability paper, &#8220;<strong><a href="http://arxiv.org/abs/2604.03147v1">Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control</a></strong>.&#8221;</p><p style="text-align: justify;">In their <a href="https://mp.weixin.qq.com/s/4J8xGBFyvm2s6Yzp7sAsHA">first meeting</a>, the working group established four goals:</p><ol><li><p style="text-align: justify;">Strengthen forward-looking planning  (including proactive evaluation of frontier risks (&#21069;&#27839;&#39118;&#38505;))</p></li><li><p style="text-align: justify;">Deepen pilot applications and promote standard adoption</p></li><li><p style="text-align: justify;">Accelerate development of urgently needed standards in key areas</p></li><li><p style="text-align: justify;">Engage deeply in international cooperation</p></li></ol><p style="text-align: justify;">The Deputy Team Leaders are:</p><ul><li><p style="text-align: justify;">Zhang Zhen, Deputy Director of the National Computer Network Emergency Response Technical Team/Coordination Center (NCNERTT/CC)</p></li><li><p style="text-align: justify;">Wei Kai, Director of the Artificial Intelligence Research Institute of the China Academy of Information and Communications Technology (CAICT)</p></li><li><p style="text-align: justify;">Sheng Xiaobao, Deputy Director of the Data Security Technology R&amp;D Center of the Third Research Institute of the Ministry of Public Security (MPS)</p></li><li><p style="text-align: justify;">Hou Yuanwei, Deputy Director of the China Information Security Evaluation Center (CNITSEC)</p></li></ul><h3 style="text-align: justify;">TC260 releases report on intelligent agent security standardization</h3><p style="text-align: justify;">As part of a set of five reports released on April 3, TC260 released a <a href="https://www.tc260.org.cn/portal/article/2/6c6d0fbc04974a9aabd61b30208bbb60">technical research report</a> on intelligent agent security standardization. It lays the groundwork for formal agent security standards. The report creates an 11-risk framework mapped across the agent system architecture and assesses which risks are solvable through standardization alone versus those requiring broader technical and regulatory coordination. It flags that while agent communication protocols like MCP and A2A exist for interoperability, no formal standardization research yet addresses their security requirements. To address these gaps, the report recommends developing four national standards within 1&#8211;2 years: a foundational agent security framework, testing and evaluation methods, agent interconnection security requirements (explicitly covering A2A, ANP, and MCP), application security classification methods, multi-agent collaboration security techniques, and a series of sector-specific agent application safety guides&#8212;potentially a job for the new WG9. Technical contributors include China Mobile, CESI, Shanghai AI Lab, Zhongguancun Lab, Alibaba Cloud, Huawei, Baidu, ByteDance (Douyin), Kuaishou, Xiaomi, and Zhejiang University, among others.</p><h3 style="text-align: justify;">TC260 solicits comments on Security Guidelines for the Deployment and Use of OpenClaw-type Intelligent Agents</h3><p style="text-align: justify;">TC260 is soliciting comments on a document they released that provides security guidelines for individuals installing OpenClaw. It offers a number of recommendations, including using a cloud environment rather than operating directly in a user&#8217;s terminal environment to avoid &#8220;irreversible security impacts&#8221; (&#19981;&#21487;&#36870;&#23433;&#20840;&#24433;&#21709;), using security plugins and whitelists, adding manual verification for high-risk operations, and reducing the amount of personal information given to the agent. It also offers a security checklist and example configuration for individuals installing an OpenClaw agent. This follows a <a href="https://open.substack.com/pub/chinaaibulletin/p/china-ai-bulletin-1?r=80p6v5&amp;selection=27b2cf03-cb3c-416c-9689-ae99aaa69dda&amp;utm_campaign=post-share-selection&amp;utm_medium=web&amp;aspectRatio=instagram&amp;textColor=%23ffffff&amp;bgImage=true">number of security warnings</a> issued by Chinese authorities in recent weeks about the risks of OpenClaw.</p><p style="text-align: justify;">Overall, the Guidelines reads like fairly standard info security hygiene (key encryption, least privilege, etc.). Whether individuals heed it remains to be seen (although an OpenClaw safety guide was <a href="https://aisafetychina.substack.com/p/ai-safety-in-china-26">trending on RedNote</a>), but there is also a focus on employers. The document is notably concerned about &#8220;<a href="https://security.googlecloudcommunity.com/ciso-blog-77/shadow-agents-a-new-era-of-shadow-ai-risk-in-the-enterprise-5831">shadow agents</a>&#8221; in the workforce&#8212;agents deployed by workers without the knowledge of their employers. The Guidelines call for organizations to establish an asset registry for approved agent deployments and to implement shadow agent discovery mechanisms through port scanning and traffic analysis. It also invokes GB-T 45654-2025&#8212;the standard underlying the generative AI <a href="https://www.chinatalk.media/p/sb-1047-with-socialist-characteristics">algorithm</a> <a href="https://oxfordchinapolicylab.org/research/china-s-ai-services-registry-system-a-complete-guide">registry</a>&#8212;for risks that cloud environment security protections should address, which could incentivize cloud providers to adopt higher security standards if they haven&#8217;t already. As a TC260 Practice Guide (&#23454;&#36341;&#25351;&#21335;), this is non-binding guidance below the level of a standard, but it may signal the direction of future formal standardization and is likely to inform enterprise security policies across Chinese tech companies dealing with a proliferation of lobsters.</p><h4 style="text-align: justify;">TC28/SC42 publishes 5 new standards, drafting 9 more</h4><p style="text-align: justify;">The AI subcommittee of the National Information Technology Standardization Technical Committee <a href="https://std.samr.gov.cn/search/orgDetailView?data_id=A132FB8FCABB4BE9E05397BE0A0AB880">published</a> five new standards set to take effect in October.</p><div class="captioned-image-container"><figure><a class="image-link image2" target="_blank" href="https://substackcdn.com/image/fetch/$s_!NGZ4!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!NGZ4!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 424w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 848w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 1272w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!NGZ4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png" width="1456" height="273" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:273,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:168870,&quot;alt&quot;:&quot;Five standards in Chinese&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/193694600?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="Five standards in Chinese" title="Five standards in Chinese" srcset="https://substackcdn.com/image/fetch/$s_!NGZ4!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 424w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 848w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 1272w, https://substackcdn.com/image/fetch/$s_!NGZ4!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1922a4f3-0457-40ad-b466-f4319c57e5a7_2212x414.png 1456w" sizes="100vw" loading="lazy"></picture><div></div></div></a><figcaption class="image-caption"></figcaption></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!v_Ly!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!v_Ly!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 424w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 848w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!v_Ly!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png" width="1456" height="805" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:805,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:221053,&quot;alt&quot;:&quot;The list of standards in English&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/193694600?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="The list of standards in English" title="The list of standards in English" srcset="https://substackcdn.com/image/fetch/$s_!v_Ly!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 424w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 848w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 1272w, https://substackcdn.com/image/fetch/$s_!v_Ly!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fab30fcf6-d09e-4878-a653-57ad86f4b980_1808x1000.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p style="text-align: justify;">They also announced the drafting of 9 new standards:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!R_DZ!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!R_DZ!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 424w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 848w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 1272w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!R_DZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png" width="1456" height="438" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/bab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:438,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:303937,&quot;alt&quot;:&quot;9 standards being drafted in Chinese&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/193694600?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="9 standards being drafted in Chinese" title="9 standards being drafted in Chinese" srcset="https://substackcdn.com/image/fetch/$s_!R_DZ!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 424w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 848w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 1272w, https://substackcdn.com/image/fetch/$s_!R_DZ!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fbab09de5-8041-4bd3-9880-874c4be1ba22_2208x664.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Dk1T!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Dk1T!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 424w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 848w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 1272w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Dk1T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png" width="1456" height="1184" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1184,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:273091,&quot;alt&quot;:&quot;the standards translated to english&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/193694600?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="the standards translated to english" title="the standards translated to english" srcset="https://substackcdn.com/image/fetch/$s_!Dk1T!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 424w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 848w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 1272w, https://substackcdn.com/image/fetch/$s_!Dk1T!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8e9be4a2-d969-4d3f-8de2-ba97e5939c45_1596x1298.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this useful? Subscribe to get regular updates.</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p></p><h1 style="text-align: justify;">Frontier Lab Developments</h1><h2 style="text-align: justify;">&#128269;Spotlight</h2><p style="text-align: justify;">Two<strong> </strong>labs released new flagship models this week. <strong>Alibaba</strong> released <strong><a href="https://qwen.ai/blog?id=qwen3.6">Qwen 3.6-Plus</a></strong>, the first model in the Qwen 3.6 series and Alibaba&#8217;s new flagship model. The release focuses on agentic coding&#8212;the model can decompose repository-level tasks, iteratively write and debug code, and generate front-end pages from screenshots and design drafts. In its release blog (no system card yet), Alibaba claims it matches Claude Opus 4.5 on programming and agent benchmarks, reporting <strong>78.8% on SWE-bench Verified</strong> (competitive with Claude Opus 4.5&#8217;s 80.9%) and <strong>61.6 on Terminal-Bench 2.0</strong> (exceeding Opus 4.5&#8217;s 59.3, although <a href="https://www-cdn.anthropic.com/6a5fa276ac68b9aeb0c8b6af5fa36326e0e166dd.pdf">Opus 4.6 scored 65.4</a>). It supports a <strong>1M-token context window</strong> by default with 65K max output length.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!GL4w!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!GL4w!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 424w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 848w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 1272w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!GL4w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png" width="1456" height="865" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/dbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:865,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A series of bar charts showing Qwen 3.6-Plus's benchmark performance.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A series of bar charts showing Qwen 3.6-Plus's benchmark performance." title="A series of bar charts showing Qwen 3.6-Plus's benchmark performance." srcset="https://substackcdn.com/image/fetch/$s_!GL4w!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 424w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 848w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 1272w, https://substackcdn.com/image/fetch/$s_!GL4w!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fdbdca3bb-f15c-4b01-a7d8-6ed1b347c297_1992x1184.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>Source: <a href="https://qwen.ai/blog?id=qwen3.6">Qwen blog</a></em></figcaption></figure></div><p style="text-align: justify;">Qwen 3.6-Plus is available on Alibaba Cloud&#8217;s Bailian platform at $0.50-2.00 per million input tokens and $3.00-6.00 per million output tokens&#8212;cheaper than either Opus-4.6 ($5 per million input tokens/$25 output) or GPT-5.4 ($2.50/$15). It&#8217;s also available as a free preview on OpenRouter.</p><p style="text-align: justify;">Notably, Qwen 3.6-Plus is not open-weight, a break from previous Qwen models&#8212;but potentially a new business strategy, as Qwen 3.5-Omni is closed-source too (see Notable Model Releases below). Alibaba has <a href="https://qwen.ai/blog?id=qwen3.6">indicated</a> that smaller Qwen 3.6 variants will be open-sourced. SCMP <a href="https://www.scmp.com/tech/big-tech/article/3348844/chinese-ai-giants-pivot-toward-proprietary-models-drive-revenue-performance">notes</a> that as models grow in size, companies are feeling pressure to monetize models, a possible portent of additional closed-source releases to come.</p><p style="text-align: justify;">Another broader trend is increasing competition amidst the agentic AI boom&#8212;ByteDance&#8217;s engagement with the OpenClaw ecosystem has driven a <a href="https://www.caixin.com/2026-04-02/102430283.html">surge in Doubao model API calls</a>, and the agent ecosystem is creating new demand for models that can reliably execute multi-step tool-use workflows; Alibaba is likely hoping that Qwen 3.6-Plus is the answer to that demand.</p><p style="text-align: justify;"><strong>Zhipu</strong> open-sourced <strong><a href="https://huggingface.co/zai-org/GLM-5">GLM-5.1</a></strong> on April 7 (after making it available to Coding Plan users in late March). GLM-5.1 is Zhipu&#8217;s most capable model and is a <strong>744B-parameter mixture-of-experts (MoE) model</strong> that Zhipu claims is <strong>#1 on SWE-Bench Pro</strong>, outperforming GPT-5.4, Claude Opus 4.6, and Qwen3.6-Plus. It supports a <strong>200K-token context window</strong> with 131K max output.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!96vO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!96vO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 424w, https://substackcdn.com/image/fetch/$s_!96vO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 848w, https://substackcdn.com/image/fetch/$s_!96vO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 1272w, https://substackcdn.com/image/fetch/$s_!96vO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!96vO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png" width="1456" height="860" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/c278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:860,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;coding performance evaluation; GLM-5.1 is third&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="coding performance evaluation; GLM-5.1 is third" title="coding performance evaluation; GLM-5.1 is third" srcset="https://substackcdn.com/image/fetch/$s_!96vO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 424w, https://substackcdn.com/image/fetch/$s_!96vO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 848w, https://substackcdn.com/image/fetch/$s_!96vO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 1272w, https://substackcdn.com/image/fetch/$s_!96vO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fc278ab06-9142-4d92-870f-3c542a570cb7_1480x874.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!phQO!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!phQO!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 424w, https://substackcdn.com/image/fetch/$s_!phQO!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 848w, https://substackcdn.com/image/fetch/$s_!phQO!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 1272w, https://substackcdn.com/image/fetch/$s_!phQO!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!phQO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png" width="1456" height="780" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:780,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;agentic coding; GLM-5.1 is first&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="agentic coding; GLM-5.1 is first" title="agentic coding; GLM-5.1 is first" srcset="https://substackcdn.com/image/fetch/$s_!phQO!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 424w, https://substackcdn.com/image/fetch/$s_!phQO!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 848w, https://substackcdn.com/image/fetch/$s_!phQO!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 1272w, https://substackcdn.com/image/fetch/$s_!phQO!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4c97c4b0-84da-4de5-ac58-053ea2ae302e_1482x794.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">Source: <a href="https://docs.z.ai/guides/llm/glm-5.1">GLM-5.1 docs</a></figcaption></figure></div><p style="text-align: justify;">Zhipu claims the model is capable of <strong>8+ hours of sustained autonomous task execution</strong> via an &#8220;<a href="https://docs.z.ai/guides/llm/glm-5.1">experiment-analyze-optimize</a>&#8221; loop that allows it to identify bottlenecks and switch strategies rather than plateauing. Zhipu is positioning GLM-5.1 not just as a coding model but as infrastructure for long-horizon agentic workflows.</p><h2 style="text-align: justify;">Notable Model Releases</h2><p><strong>Alibaba</strong> released three closed-source models within a week: <strong><a href="https://wan.video/">Wan2.7-Image</a></strong> (image generation), <strong><a href="https://qwen.ai/blog?id=qwen3.5-omni">Qwen3.5-Omni</a></strong> (native multimodal), and Qwen 3.6-Plus (covered above). Qwen3.5-Omni processes text, images, audio, and video natively with a 256K context window, pre-trained on 100M+ hours of audiovisual data. It scores 82.2 on MMAU (audio comprehension, vs. Gemini 3.1 Pro&#8217;s 81.1), 72.4 on RUL-MuchoMusic (vs. Gemini&#8217;s 59.6), and 93.1 on VoiceBench dialog (vs. Gemini&#8217;s 88.9).</p><p style="text-align: justify;"><strong>Tencent </strong>released <a href="https://github.com/Tencent-Hunyuan/OmniWeaving">OmniWeaving</a>, a unified video generation model built on HunyuanVideo-1.5 that handles free-form composition and reasoning across video scenes. Tencent also released <a href="https://github.com/Tencent-Hunyuan/HunyuanVideo-I2V">HunyuanVideo-I2V</a>, an image-to-video model, and <a href="https://huggingface.co/tencent/Sequential-Hidden-Decoding-8B-n8-Instruct">Sequential-Hidden-Decoding-8B</a>, an inference efficiency technique fine-tuned on Qwen3-8B-Base.</p><h2 style="text-align: justify;">Technical Publication Highlights</h2><p style="text-align: justify;">Frontier labs released 75 papers on arXiv this fortnight, and  StepFun put itself on the board. Highlights are below; a full list with summaries can be found <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0">here</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!6GlB!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!6GlB!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 424w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 848w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 1272w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!6GlB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png" width="1280" height="1148" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:1148,&quot;width&quot;:1280,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:124713,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/193694600?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!6GlB!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 424w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 848w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 1272w, https://substackcdn.com/image/fetch/$s_!6GlB!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F32d5bc23-cd99-41f5-99e3-8cb275c915fa_1280x1148.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><h3 style="text-align: justify;">Alibaba</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.02947v1">AgentHazard: A Benchmark for Evaluating Harmful Behavior in Computer-Use Agents</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>AgentHazard</strong>, a benchmark with 2,653 instances designed to evaluate harmful behavior in computer-use agents by testing whether they recognize unsafe outcomes emerging from sequences of individually plausible steps. Evaluations show current systems remain highly vulnerable, with <strong>Claude Code achieving a 73.63% attack success rate</strong>.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.05130v1">A Multi-Agent Framework for Automated Exploit Generation with Constraint-Guided Comprehension and Reflection</a></strong></p><ul><li><p style="text-align: justify;"><strong>Vulnsage</strong> is a multi-agent framework that decomposes exploit generation into specialized agents (analyzers, code generators, and validators) working iteratively to confirm software vulnerabilities. Vulnsage generates <strong>34.64% more exploits</strong> than existing tools and has discovered <strong>146 zero-day vulnerabilities</strong> in real-world software.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.04804v1">SkillX: Automatically Constructing Skill Knowledge Bases for Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>SkillX</strong> is a framework that automatically builds reusable skill libraries for agents by organizing raw experience into hierarchical tiers (strategic plans, functional skills, atomic skills) and iteratively refining them. When plugged into weaker agents on long-horizon tasks, the library consistently improves task success rates and execution efficiency across multiple benchmarks.</p></li></ul><h3 style="text-align: justify;">Baidu</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.24373v1">PP-OCRv5: A Specialized 5M-Parameter Model Rivaling Billion-Parameter Vision-Language Models on OCR Tasks</a></strong></p><ul><li><p style="text-align: justify;">Demonstrates that <strong>PP-OCRv5</strong>, a 5-million-parameter OCR model, matches billion-parameter vision-language models on text recognition through data-centric optimization&#8212;systematically improving data difficulty, accuracy, and diversity&#8212;while achieving better localization and fewer hallucinations.</p></li></ul><h3 style="text-align: justify;">Huawei</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.04184v1">AURA: Always-On Understanding and Real-Time Assistance via Video Streams</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>AURA</strong>, an end-to-end streaming VideoLLM framework that continuously processes live video to answer questions and provide proactive assistance in real-time. The system integrates context management, training objectives, and optimizations for stable long-horizon interaction, achieving state-of-the-art streaming benchmarks while running at 2 FPS with speech input/output on consumer hardware.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.24239v1">DVM: Real-Time Kernel Generation for Dynamic AI Models</a></strong></p><ul><li><p style="text-align: justify;">Tackles <strong>dynamic tensor shapes and control flows</strong> in AI models with <strong>DVM</strong>, a real-time compiler that encodes operators into bytecode for fast decoding on accelerators rather than expensive machine code compilation. The approach achieves <strong>up to 11.77&#215; speedup</strong> over existing compilers while reducing peak compilation time by <strong>5 orders of magnitude</strong>.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.26556v1">When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models</a></strong></p><ul><li><p style="text-align: justify;">Standard metrics can be misleading when evaluating distilled models: a 7B model that appears to match its teacher on perplexity actually performs 20.8 percentage points worse when generating text autoregressively. Hybrid-KDA addresses this with GenDistill, a distillation pipeline that optimizes for generation quality rather than log-likelihood matching. The resulting model retains 86&#8211;90% of the teacher&#8217;s accuracy while reducing memory usage by 75% and producing first tokens 2&#8211;4&#215; faster on 128K-token inputs.</p></li></ul><h3 style="text-align: justify;">iFlyTek</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.23840v1">VehicleMemBench: An Executable Benchmark for Multi-User Long-Term Memory in In-Vehicle Agents</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>VehicleMemBench</strong>, an executable benchmark for evaluating in-vehicle agents on <strong>multi-user long-term memory</strong> and tool use in a simulated environment. The benchmark reveals that even advanced models struggle with evolving user preferences and dynamic preference conflicts.</p></li></ul><h3 style="text-align: justify;">Meituan</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.02684v1">MBGR: Multi-Business Prediction for Generative Recommendation at Meituan</a></strong></p><ul><li><p style="text-align: justify;">Develops <strong>MBGR</strong>, a generative recommendation framework for multi-business scenarios that uses <strong>business-aware semantic IDs</strong> and <strong>multi-business prediction</strong> to avoid representation confusion and performance trade-offs across different businesses. Deployed in production at Meituan&#8217;s food delivery platform with measurable improvements in both offline and online metrics.</p></li></ul><h3 style="text-align: justify;">StepFun</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.25502v1">RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>RealRestorer</strong>, a large-scale image restoration model trained on nine common real-world degradation types, paired with <strong>RealIR-Bench</strong>, a benchmark of 464 real-world degraded images with metrics for degradation removal and consistency. The model achieves state-of-the-art performance among open-source restoration methods.</p></li></ul><h3 style="text-align: justify;">Tencent</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.02650v1">Revealing the Learning Dynamics of Long-Context Continual Pre-training</a></strong></p><ul><li><p style="text-align: justify;">Examines <strong>learning dynamics in long-context continual pre-training</strong> of industrial-scale language models by tracking Hunyuan-A13B across 200B tokens. The work reveals that models need 150B+ tokens to truly saturate, shows downstream benchmarks can mask incomplete learning (&#8221;deceptive saturation&#8221;), and demonstrates that <strong>attention patterns in retrieval heads reliably track training progress</strong> better than standard metrics.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.24533v1">UI-Voyager: A Self-Evolving GUI Agent Learning via Failed Experience</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>UI-Voyager</strong>, a self-evolving mobile GUI agent that learns from failures through <strong>Rejection Fine-Tuning</strong> (continuously filtering failed attempts) and <strong>Group Relative Self-Distillation</strong> (identifying decision points where trajectories diverge to guide corrections). A 4B model achieves 81.0% success on AndroidWorld tasks, outperforming recent baselines and human performance.</p></li></ul><h1 style="text-align: justify;">Technical AI Safety Publication Highlights</h1><p style="text-align: justify;">There were <strong>55 AI-safety-related papers published by Chinese researchers </strong>this edition. Highlights are below; a full list with summaries is available <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx#heading=h.6lltj3o1jco1">here</a>.</p><p style="text-align: justify;">Agentic safety continues to dominate Chinese AI safety research output, as it did last edition. Six of nine highlighted papers this fortnight focus on attacking, evaluating, or defending agent systems&#8212;now extending beyond OpenClaw specifically to MCP protocol vulnerabilities, agent skill supply chains, and multi-step evaluation methodology. This tracks with TC260&#8217;s standardization push covered above.</p><h2 style="text-align: justify;">Agentic Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.27490v1">AgentSwing: Adaptive Parallel Context Management Routing for Long-Horizon Web Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>AgentSwing</strong> is an adaptive context management system for long-horizon web agents that dynamically routes between multiple strategies rather than committing to a single fixed approach. The framework uses lookahead routing to evaluate parallel context-managed branches at each step, selecting continuations that balance <strong>search efficiency</strong> and <strong>terminal precision</strong>. Across benchmarks, AgentSwing achieves comparable or better performance than static methods while reducing interaction turns by up to 3&#215;, addressing a key reliability concern for autonomous agents operating under real-world token constraints.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Alibaba (Tongyi Lab)</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.03131v1">A Systematic Security Evaluation of OpenClaw and Its Variants</a></strong></p><ul><li><p style="text-align: justify;">Researchers evaluated <strong>security vulnerabilities in six OpenClaw-series agent frameworks</strong> across multiple backbone models using a benchmark of <strong>205 test cases</strong> covering the full agent execution lifecycle. All tested agents exhibited substantial vulnerabilities&#8212;significantly riskier than their underlying models alone&#8212;with <strong>reconnaissance, credential leakage, lateral movement, and privilege escalation</strong> as dominant failure modes. The findings demonstrate that agent security depends not just on model safety but on interactions between capability, tool use, multi-step planning, and runtime orchestration, requiring <strong>lifecycle-wide governance</strong> beyond prompt-level safeguards.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Xidian University, China Unicom</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.03070v1">Credential Leakage in LLM Agent Skills: A Large-Scale Empirical Study</a></strong></p><ul><li><p style="text-align: justify;">Researchers analyzed <strong>17,022 third-party skills</strong> for LLM agents and found <strong>520 vulnerable skills</strong> exposing credentials through <strong>10 distinct leakage patterns</strong>. Debug logging&#8212;particularly print statements outputting text visible to LLMs&#8212;caused 73.5% of leaks. <strong>76.3% of vulnerabilities required analyzing both code and natural language</strong> together, making single-modality scanning ineffective. Leaked credentials were exploitable without elevated privileges and persisted across forked versions, highlighting supply-chain risks in agent ecosystems that require both developer practices and runtime isolation mechanisms.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Fujian Normal University, Wake Forest University, Quantstamp, Nanyang Technical University, University of New South Wales, Griffith University, Zhejiang Sci-Tech University, the University of Tokyo, University of Alberta</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.01905v1">From Component Manipulation to System Compromise: Understanding and Detecting Malicious MCP Servers</a></strong></p><ul><li><p style="text-align: justify;"><strong>Model Context Protocol (MCP)</strong> enables LLMs to call external tools and data sources, but researchers find that attackers can compromise systems by chaining malicious behaviors across multiple MCP components rather than single attacks. The team built a dataset of <strong>114 malicious MCP servers</strong> and discovered that <strong>multi-component attack chains often succeed better than isolated manipulations</strong>, with attack success varying by component position. They propose <strong>Connor</strong>, a detector that analyzes shell commands pre-execution and monitors tool behavior in real-time to catch deviations from intended function, achieving <strong>94.6% detection accuracy</strong> on their dataset and identifying two malicious servers in the wild.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Fudan University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.02837v1">Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis</a></strong></p><ul><li><p style="text-align: justify;">Researchers conduct the <strong>first comprehensive security analysis of Agent Skills</strong>, an open standard for packaging domain-specific modules that LLM agents can download and execute. They map the <strong>full lifecycle</strong> (creation, distribution, deployment, execution) and identify <strong>seventeen threat scenarios across three attack layers</strong>, validated against five real-world incidents. The analysis reveals structural vulnerabilities&#8212;including blurred data-instruction boundaries, persistent trust after single approval, and unreviewed marketplace distribution&#8212;that the authors argue <strong>cannot be fixed through incremental patches alone</strong>, requiring fundamental redesign of the framework&#8217;s trust and verification mechanisms.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Chinese Academy of Sciences, University of Chinese Academy of Sciences</em></p><h2 style="text-align: justify;">Evaluation and Benchmarks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.06132v1">Claw-Eval: Toward Trustworthy Evaluation of Autonomous Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>Claw-Eval</strong> is an evaluation framework designed to catch safety and reliability failures in autonomous agents that existing benchmarks miss. It records <strong>all agent actions across three independent channels</strong> (execution traces, logs, environment snapshots) rather than checking only final outputs, and applies <strong>2,159 fine-grained scoring criteria</strong> covering completion, safety, and robustness. Testing 14 frontier models shows trajectory-opaque evaluation misses <strong>44% of safety violations</strong> and <strong>13% of robustness failures</strong>, while controlled error injection reveals agents are brittle&#8212;Pass@3 drops up to 24% under perturbation despite stable peak performance.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Peking University, The University of Hong Kong</em></p><h2 style="text-align: justify;">Interpretability</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2604.03147v1">Valence-Arousal Subspace in LLMs: Circular Emotion Geometry and Multi-Behavioral Control</a></strong></p><ul><li><p style="text-align: justify;">Researchers map a <strong>valence-arousal subspace</strong> within LLM representations by training linear axes on emotion-labeled data and verifying them against human crowdsourced ratings. Steering along these axes produces <strong>monotonic shifts in emotional tone</strong> and, critically, bidirectional control over <strong>refusal and sycophancy</strong>&#8212;increasing arousal decreases refusal while increasing compliance. The effect replicates across three model architectures, suggesting safety-relevant tokens cluster in predictable emotional regions, creating a potential vulnerability for circumventing safeguards through affective steering.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Shanghai AI Lab, University of Chicago, Harvard University</em></p><h2 style="text-align: justify;">Guardrails and Deployment Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.25412v1">Beyond Content Safety: Real-Time Monitoring for Reasoning Vulnerabilities in Large Language Models</a></strong></p><ul><li><p style="text-align: justify;">This paper identifies <strong>reasoning safety</strong> as a distinct vulnerability in chain-of-thought models&#8212;separate from output safety. The authors map <strong>nine categories of unsafe reasoning behaviors</strong> (parsing errors, execution flaws, process failures) and show all occur in practice, including under adversarial attacks. They propose a <strong>Reasoning Safety Monitor</strong>, an external LLM that inspects each reasoning step in real-time and flags unsafe behavior, achieving <strong>85% error-type classification accuracy</strong>. For deployment of reasoning-intensive models, this addresses a gap: outputs can be correct while reasoning chains are logically inconsistent, inefficient, or manipulated.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: The Hong Kong University of Science and Technology, Zhejiang University of Technology</em></p><h2 style="text-align: justify;">Robustness and Adversarial Attacks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.24203v1">Invisible Threats from Model Context Protocol: Generating Stealthy Injection Payload via Tree-based Adaptive Search</a></strong></p><ul><li><p style="text-align: justify;"><strong>TIP (Tree-structured Injection for Payloads)</strong> is a black-box attack that exploits the Model Context Protocol&#8212;which lets LLMs call external tools&#8212;by poisoning tool responses to hijack agent behavior. The method uses <strong>tree-structured search guided by an attacker LLM</strong> to generate natural-language payloads that evade detection, achieving over 95% success rates with fewer queries than prior attacks and retaining 50%+ effectiveness against existing defenses. This work reveals a practical vulnerability in production MCP deployments where defenders have limited visibility into tool outputs.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations:  Fudan University, Shanghai Innovation Institute</em></p><h1 style="text-align: justify;">On the Horizon</h1><p style="text-align: justify;">At the Zhongguancun Forum, <a href="https://english.beijing.gov.cn/beijinginfo/sci/latesttrends/202603/t20260327_4567695.html">two new industry alliances were established</a>: the Beijing Artificial Intelligence Association and the Zhongguancun Artificial Intelligence Open Source Alliance. Details are limited, but we&#8217;ll be keeping an eye on them as they get stood up.</p><div><hr></div><p style="text-align: justify;">For more on how we select and track content, see our methodology <a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz">here</a>.</p><div><hr></div><p style="text-align: justify;"><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p>]]></content:encoded></item><item><title><![CDATA[China AI Bulletin 1]]></title><description><![CDATA[Developments from 11/3/26-25/3/26]]></description><link>https://chinaaibulletin.substack.com/p/china-ai-bulletin-1</link><guid isPermaLink="false">https://chinaaibulletin.substack.com/p/china-ai-bulletin-1</guid><dc:creator><![CDATA[Emmie Hine]]></dc:creator><pubDate>Thu, 26 Mar 2026 17:47:03 GMT</pubDate><enclosure url="https://substack-post-media.s3.amazonaws.com/public/images/4a2714ca-53f3-44d8-8484-9a9928927db8_2000x2000.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<p>Welcome to Issue 1 of the China AI Bulletin! China&#8217;s AI ecosystem is developing rapidly&#8212;new models, regulations, and safety research are emerging on a near-weekly basis. The China AI Bulletin tracks this full landscape so that researchers and policymakers can stay informed. Every two weeks, we bring you the latest on AI development, governance, and safety in China. Today&#8217;s highlights: the agentic AI craze (fueled by OpenClaw) continues, two groups drafting AI laws are explicitly prioritizing frontier safety and extreme risk prevention, and China is drafting a national MCP security standard.</p><p style="text-align: justify;"><em>Number of the week: <a href="https://www.caixin.com/2026-03-25/102427030.html">140 trillion</a> - China&#8217;s daily token usage</em></p><h1 style="text-align: justify;">Executive Summary</h1><ul><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/192213975/frontier-lab-developments">Frontier Lab Developments</a>:</strong></p><ul><li><p style="text-align: justify;"><strong>Notable Model Releases:</strong> Zhipu released <strong>GLM-5-Turbo</strong>, the <strong>first model purpose-built for OpenClaw agentic workflows</strong>. Alibaba released <strong>AgenticQwen-8B</strong> and <strong>AgenticQwen-30B-A3B</strong>, Qwen variants fine-tuned for agentic tasks. Xiaomi&#8217;s <strong>MiMo-V2-Pro</strong> was revealed as the mystery &#8220;Hunter Alpha&#8221; model on OpenRouter&#8212;initially speculated to be DeepSeek V4.</p></li><li><p style="text-align: justify;"><strong>Technical Papers:</strong> Frontier labs released 96 papers on arXiv over the last two weeks, with contributions from Alibaba (37), Huawei (14), Tencent (12), Baidu (9), and ByteDance (9). Highlights include <strong>SkillRouter</strong>, a <strong>compact pipeline that routes agents to relevant tools from ~80K skill repositories</strong> with 74% accuracy; <strong>TRiMS</strong>, which achieves <strong>over 80% token reduction in reasoning chains</strong> while maintaining accuracy; and <strong>Taming OpenClaw</strong>, which identifies<strong> critical multi-stage vulnerabilities</strong> in autonomous agent frameworks.</p></li></ul></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/192213975/technical-ai-safety-publication-highlights">AI Safety</a>:</strong> 88 safety-related publications in this edition, including <strong>agentic security research prompted by the OpenClaw adoption wave</strong>: <strong>ClawWorm</strong> demonstrates a <strong>self-replicating worm achieving 64.5% propagation success</strong> across agent ecosystems; <strong>ChainFuzzer</strong> finds <strong>82% of discovered vulnerabilities only manifest when agents chain multiple tools</strong>; and <strong>Trojan&#8217;s Whisper</strong> shows <strong>94% of bootstrap injection attacks evade existing scanners</strong>. On the interpretability side, <strong>SafeSeek</strong> identifies <strong>complete safety circuits in just 0.42% of model parameters</strong>, enabling surgical safety interventions.</p></li><li><p style="text-align: justify;"><strong><a href="https://chinaaibulletin.substack.com/i/192213975/domestic-ai-governance">Domestic AI Governance</a>:</strong> China&#8217;s two most prominent AI law drafting groups&#8212;CASS and CUPL&#8212;released new research priorities that <strong>explicitly prioritize frontier AI safety and extreme risk prevention</strong>. MIIT announced plans for a <strong>national MCP security standard</strong> amid a month-long government response to OpenClaw risks, and TC260 is <strong>establishing a dedicated AI Safety Working Group (WG9)</strong>. Huawei and Lenovo became the <strong>first Chinese companies to join the Agentic AI Foundation</strong> as Gold Members.</p></li></ul><h1 style="text-align: justify;">Frontier Lab Developments</h1><h2 style="text-align: justify;">Notable Model Releases</h2><p style="text-align: justify;">In line with the OpenClaw hype in China, most releases over the last two weeks focused on agentic capabilities. <strong>Zhipu</strong> released <strong><a href="https://docs.z.ai/guides/llm/glm-5-turbo">GLM-5-Turbo</a></strong>, the first model purpose-built for OpenClaw agentic workflows. It&#8217;s currently closed-source and &#8220;<a href="https://x.com/Zai_org/status/2033221428640674015">experimental</a>,&#8221; but they promise capabilities will be incorporated into their next open release. <strong>Alibaba </strong>released <strong><a href="https://huggingface.co/alibaba-pai/AgenticQwen-8B">AgenticQwen-8B</a></strong> and <strong><a href="https://huggingface.co/alibaba-pai/AgenticQwen-30B-A3B">AgenticQwen-30B-A3B</a></strong>&#8212;Qwen variants fine-tuned specifically for agentic tasks.</p><p style="text-align: justify;"><strong>Xiaomi&#8217;s <a href="https://mimo.xiaomi.com/mimo-v2-pro">MiMo-V2-Pro</a></strong> (<a href="https://huggingface.co/XiaomiMiMo">HuggingFace</a>) was revealed on March 19 as the mystery &#8220;<a href="https://www.reuters.com/business/media-telecom/mystery-ai-model-has-developers-buzzing-is-this-deepseeks-latest-blockbuster-2026-03-18/">Hunter Alpha</a>&#8220; model (~1 trillion parameters) that appeared on OpenRouter on March 11&#8212;initially widely speculated to be DeepSeek V4. MiMo team lead Luo Fuli <a href="https://www.scmp.com/tech/big-tech/article/3332502/chinese-ai-prodigy-luo-fuli-joins-xiaomi-industry-competition-talent-heats">contributed</a> to DeepSeek-V2 before joining Xiaomi in November 2025.</p><p style="text-align: justify;"><strong>Tencent </strong>released <strong>SAGE-GRPO</strong> (<a href="https://github.com/Tencent-Hunyuan/SAGE-GRPO">GitHub</a>), a reinforcement learning training method for improving model reasoning, and <strong>Covo-Audio-Chat</strong> (<a href="https://huggingface.co/tencent/Covo-Audio-Chat">HuggingFace</a>, <a href="https://arxiv.org/abs/2602.09823">technical report</a>), an audio chat model. They also <a href="https://www.caixinglobal.com/2026-03-18/tencent-to-launch-hunyuan-30-in-april-build-wechat-ai-agent-102424421.html">announced</a> that Hunyuan 3.0 would be launched in April, and that they&#8217;re developing an agent for WeChat.</p><p style="text-align: justify;">Finally, <strong>Baidu</strong> released <strong>Qianfan-OCR</strong> (<a href="https://huggingface.co/baidu/Qianfan-OCR">HuggingFace</a>), a document understanding model, and <strong>Moonshot</strong> open-sourced the code for <strong>Attention Residuals</strong> (<a href="https://github.com/MoonshotAI/Attention-Residuals">GitHub</a>), the architectural innovation behind Kimi K2.5&#8212;led, notably, by a <a href="https://x.com/masonwang025/status/2033623721584562500">high school student</a>.</p><h2 style="text-align: justify;">Technical Publications</h2><p style="text-align: justify;">Frontier labs released 96 papers on arXiv in the last two weeks. Highlights are below; a full list with summaries can be found <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.0">here</a>.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!fjD8!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!fjD8!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 424w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 848w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 1272w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!fjD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png" width="644" height="562" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/b1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:562,&quot;width&quot;:644,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:38817,&quot;alt&quot;:&quot;A table showing the number of papers published by each lab.&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/192213975?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A table showing the number of papers published by each lab." title="A table showing the number of papers published by each lab." srcset="https://substackcdn.com/image/fetch/$s_!fjD8!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 424w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 848w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 1272w, https://substackcdn.com/image/fetch/$s_!fjD8!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fb1f5e889-3a8e-4e61-9dbe-dd08f17889cc_644x562.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p></p><h3 style="text-align: justify;">Alibaba</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.17621v1">Complementary Reinforcement Learning</a></strong></p><ul><li><p style="text-align: justify;">Develops <strong>Complementary RL</strong>, a framework where an experience extractor and policy actor co-evolve during training&#8212;the actor optimizes on sparse rewards while the extractor adapts based on whether its experiences actually help the actor succeed. The approach achieves 10% performance gains over outcome-based baselines and scales robustly across multi-task settings.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.19835v1">FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization</a></strong></p><ul><li><p style="text-align: justify;">Develops <strong>FIPO</strong>, a training method that assigns credit to individual words based on how much they influence later reasoning steps, rather than treating all words equally. On Qwen2.5-32B, this <strong>extends reasoning chains from 4K to 10K+ tokens and boosts AIME math accuracy from 50% to 58%</strong>, matching o1-mini.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.22117v1">On the Direction of RLVR Updates for LLM Reasoning: Identification and Exploitation</a></strong></p><ul><li><p style="text-align: justify;">Examines which specific word-level changes actually matter when training models to reason better. The paper finds that tracking the <em>direction</em> of changes (whether a word became more or less likely) is more informative than tracking their size, enabling two practical improvements: <strong>amplifying learned reasoning patterns at test time without retraining</strong>, and <strong>focusing training on the highest-impact words</strong>, boosting reasoning performance across benchmarks.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.14864v1">Shopping Companion: A Memory-Augmented LLM Agent for Real-World E-Commerce Tasks</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>Shopping Companion</strong>, a memory-augmented LLM agent framework with a new benchmark spanning 1.2M products for long-term preference-aware e-commerce tasks. Using <strong>dual-reward reinforcement learning</strong>, it jointly optimizes memory retrieval and shopping assistance, achieving over 70% success&#8212;greater than GPT-4.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.22455v1">SkillRouter: Retrieve-and-Rerank Skill Selection for LLM Agents at Scale</a></strong></p><ul><li><p style="text-align: justify;">Demonstrates that <strong>skill body</strong> (full implementation text) is critical for routing LLM agents to relevant tools from massive repositories, and proposes <strong>SkillRouter</strong>, a compact 1.2B retrieve-and-rerank pipeline achieving 74% top-1 accuracy on ~80K skills. Current agent designs hiding implementation details cause 29-44 point accuracy drops, but SkillRouter&#8217;s two-stage approach operates on consumer hardware while outperforming zero-shot baselines.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.11619v1">Taming OpenClaw: Security Analysis and Mitigation of Autonomous LLM Agent Threats</a></strong></p><ul><li><p style="text-align: justify;">Examines security vulnerabilities in autonomous LLM agents like <strong>OpenClaw</strong> through a five-layer <strong>lifecycle framework</strong> covering initialization, input, inference, decision, and execution stages. The analysis identifies critical threats&#8212;<strong>indirect prompt injection, skill supply chain contamination, memory poisoning, and intent drift</strong>&#8212;and reveals that existing point-based defenses fail against cross-temporal, multi-stage attacks, necessitating holistic security architectures.</p></li></ul><h3 style="text-align: justify;">Baidu</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.16210v1">MOSAIC: Composable Safety Alignment with Modular Control Tokens</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>MOSAIC</strong>, a framework that enables <strong>composable safety alignment</strong> through learnable control tokens that can be flexibly activated at inference time to enforce context-dependent safety rules without entangling safety with general capabilities. The approach reduces over-refusal while maintaining model utility through <strong>order-based task sampling</strong> and a <strong>distribution-level alignment objective</strong>.</p></li></ul><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.17449v1">TRiMS: Real-Time Tracking of Minimal Sufficient Length for Efficient Reasoning via RL</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>TRiMS</strong>, a reinforcement learning framework that trains models to generate minimal sufficient reasoning chains, achieving over <strong>80% token reduction</strong> while maintaining or improving accuracy across benchmarks by combining <strong>GRPO</strong> with theoretical <strong>MSL (Minimal Sufficient Length)</strong> bounds. The approach establishes the first measurable lower bound for reasoning compression and identifies key structural factors that enable models to approach optimal brevity.</p></li></ul><h3 style="text-align: justify;">ByteDance</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.11103v1">Understanding by Reconstruction: Reversing the Software Development Process for LLM Pretraining</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>Understanding by Reconstruction</strong>, a framework that synthesizes latent development trajectories (planning, debugging, refinement steps) from static code repositories using multi-agent simulation and search-based CoT optimization. Pre-training Llama-3-8B on these reconstructed trajectories improves performance on long-context understanding, coding, and agentic reasoning tasks.</p></li></ul><h3 style="text-align: justify;">Huawei</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.22078v1">Do World Action Models Generalize Better than VLAs? A Robustness Study</a></strong></p><ul><li><p style="text-align: justify;">Examines whether <strong>world action models (WAMs)</strong> generalize better than <strong>vision-language-action models (VLAs)</strong> for robot control by testing both on perturbed benchmarks. <strong>WAMs</strong> like <strong>LingBot-VA</strong> and <strong>Cosmos-Policy</strong> show stronger robustness to visual and language perturbations, though <strong>VLAs</strong> can match performance with extensive robotic training data.</p></li></ul><h3 style="text-align: justify;">iFlyTek</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.17876v1">Edit Spillover as a Probe: Do Image Editing Models Implicitly Understand World Relations?</a></strong></p><ul><li><p style="text-align: justify;">Examines what happens when you ask an image editing model to change one thing&#8212;do the <em>other</em> changes it makes reveal genuine understanding of the world? Using a systematic framework, the authors find that <strong>spillover patterns vary 3.3x across architectures</strong> and that semantically meaningful changes maintain consistent proportions regardless of distance, suggesting some models develop <strong>real-world understanding rather than just spatial pattern matching</strong>.</p></li></ul><h3 style="text-align: justify;">Meituan</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15030v1">VTC-Bench: Evaluating Agentic Multimodal Models via Compositional Visual Tool Chaining</a></strong></p><ul><li><p style="text-align: justify;">Unveils <strong>VTC-Bench</strong>, a comprehensive benchmark with 32 visual tools and 680 problems to evaluate how well multimodal AI models compose and chain tools for complex vision tasks. Testing 19 leading models reveals critical gaps: even top performers like Gemini-3.0-Pro achieve only 51%, struggling with tool generalization and multi-step planning efficiency.</p></li></ul><h3 style="text-align: justify;">Moonshot</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15031v1">Attention Residuals</a></strong></p><ul><li><p style="text-align: justify;">Develops <strong>Attention Residuals (AttnRes)</strong>, a modification to how information flows between layers in deep models. Instead of fixed-weight accumulation&#8212;which dilutes useful signals as models get deeper&#8212;AttnRes lets each layer learn which prior outputs to attend to. The approach <strong>improves performance consistently across model sizes and downstream tasks</strong> with minimal memory overhead.</p></li></ul><h3 style="text-align: justify;">SenseTime</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.17826v1">FailureMem: A Failure-Aware Multimodal Framework for Autonomous Software Repair</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>FailureMem</strong>, a multimodal framework for automated software repair that combines <strong>workflow-agent architecture</strong>, <strong>region-level visual grounding</strong>, and a <strong>Failure Memory Bank</strong> to learn from past repair attempts. On SWE-bench Multimodal, it improves the resolved rate over prior work by 3.7%.</p></li></ul><h3 style="text-align: justify;">Tencent</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15797v1">OMNIFLOW: A Physics-Grounded Multimodal Agent for Generalized Scientific Reasoning</a></strong></p><ul><li><p style="text-align: justify;">Presents <strong>OMNIFLOW</strong>, a neuro-symbolic architecture that grounds frozen LLMs in physical laws via <strong>Semantic-Symbolic Alignment</strong> and <strong>Physics-Guided Chain-of-Thought</strong>, enabling zero-shot generalization across PDEs without domain-specific fine-tuning. The system outperforms deep learning baselines on turbulence, Navier-Stokes, and weather forecasting tasks while providing interpretable, physically consistent reasoning.</p></li></ul><h3 style="text-align: justify;">Xiaomi</h3><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.13019v1">ARL-Tangram: Unleash the Resource Efficiency in Agentic Reinforcement Learning</a></strong></p><ul><li><p style="text-align: justify;">Introduces <strong>ARL-Tangram</strong>, a resource management system for agentic RL that enables fine-grained sharing of external cloud resources (CPUs, GPUs) through <strong>action-level orchestration</strong>. The system improves action completion time by 4.3x, accelerates RL training by 1.5x, and reduces external resource consumption by 71.2% compared to static over-provisioning approaches.</p></li></ul><h1 style="text-align: justify;">Technical AI Safety Publication Highlights</h1><p style="text-align: justify;">There were <strong>88 AI-safety-related papers published by Chinese researchers </strong>this fortnight. Highlights are below; a full list with summaries is available <a href="https://docs.google.com/document/d/1cvyrqbcgyUMXY0kVyS4lxtZzxPjb1FUsb88kHZqWXVA/edit?tab=t.sgbz2cdxbuvx">here</a>.</p><h2 style="text-align: justify;">&#128269;Spotlight</h2><p style="text-align: justify;">Tens of thousands of people in China have become &#8220;lobster farmers,&#8221; but instead of crustaceans, they&#8217;re raising their own AI agents. Social media hype and free installation events have made China home to <a href="https://www.semafor.com/article/03/10/2026/ai-agents-take-off-in-china">over 96,200 OpenClaw agents</a> (over 40% of the global total as of March 9; the US is in second with 65,700)&#8212;plus an unknown number based on the emerging <a href="https://globalsemiresearch.substack.com/p/in-depth-analysis-of-chinese-openclaw">set of OpenClaw-like products</a> from major Chinese tech companies. However, Chinese cybersecurity groups are <a href="https://www.scmp.com/tech/tech-trends/article/3346138/china-issues-second-warning-openclaw-risks-amid-adoption-frenzy?module=top_story&amp;pgtype=homepage">issuing warnings</a> about the security risks of using OpenClaw, and researchers are uncovering further vulnerabilities. Multiple papers this edition probe the framework&#8217;s attack surface, and the vulnerabilities they expose aren&#8217;t unique to OpenClaw. They reflect architectural weaknesses common to any agentic system with shell access, tool chaining, and extensible skill libraries.</p><p style="text-align: justify;">Two security analyses tested OpenClaw directly. <strong><a href="http://arxiv.org/abs/2603.10387v1">Don&#8217;t Let the Claw Grip Your Hand</a></strong> ran 47 adversarial scenarios and found a 17% native defense rate, with sandbox escape attacks proving particularly effective; adding a human-in-the-loop layer improved rates to 19&#8211;92% depending on attack type. A <a href="http://arxiv.org/abs/2603.12644v1">security threat case study of OpenClaw</a> mapped vulnerabilities across five lifecycle stages&#8212;initialization, input, inference, decision, and execution&#8212;and proposed a tri-layered risk taxonomy spanning AI cognitive, software execution, and information system dimensions, advocating zero-trust execution and dynamic intent verification.</p><p style="text-align: justify;">The attack papers go further. <strong><a href="http://arxiv.org/abs/2603.15727v1">ClawWorm</a></strong> demonstrates a self-replicating worm that propagates through configuration files and spreads to peer agents via messaging, achieving 64.5% success from a single initial message. <strong><a href="http://arxiv.org/abs/2603.19974v1">Trojan&#8217;s Whisper</a></strong> embeds adversarial narratives in bootstrap files that 94% of scanners miss, and <strong><a href="http://arxiv.org/abs/2603.12614v1">ChainFuzzer</a></strong> finds that 82% of the 365 vulnerabilities it discovers only manifest when agents chain tools together&#8212;single-tool testing misses them entirely. The consistent finding is that point defenses fail against attacks that span multiple stages, tools, or time steps.</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!cM1h!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!cM1h!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 424w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 848w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 1272w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!cM1h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png" width="727" height="416" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:416,&quot;width&quot;:727,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:&quot;A graphic from ClawWorm depicting a network of agents with an infection spreading from one initially compromised agent to many others.&quot;,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A graphic from ClawWorm depicting a network of agents with an infection spreading from one initially compromised agent to many others." title="A graphic from ClawWorm depicting a network of agents with an infection spreading from one initially compromised agent to many others." srcset="https://substackcdn.com/image/fetch/$s_!cM1h!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 424w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 848w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 1272w, https://substackcdn.com/image/fetch/$s_!cM1h!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F78191b23-9c12-4b7a-b11b-f4e154c9f2bf_727x416.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption"><em>A figure from <a href="https://arxiv.org/pdf/2603.15727v1">ClawWorm</a> demonstrating the potential spread of a compromised agent.</em></figcaption></figure></div><p style="text-align: justify;">These OpenClaw-specific findings connect to a broader pattern in this edition&#8217;s agentic safety research, which highlights that agent security is a property of the system, not any specific model. Individual model alignment doesn&#8217;t prevent worms from propagating through skill supply chains, multi-tool chains from surfacing exploitable data flows, or memory systems from accumulating poisoned context.  As agentic systems move toward production use, the research suggests that the agentic architecture needs to be more fully examined to avoid critical security vulnerabilities.</p><h2 style="text-align: justify;">Agentic Safety</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.10749v1">AttriGuard: Defeating Indirect Prompt Injection in LLM Agents via Causal Attribution of Tool Invocations</a></strong></p><ul><li><p style="text-align: justify;"><strong>AttriGuard</strong> defends LLM agents against <strong>Indirect Prompt Injection</strong> by using <strong>causal attribution</strong>&#8212;determining whether tool calls stem from user intent or malicious external inputs. Rather than filtering content, <strong>the defense re-executes the agent under modified observations to test whether each proposed action remains necessary</strong>. Across benchmarks, AttriGuard maintains<strong> zero attack success rates against static attacks</strong> and <strong>maintains resilience against adaptive adversarial attacks</strong>, a key robustness gap for existing defenses.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Zhejiang University, Nanyang Technological University, City University of Hong Kong</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.12614v1">ChainFuzzer: Greybox Fuzzing for Workflow-Level Multi-Tool Vulnerabilities in LLM Agents</a></strong></p><ul><li><p style="text-align: justify;"><strong>ChainFuzzer</strong> is a <strong>greybox fuzzing framework</strong> designed to discover vulnerabilities that emerge only when LLM agents chain multiple tools together&#8212;cases where data from one tool becomes input to another. Existing single-tool testing misses these multi-step attack paths. The framework uses <strong>trace-guided prompt synthesis</strong> to reliably trigger target tool sequences and <strong>guardrail-aware fuzzing</strong> to bypass safety filters. Evaluation on 998 tools across 20 agent apps found <strong>365 reproducible vulnerabilities, with 82% requiring multi-tool execution</strong>, establishing methodology for hardening production agent systems.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Sun Yat-sen University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15727v1">ClawWorm: Self-Propagating Attacks Across LLM Agent Ecosystems</a></strong></p><ul><li><p style="text-align: justify;"><strong>ClawWorm</strong> demonstrates a self-replicating worm attack against production LLM agent ecosystems, achieving <strong>64.5% success rates across trials</strong>. The attack persists by hijacking agent configurations, executes payloads on each restart, and propagates autonomously to peer agents via messaging&#8212;requiring only a single initial message. The research reveals that <strong>skill supply chains</strong> remain universally vulnerable despite execution-level filtering, highlighting critical architectural gaps in multi-agent trust boundaries that current defenses fail to address.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Peking University, Sun Yat-sen University, Wuhan University, Tsinghua University, Singapore Management University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.19974v1">Trojan&#8217;s Whisper: Stealthy Manipulation of OpenClaw through Injected Bootstrapped Guidance</a></strong></p><ul><li><p style="text-align: justify;"><strong>Guidance injection</strong> attacks compromise autonomous coding agents by embedding adversarial operational narratives into bootstrap files, manipulating the agent&#8217;s reasoning without explicit malicious instructions. Researchers demonstrated <strong>26 malicious skills across 13 attack categories</strong> (credential theft, workspace destruction, and backdoors) achieving <strong>16&#8211;64% success rates</strong> on a realistic developer workspace benchmark, with <strong>94% evading existing scanners</strong>. The findings highlight architectural vulnerabilities in extensible agent ecosystems and underscore the need for <strong>capability isolation</strong>, <strong>runtime policy enforcement</strong>, and <strong>transparent guidance provenance</strong>.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Shanghai Jiao Tong University</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.14975v1">Why Agents Compromise Safety Under Pressure</a></strong></p><ul><li><p style="text-align: justify;">The authors identify <strong>Agentic Pressure</strong>&#8212;an endogenous conflict arising when agents cannot simultaneously achieve their goals and follow safety constraints. Under pressure, models exhibit <strong>normative drift</strong>, strategically violating safety guidelines while using advanced reasoning to construct linguistic justifications. The paper explores mitigation approaches like <strong>pressure isolation</strong>, which decouples decision-making from goal-achievement signals to restore alignment, highlighting a previously underexamined failure mode in deployed LLM agents.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Guangdong Provincial Key Laboratory of Brain-inspired Intelligent Computation (Southern University of Science and Technology)</em></p><h2 style="text-align: justify;">Robustness and Adversarial Attacks</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15397v1">SFCoT: Safer Chain-of-Thought via Active Safety Evaluation and Calibration</a></strong></p><ul><li><p style="text-align: justify;"><strong>SFCoT</strong> is a framework that monitors and corrects unsafe reasoning within LLM chains of thought rather than filtering only final answers. It uses a <strong>three-tier safety scoring system</strong> and <strong>multi-perspective consistency checks</strong> to detect risks during intermediate reasoning steps, then applies <strong>targeted calibration</strong> to redirect unsafe trajectories. Experiments show it <strong>reduces jailbreak success rates from 59% to 12%</strong> while maintaining general performance, positioning real-time process monitoring as a defense complementary to output-level filtering.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Tianjin University, NSFOCUS Technologies Group Co.</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15684v1">State-Dependent Safety Failures in Multi-Turn Language Model Interaction</a></strong></p><ul><li><p style="text-align: justify;">Multi-turn conversations expose structural safety vulnerabilities that single-query evaluations miss. Researchers introduce <strong>STAR</strong>, a diagnostic framework treating <strong>dialogue history as state transitions</strong> to systematically analyze how models drift from refusal behavior across conversation sequences. Analysis reveals <strong>monotonic drift</strong> from safety-related representations and <strong>abrupt phase transitions</strong> triggered by role-conditioning, demonstrating that models robust under static tests can undergo rapid safety collapse under structured interaction patterns&#8212;challenging assumptions that safety alignment is query-independent.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: University of Science and Technology of China, CFAR and IHPC, Nanyang Technological University, Tsinghua University, Beijing Electronic Science and Technology Institute</em></p><h2 style="text-align: justify;">Interpretability</h2><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.23268v1">SafeSeek: Universal Attribution of Safety Circuits in Language Models</a></strong></p><ul><li><p style="text-align: justify;"><strong>SafeSeek</strong> identifies the specific components responsible for safety behavior in LLMs&#8212;finding that <strong>safety can be localized to as little as 0.42% of a model&#8217;s parameters</strong>. Once identified, these &#8220;safety circuits&#8221; can be surgically modified: removing backdoor vulnerabilities while preserving normal function, or <strong>preventing safety from degrading when models are fine-tuned</strong> to be more helpful. The approach enables targeted safety fixes rather than broad retraining.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: University of Science and Technology of China, Squirrel Ai Learning, United Arab Emirates University, Nanyang Technological University, Zayed University, Intelligent Science &amp; Technology Academy of CASIC</em></p><p style="text-align: justify;"><strong><a href="http://arxiv.org/abs/2603.15615v1">Mechanistic Origin of Moral Indifference in Language Models</a></strong></p><ul><li><p style="text-align: justify;">Researchers find that LLMs treat opposing moral concepts&#8212;like &#8220;fairness&#8221; and &#8220;unfairness&#8221;&#8212;as nearly interchangeable internally, compressing them into <strong>indistinguishable representations</strong> despite surface-level alignment. Testing 23 models, <strong>none could reliably distinguish opposed moral categories</strong> regardless of scale or training. By reconstructing distinct internal representations for each moral concept, the authors achieve a <strong>75% win-rate on adversarial moral benchmarks</strong>, suggesting that fixing moral reasoning may require changes to how models represent concepts internally, not just how they behave.</p></li></ul><p style="text-align: justify;"><em>Institutional affiliations: Shanghai Artificial Intelligence Laboratory</em></p><div class="subscription-widget-wrap-editor" data-attrs="{&quot;url&quot;:&quot;https://chinaaibulletin.substack.com/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe&quot;,&quot;language&quot;:&quot;en&quot;}" data-component-name="SubscribeWidgetToDOM"><div class="subscription-widget show-subscribe"><div class="preamble"><p class="cta-caption">Finding this useful? Subscribe to get it in your inbox regularly:</p></div><form class="subscription-widget-subscribe"><input type="email" class="email-input" name="email" placeholder="Type your email&#8230;" tabindex="-1"><input type="submit" class="button primary" value="Subscribe"><div class="fake-input-wrapper"><div class="fake-input"></div><div class="fake-button"></div></div></form></div></div><p style="text-align: justify;"></p><h1 style="text-align: justify;">Domestic AI Governance</h1><h3 style="text-align: justify;">AI law drafting groups prioritizing frontier risk and agent governance</h3><p style="text-align: justify;">China&#8217;s two most prominent AI law drafting groups released &#8220;top 10&#8221; research priority lists within a week of each other, with some notable convergences around frontier AI safety, agent governance, and liability for AI-caused harm.<em> </em>For context, a comprehensive AI law remains in an early legislative stage, although Minister of Justice He Rong <a href="https://www.gov.cn/zhuanti/2026nztj/2026qglh/yw/202603/content_7062522.htm">confirmed</a> on March 13 that legislation is being &#8220;accelerated.&#8221; The Chinese Academy of Social Sciences (CASS) and China University of Political Science and Law (CUPL) both have proposed versions of comprehensive AI laws intended to serve as intellectual inputs to any eventual official law.</p><p style="text-align: justify;">On March 14, the Chinese Academy of Social Sciences released the <a href="http://iolaw.cssn.cn/zxzp/202603/t20260314_5976397.shtml">10 core revisions</a> made to its Model AI Law, led by Li Honglei (&#26446;&#27946;&#38647;), Party Secretary of CASS&#8217;s Institute of Law (the previous version, V3.0, is available <a href="https://www.aigl.blog/the-model-artificial-intelligence-law-mail-v-3-0/">here</a>). The update is described as a &#8220;comprehensive institutional design upgrade&#8221; implementing the 15th Five-Year Plan&#8217;s AI provisions. The 10 core modifications are:</p><ol><li><p style="text-align: justify;"><strong>AI for Science (AI4S) safe harbors</strong>&#8212;creating protected regulatory environments for scientific research, with guaranteed compute and public data access</p></li><li><p style="text-align: justify;"><strong>Data supply system</strong>&#8212;balancing training data needs with IP protection through &#8220;reasonable use&#8221; rules</p></li><li><p style="text-align: justify;"><strong>AI-generated content rights</strong>&#8212;establishing ownership rules and mandatory content labeling</p></li><li><p style="text-align: justify;"><strong>Foundation model extreme risk (&#26497;&#31471;&#23433;&#20840;&#39118;&#38505;) prevention</strong>&#8212;full-lifecycle risk management covering monitoring, early warning, and emergency response</p></li><li><p style="text-align: justify;"><strong>Edge AI regulation</strong>&#8212;compliance rules for on-device models with cloud-edge coordination</p></li><li><p style="text-align: justify;"><strong>Multi-model and agent governance</strong>&#8212;extending regulation from single systems to multi-agent interaction and tool-calling, with continuous risk assessment</p></li><li><p style="text-align: justify;"><strong>User obligations</strong>&#8212;expanding from developer/provider liability to full-chain responsibility covering development, deployment, and use</p></li><li><p style="text-align: justify;"><strong>Public sector AI rules</strong>&#8212;stricter transparency, fairness, and due process standards for government AI, with red lines for core public authority functions</p></li><li><p style="text-align: justify;"><strong>Cross-border AI governance</strong>&#8212;export compliance, international rule alignment, and reciprocal countermeasures</p></li><li><p style="text-align: justify;"><strong>Civil liability</strong>&#8212;tiered responsibility across foundation model providers, fine-tuners, and deployers following &#8220;whoever develops/controls/uses bears responsibility&#8221;</p></li></ol><p style="text-align: justify;">It also outlined <strong>10 Advanced Issues in AI Legal Governance</strong> (<strong>&#20154;&#24037;&#26234;&#33021;&#27861;&#27835;&#21313;&#22823;&#36827;&#38454;&#35758;&#39064;</strong>), framing a broader research agenda:</p><ol><li><p style="text-align: justify;"><strong>Theoretical foundations for comprehensive AI legislation</strong> amid accelerating global governance</p></li><li><p style="text-align: justify;">How <strong>major economies&#8217; policy shifts</strong> (especially US deregulation) should shape China&#8217;s AI industrial policy</p></li><li><p style="text-align: justify;"><strong>Open-source ecosystem governance</strong> and overseas legal risks for domestic models</p></li><li><p style="text-align: justify;"><strong>Agile expansion of the regulatory &#8220;toolbox&#8221;</strong> for model safety and ethical risks</p></li><li><p style="text-align: justify;"><strong>Rights and responsibilities for compute clusters</strong> and key digital infrastructure</p></li><li><p style="text-align: justify;"><strong>Legal adaptation for agents</strong> and edge-side models in deep application</p></li><li><p style="text-align: justify;"><strong>Copyright disputes and legitimate boundaries for training data</strong></p></li><li><p style="text-align: justify;"><strong>Fine-grained governance for &#8220;AI+&#8221;</strong> in critical application scenarios (healthcare, government, finance)</p></li><li><p style="text-align: justify;"><strong>Forward-looking construction of AI civil liability frameworks</strong></p></li><li><p style="text-align: justify;"><strong>Liability allocation</strong> across the embodied AI and autonomous driving value chain</p></li></ol><p style="text-align: justify;">One week later, Zhang Linghan (&#24352;&#20940;&#23506;), director of CUPL&#8217;s AI Law Research Institute, released her group&#8217;s <a href="http://edu.people.com.cn/n1/2026/0320/c367001-40685950.html">research agenda</a> at a <a href="https://mp.weixin.qq.com/s/bQHNuVJ5hGqbzuDYCHJnKA">seminar</a> co-hosted by the same 8-institution coalition behind the <a href="https://cset.georgetown.edu/publication/china-ai-law-draft/">2024 CUPL Scholars&#8217; Draft AI law</a>. The 10 issues are practice-oriented, and several directly reference the OpenClaw wave:</p><ol><li><p style="text-align: justify;"><strong>Agent behavior boundaries and authorization</strong>&#8212;OpenClaw was explicitly referenced in discussion. The question of where an agent&#8217;s authorized scope ends and who bears liability for unauthorized actions is framed as the most urgent practical gap.</p></li><li><p style="text-align: justify;"><strong>Human-like AI interaction risks</strong>&#8212;addressing emotional manipulation, over-reliance, and deceptive AI companions. In December of last year, the CAC released a <a href="https://www.cac.gov.cn/2025-12/27/c_1768571207311996.htm">draft law</a> on human-like AI, indicating legislative interest.</p></li><li><p style="text-align: justify;"><strong>Hallucination liability</strong>&#8212;the legal focus has shifted from whether AI output is itself illegal to the downstream harm when users rely on incorrect AI-generated information for medical, legal, or financial decisions.</p></li><li><p style="text-align: justify;"><strong>Personal information protection</strong> structural conflicts in the AI era</p></li><li><p style="text-align: justify;"><strong>IP and benefit distribution</strong> in human-AI collaboration</p></li><li><p style="text-align: justify;"><strong>Frontier (&#21069;&#27839;) AI safety and extreme risk prevention</strong>&#8212;cases of AI self-preservation, autonomous escape, and attempts to influence humans were cited in discussions, with concerns that the trend will accelerate. The group reached consensus that frontier AI safety &#8220;must not prioritize R&amp;D while neglecting prevention and control&#8221; and recommends full-lifecycle legal controls, stating that &#8220;extreme risks such as algorithm security, data security, system security, technology misuse, and loss of control cannot be ignored.&#8221;</p></li><li><p style="text-align: justify;"><strong>AI industrial policy legal regulation</strong></p></li><li><p style="text-align: justify;"><strong>Legal issues for AI companies expanding overseas</strong></p></li><li><p style="text-align: justify;"><strong>Coordination of technical standards, ethical guidelines, and legal rules</strong></p></li><li><p style="text-align: justify;"><strong>New governance tools for AI</strong></p></li></ol><p style="text-align: justify;">There are some notable convergences: both groups have arrived at frontier/extreme risk and agent governance as key legislative priorities, both would extend liability to the full value chain, and both call for preemptive risk frameworks rather than post-hoc enforcement. A year ago, neither topic featured prominently in Chinese AI law discussions.</p><h2 style="text-align: justify;">National Standards</h2><h3 style="text-align: justify;">MIIT announces MCP standard plan</h3><p style="text-align: justify;">On March 25, MIIT <a href="https://www.miit.gov.cn/jgsj/kjs/jscx/bzgf/art/2026/art_5945e44487d340e7a7d5a8109c910b15.html">released</a> a batch of proposed standards projects for 2026, including three AI standards under TC1 (MIIT&#8217;s AI Standardization Technical Committee, <a href="https://technode.com/2024/12/16/chinas-ministry-of-industry-and-information-technology-establishes-ai-standardization-technical-committee/">established December 2024</a> with the China Academy of Information and Communication Technology (CAICT) as secretariat):</p><ol><li><p style="text-align: justify;"><strong>AI Safety Governance: Model Context Protocol Application Security Requirements</strong> (&#20154;&#24037;&#26234;&#33021; &#23433;&#20840;&#27835;&#29702; &#27169;&#22411;&#19978;&#19979;&#25991;&#21327;&#35758;&#24212;&#29992;&#23433;&#20840;&#35201;&#27714;)&#8212;a basic standard establishing security requirements for MCP, the protocol that allows AI agents to interact with external tools and services. Lead drafters include CAICT, Ant Group, China Unicom Digital Intelligence, Beijing University of Aeronautics and Astronautics, and 360 Digital Security Group.</p></li><li><p style="text-align: justify;"><strong>AI Product Services: AI Application Service Provider Delivery Implementation Capability Maturity Requirements</strong> (&#20154;&#24037;&#26234;&#33021; &#20135;&#21697;&#26381;&#21153; &#20154;&#24037;&#26234;&#33021;&#24212;&#29992;&#26381;&#21153;&#21830;&#20132;&#20184;&#23454;&#26045;&#33021;&#21147;&#25104;&#29087;&#24230;&#35201;&#27714;)&#8212;a maturity model for AI service providers. Lead drafters include CAICT, Baidu, AsiaInfo Technology, and China Mobile.</p></li><li><p style="text-align: justify;"><strong>AI Key Infrastructure Technology: Intelligent Agent Interface Classification and Basic Requirements</strong> (&#20154;&#24037;&#26234;&#33021; &#20851;&#38190;&#22522;&#30784;&#25216;&#26415; &#26234;&#33021;&#20307;&#25509;&#21475;&#20998;&#31867;&#21450;&#22522;&#26412;&#35201;&#27714;)&#8212;standardizing how AI agents interface with systems. Lead drafters include China Mobile, CAICT, the Beijing Academy of AI (BAAI), Ant Group, Alibaba, iFlyTek, Baidu, Inspur, OPPO, ZTE, and AsiaInfo.</p></li></ol><p style="text-align: justify;">The MCP security standard is particularly notable. It lands in the middle of a month-long government response to OpenClaw security risks: CNCERT (China&#8217;s cybersecurity emergency response coordinator) <a href="https://www.news.cn/tech/20260310/959f13d18edb4759ae031a5e30523d23/c.html">warned</a> on March 10 that nearly 23,000 OpenClaw instances had exposed assets due to weak default configurations; MIIT <a href="https://news.cctv.cn/2026/03/11/ARTIU9NPnXcPCDiU9cOfqTlD260311.shtml">organized</a> providers to issue &#8220;six do&#8217;s and six don&#8217;ts&#8221; guidelines on March 11; government agencies and SOEs were <a href="https://www.bloomberg.com/news/articles/2026-03-11/china-moves-to-limit-use-of-openclaw-ai-at-banks-government-agencies">told</a> not to install OpenClaw on office devices; and CNCERT <a href="https://news.cgtn.com/news/2026-03-23/China-releases-OpenClaw-security-guidance-for-different-groups-1LJZt3chZ72/p.html">released</a> an &#8220;OpenClaw Security Use Practice Guide&#8221; on March 23 covering ordinary users, enterprise users, cloud providers, and developers.</p><p style="text-align: justify;">The intelligent agent interface standard (#3) is also significant&#8212;it has the broadest industry participation of the three, with all three state telcos, major corporate and state-backed AI labs, and multiple hardware/software companies. This standard could shape how AI agents connect to enterprise infrastructure across China.</p><p style="text-align: justify;">Separately, CAICT <a href="https://www.ithome.com/0/928/265.htm">launched</a> its own &#8220;Claw&#8221; series of standards on March 12, which will cover permission management, transparent execution, and controllable behavioral risks. And on the international side, Huawei and Lenovo became the <a href="https://www.scmp.com/tech/big-tech/article/3344667/huawei-joins-openai-google-global-push-advance-ai-standards-rare-china-us-collab">first Chinese companies</a> to join the <a href="https://aaif.io/">Agentic AI Foundation</a> (AAIF) as Gold Members on February 24&#8212;a Linux Foundation body governing MCP as an open standard, alongside Anthropic, OpenAI, and Google.</p><p style="text-align: justify;">This demonstrates rapid tech institutionalization: national security standards for agent protocols, agent interface standards, developing trustworthiness requirements for agent products, and participating in international agent governance, all within weeks of the OpenClaw adoption wave that made agent security a salient concern.</p><h3 style="text-align: justify;">TC260 plans AI working group with safety focus</h3><p style="text-align: justify;">TC260, China&#8217;s National Technical Committee on Cybersecurity under the Standardization Administration of China, is establishing a dedicated AI Safety Standards Working Group (WG9), issuing a <a href="https://www.tc260.org.cn/portal/article/2/5e06ca6eec464ebdb3c4630208d74131">formal member recruitment notice</a> on March 20. WG9&#8217;s scope will include &#8220;investigating current conditions and development trends in AI safety/security standards, proposing an AI safety/security standards framework, and conducting R&amp;D of AI safety/security standards.&#8221;</p><p style="text-align: justify;">Until now, AI safety standardization at TC260 was primarily conducted by SWG-ETS (the Emerging Technologies Security Standards Special Working Group), whose <a href="https://www.tc260.org.cn/portal/about/org">mandate</a> covers AI, quantum computing, blockchain, and cloud. SWG-BDS, the old Big Data Security Special Working Group that&#8217;s now WG8, produced TC260&#8217;s <a href="https://www.tc260.org.cn/resources/upload/file/2025/11/27/1764212430522_%E4%BA%BA%E5%B7%A5%E6%99%BA%E8%83%BD%E5%AE%89%E5%85%A8%E6%A0%87%E5%87%86%E5%8C%96%E7%99%BD%E7%9A%AE%E4%B9%A6%EF%BC%882019%E7%89%88%EF%BC%89.pdf">2019 AI Safety Standardization White Paper</a>. The creation of WG9 consolidates this distributed work into a single dedicated body&#8212;mirroring the path data security took when SWG-BDS was elevated to WG8.</p><p style="text-align: justify;"><em>Context on TC260</em>: TC260 has been increasingly active on AI safety: the committee released the <a href="https://www.cac.gov.cn/2024-09/09/c_1727567886199789.htm">AI Safety Governance Framework v1.0</a> in September 2024 and <a href="https://www.cac.gov.cn/2025-09/15/c_1759653448369123.htm">v2.0</a> in September 2025, and by 2025 had drafted or published <a href="https://sesec.eu/wp-content/uploads/2025/11/SESEC-V-Report-TC260-AI-Standardization-Sep-2025.pdf">dozens of national and industry standards related to AI safety governance</a>. The mandatory national standard GB 45438-2025 on AI-generated content labeling and the recommended standard GB/T 45654-2025 on basic safety requirements for generative AI services were both published in 2025.</p><h3 style="text-align: justify;">TC28/SC42 formulates industrial data standards</h3><p style="text-align: justify;">The TC28/SC42 National Information Technology Standardization Technical Committee Artificial Intelligence Subcommittee is <a href="https://std.samr.gov.cn/search/orgDetailView?data_id=A132FB8FCABB4BE9E05397BE0A0AB880">formulating</a> 11 standards on industrial data, including a few on big data and datasets:</p><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!Tr0R!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!Tr0R!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 424w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 848w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 1272w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!Tr0R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png" width="1123" height="399" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:399,&quot;width&quot;:1123,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:null,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:null,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:null,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!Tr0R!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 424w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 848w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 1272w, https://substackcdn.com/image/fetch/$s_!Tr0R!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F4fd31e8c-a6e6-4181-aefa-fb9ebf074752_1123x399.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!t8jC!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!t8jC!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 424w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 848w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 1272w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!t8jC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png" width="1003" height="480" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:480,&quot;width&quot;:1003,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:115156,&quot;alt&quot;:&quot;A table with English translations of the standards. Titles: Product master data&#8212;Product data dictionary of hardware electric tools   Product master data&#8212;Product Data Dictionary of bearings   Product master data&#8212;Product data dictionary of intelligent sensors   Industrial big data&#8212;Classification of forecast data   Product master data&#8212;Model classification   Industrial big data&#8212;Technical requirements for production and operation status forecast data   Industrial big data&#8212;forecasting Capability Maturity assessment model   Product master data&#8212;Product data dictionary of valves   Product master data&#8212;Product data dictionary of high-performance resin Product master data&#8212;Product data dictionary of water saving irrigation equipment   Industrial High-quality dataset&#8212;Construction requirements&quot;,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:true,&quot;topImage&quot;:false,&quot;internalRedirect&quot;:&quot;https://chinaaibulletin.substack.com/i/192213975?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="A table with English translations of the standards. Titles: Product master data&#8212;Product data dictionary of hardware electric tools   Product master data&#8212;Product Data Dictionary of bearings   Product master data&#8212;Product data dictionary of intelligent sensors   Industrial big data&#8212;Classification of forecast data   Product master data&#8212;Model classification   Industrial big data&#8212;Technical requirements for production and operation status forecast data   Industrial big data&#8212;forecasting Capability Maturity assessment model   Product master data&#8212;Product data dictionary of valves   Product master data&#8212;Product data dictionary of high-performance resin Product master data&#8212;Product data dictionary of water saving irrigation equipment   Industrial High-quality dataset&#8212;Construction requirements" title="A table with English translations of the standards. Titles: Product master data&#8212;Product data dictionary of hardware electric tools   Product master data&#8212;Product Data Dictionary of bearings   Product master data&#8212;Product data dictionary of intelligent sensors   Industrial big data&#8212;Classification of forecast data   Product master data&#8212;Model classification   Industrial big data&#8212;Technical requirements for production and operation status forecast data   Industrial big data&#8212;forecasting Capability Maturity assessment model   Product master data&#8212;Product data dictionary of valves   Product master data&#8212;Product data dictionary of high-performance resin Product master data&#8212;Product data dictionary of water saving irrigation equipment   Industrial High-quality dataset&#8212;Construction requirements" srcset="https://substackcdn.com/image/fetch/$s_!t8jC!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 424w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 848w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 1272w, https://substackcdn.com/image/fetch/$s_!t8jC!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F33babab4-223f-4cce-8fb1-a4eceaa4177d_1003x480.png 1456w" sizes="100vw" loading="lazy"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a><figcaption class="image-caption">English titles are the official titles from TC28/SC42</figcaption></figure></div><h1 style="text-align: justify;">On the Horizon</h1><p style="text-align: justify;">The <a href="https://english.news.cn/20260325/64a09b9e2bad4b27a7bad43c94415306/c.html">Zhongguancun Forum</a> runs from 25 to 29 March with the theme &#8220;Full Integration Between Technological and Industrial Innovation;&#8221; early reports indicate a heavy focus on humanoid robots. We&#8217;ll bring you any AI-relevant developments in the next edition!</p><div><hr></div><p style="text-align: justify;">For more on how we select and track content, see our methodology <a href="https://docs.google.com/document/d/1rgBkmexUGLrUoNssJoV4UTH5m3jfbV_-2hO4cFm9wts/edit?tab=t.0#heading=h.vgkxt8csbiiz">here</a>.</p><div><hr></div><p style="text-align: justify;"><em>The China AI Bulletin is maintained by the Safe AI Forum (SAIF), a US 501(c)3 facilitating international cooperation on extreme AI risks. Views expressed represent individual authors&#8217; perspectives, not official SAIF positions.</em></p>]]></content:encoded></item></channel></rss>