China’s AI-safety trajectory is not necessarily a delayed version of America’s
The WSJ article, “China Also Thinks AI Could Kill Us. But First, It Wants to Match the U.S.,” argues that some Chinese researchers see safety work on the most capable AI systems as less urgent because Chinese labs perceive themselves to be trailing U.S. labs.
I see two possible arguments for this view. First, scarcer computing resources mean Chinese labs have less spare compute to devote to safety. Second, less capable models may be involved in fewer safety incidents.
But the second point may not always hold. In this brief note, I argue that China’s AI-safety trajectory need not be a delayed version of America’s, and that safety work in China could make a distinct contribution.
I give three reasons: the incidents may differ, the two countries may disclose them differently, and their governments may respond differently. The relationships shown in the diagrams are hypothetical.
Baseline model
Figure 1: A baseline model in which safety incidents rise later in China than in the US.
Figure 1 illustrates the baseline model I want to question.
In this model, Chinese models trail US models in capability by a given number of months and are involved in similar safety incidents after a delay.1 As such, China has little independent safety work to do because US researchers will have already identified the problems and developed solutions for the capability levels that Chinese models would eventually reach.
Alternative model: safety incidents need not rise with capability in the same way in the US and China
Figure 2: Hypothetical: safety incidents rise differently with AI capability in the US and China.
The baseline mental model is misleading.
Firstly, the US and China need not follow the same relationship between capability and safety incidents. Hypothetically, Chinese models could be involved in a particular kind of safety incident at a lower overall capability level than US models. For example, they might escape a sandbox, a restricted computing environment, at a lower capability level because cybersecurity at a Chinese company could be worse. For another example, open-weight models make their numerical parameters available for others to modify. Further training can weaken safety safeguards. Differences in how people use models could also affect which safety incidents occur.
Secondly, the relationship (slope) between safety incidents and capability could also be steeper for one country’s models. A similar increase in capability could accompany a larger increase in safety incidents. A persistent capability gap would therefore not guarantee a persistent gap in incident frequency or severity.
Lastly, even when models pose similar overall risks, the kinds of safety incidents could differ. This example is entirely hypothetical: US models might more often take actions without authorization, while Chinese models might exploit flaws in their reward system rather than complete the intended task. A method that addresses one problem may not address the other.
Evidence and examples
The earliest reports I know of involving US and Chinese models acting beyond their intended sandbox boundaries concern Claude Mythos Preview and ROME, respectively. I also include a later, related example involving Kimi K3. The Chinese models have lower estimated capabilities than Mythos Preview.
US models
Claude Mythos Preview
Anthropic reported that an earlier version of Claude Mythos Preview was put in a secured container and explicitly told to escape and contact the researcher. It developed a multi-step exploit, obtained broader internet access, emailed the researcher, and posted details of the exploit publicly without being asked.
ECI (Epoch Capabilities Index): my rough estimate is about 161.
Chinese models
ROME
The ROME team reported that, during RL, an AI agent opened a remote-access connection outside its intended environment and used computing resources for cryptocurrency mining. Neither action was requested or necessary for its tasks.
ECI: my rough estimate is about 142.
Kimi K3
Frontier Security reported that Kimi K3 retrieved a cybersecurity benchmark’s published solution from GitHub instead of solving the task itself. GitHub was reachable through a network allowlist intended for package maintenance, although most websites were blocked. The model was not instructed to escape the sandbox.
ECI: approximately 158 for Kimi K3.
My rough ECI estimates for Mythos Preview and ROME are based mainly on their reported benchmark results and comparisons with models I judged to have similar capabilities.
These cases do not establish when either country first encountered such incidents, or which models are more likely to produce them. Delayed or missing disclosure means earlier incidents may have gone unreported.
I would love to see more rigorous research on how safety incidents relate to capability across companies and countries.
The two countries may disclose safety incidents differently
Figure 3: Hypothetical differences in public reports of safety incidents.
Public reports give us an incomplete picture of underlying risks. Even if two sets of models were involved in similar safety incidents, differences in evaluation and disclosure could produce different numbers of public reports.
For example, independent evaluators may have different access to labs, and disclosure rules may differ. Figure 3 illustrates how those differences could affect reporting; the direction and size of the gap are hypothetical.
Evidence and examples
Independent evaluation
METR at Anthropic
METR described a researcher spending three weeks testing Anthropic’s monitoring systems in a post published in March 2026. This is an example of an independent evaluator working inside a lab. There is no known embedding of third-party evaluators at Chinese AI companies yet.
Disclosure rules
China’s vulnerability regulations
Chinese vulnerability disclosure rules are more stringent (in terms of CVEs). China’s Regulations on the Management of Security Vulnerabilities in Network Products restrict public disclosure of software and hardware security vulnerabilities (see Article 9). Such rules could affect the public information available about some AI-system vulnerabilities.
Given safety incidents, the Chinese government’s policy responses could differ from the US government’s
Figure 4: Hypothetical policy response differences between the two countries.
Government responses may also differ in two ways: how quickly officials recognize a problem and decide to act, and how quickly an agreed response is carried out.
Recognizing problems: China might have a higher threshold for what is deemed a “problem”.
Policy implementation: China generally acts quickly after a decision.
Figure 4 illustrates one possible outcome. A slowdown threshold is the level of safety incidents at which policymakers decide to intervene. In this scenario, the US starts slowing the rise in incidents after a short delay, at a lower incident level than China. China responds later and slows the rise more sharply, but too late to stay below the hypothetical catastrophe threshold.
Different assumptions could reverse this outcome. If China decided to intervene at a slightly lower incident level and the US took sufficiently long to implement its response, China might in the end have responded more promptly than the US.
Evidence and examples
Policy implementation
Education policy
The central “Double Reduction” policy restricting after-school academic tutoring was issued on July 24, 2021. By December, the Ministry of Education reported that the number of offline academic tutoring institutions had fallen 83.8%, and online institutions 84.1%.
Wuhan lockdown
For Covid-19, Wuhan announced major transport restrictions around 2 a.m. on January 23, 2020, with closures taking effect at 10 a.m. By January 29, all mainland provincial-level jurisdictions had activated the highest emergency-response level, according to the government’s account.
Recognizing problems
Early Covid-19 warnings
Chinese authorities suppressed early COVID-19 warnings. A whistleblower warned about COVID 19 on December 30 but was reprimanded. Wuhan’s transport shutdown began on January 23, more than three weeks after those warnings.
Wenzhou train collision
Before the July 2011 Wenzhou train crash, the Ministry of Railways had prioritized construction speed over safety (e.g. allowing seriously defective train-control equipment into service without field testing). The State Council ordered nationwide high-speed rail safety inspections only after the crash.
In this article, I illustrated several reasons why Chinese contributions to AI safety are important and why Chinese researchers should not be seen as merely following their US counterparts. I want to stress I am not sure about the specifics of the models here, but I do believe the baseline model is too simplistic. I look forward to more work being done to enlighten us about the current situation of AI safety in China.
Footnotes
“Safety incidents” can refer to either the number of incidents or their severity. All relationships and thresholds shown in the diagrams are hypothetical, not measured.↩︎