
Chinese and Japanese model developers are claiming parity with Anthropic’s flagship model, known internally as Mythos, on safety benchmarks, according to reports in the Wall Street Journal and The Information, a result that challenges the assumption that frontier safety capability belongs to American labs. Several new models released by Chinese companies, and at least one from a Japanese developer, are said to score at or near Anthropic’s level in third-party tests of harmful-output resistance and refusal behavior.
The claims are the latest sign that the AI competition is changing dimensions. For most of the past three years, the race was defined by raw capability: who could train the biggest model, score highest on reasoning tests, and ship the most impressive product first. Safety was treated as a Western specialty, a function of research culture and regulatory pressure that Chinese labs were assumed to deprioritize. The new benchmark results complicate that picture, according to researchers who have reviewed them, and they have landed at a moment when safety has become a commercial differentiator as well as a technical one.
Anthropic has built much of its identity around safety. The company was founded in 2021 by former OpenAI researchers who argued that AI development needed stronger guardrails, and it has made safety a marketing position as well as a technical one, publishing detailed policies and committing publicly to responsible scaling practices. Its Mythos model, introduced last year, was marketed in part on its resistance to manipulation and its refusal to assist with harmful requests. The company’s enterprise business has grown on the strength of that positioning, with customers citing safety documentation as a reason to choose Claude over alternatives.
The Chinese results do not mean the models are identical. Safety testing is a narrow window: benchmarks measure whether a model refuses a defined set of harmful requests, not whether its underlying reasoning is safe in every context. Researchers caution that benchmark scores can be engineered, with developers tuning models specifically to pass evaluations, and that real-world safety depends on systems and oversight that do not show up in a test score. Even with those caveats, the convergence is notable, because it mirrors the path capability took two years ago.
Still, the direction of travel matters. Chinese labs have spent the past two years closing the capability gap with American models, a process driven by open-weight releases from companies such as Alibaba and DeepSeek that made frontier-class models available for anyone to study and improve. Safety appears to be following the same path: as the technology becomes commoditized, the practices around it spread with it, and a lead that once looked structural starts to look temporary. The pattern is consistent across every layer of the stack, from chips to models to the safety teams around them.
The development has landed at an awkward moment for the American industry. U.S. policymakers have justified export controls and access restrictions partly on the argument that American labs are uniquely responsible, and that frontier AI should stay close to the institutions best equipped to govern it. If Chinese models match American models on safety as well as capability, that argument loses some of its force, and the case for restricting access becomes harder to make. Officials have not publicly responded to the benchmark claims, but the reports have circulated inside agencies that oversee export policy, according to people familiar with the matter.
Analysts said the competitive stakes are straightforward: safety is becoming a product feature, not a moral position. Enterprises choosing between model providers increasingly ask about red-team testing, refusal behavior and compliance tooling, and a lab that can document strong safety performance has a commercial advantage. Anthropic has built its enterprise business partly on that basis; if competitors can match the documentation, the differentiation narrows, and the buying decision shifts back to price and integration quality.
For the Chinese industry, the results are a matter of practical necessity as much as pride. Chinese regulators have imposed their own content rules on AI products, and labs building models for the domestic market have had to solve refusal and filtering problems at scale. Some of that work has produced techniques that translate directly to Western safety benchmarks, according to researchers who study the models. The same engineering muscle that compressed training costs is now being applied to alignment, and the early returns are visible in the published scores.
The most consequential question is what the trend means for the pace of the race itself. If safety is converging along with capability, then the distinction between responsible and irresponsible AI development blurs, and the market is left competing on speed, price and distribution. That may be good for users, who get safer models faster, and uncomfortable for the companies that built their brands on the opposite promise. Anthropic’s answer, according to people familiar with its thinking, is that benchmarks are a floor rather than a ceiling, and that the depth of its safety work will remain visible only over time.


