The ranking that software engineers refresh obsessively changed its top tier this week. Qwen3.7-Max, Alibaba’s newest flagship model, scored 1541 on Code Arena, a widely cited third-party benchmark for AI coding, placing it second among all models — behind only Anthropic’s Claude series — and ahead of GPT-5.5, Gemini-3.5-Flash, GLM-5.1 and Kimi-K2.6. Among the large model developers, Alibaba now holds the No. 2 slot.
The result is the latest marker of a shift that has been building for two years: the gap between Western and Chinese frontier models, once considered unbridgeable, has narrowed to the point where the top of the leaderboard is genuinely contested. Qwen, Alibaba’s model family, has been the vehicle for that climb. Its open-weight releases are among the most downloaded in the world, used by developers who run models on their own hardware instead of through an API, and the series has become a fixture of the open-source AI ecosystem.
Code Arena tests what working programmers actually care about: whether a model can write, fix and extend real software under realistic conditions. Models work through engineering tasks drawn from actual codebases — fixing bugs, implementing features, writing tests — and scores reflect the share of tasks completed successfully. A spot at the top of that ranking translates directly into adoption, because coding assistants have become the fastest-growing category of AI products, and developers choose tools on measurable ability.
The competitive picture has narrowed into a two-tier structure. Anthropic’s Claude series sits at the top, powered by agentic tooling that lets its models work through complex coding workflows. Below it, a pack trades places with every release: OpenAI’s GPT series, Google’s Gemini, and China’s Qwen, GLM and Kimi. Each new launch reshuffles the order, and no member of the pack stays on top for long. Qwen’s climb this week is Alibaba’s claim to lead that pack.
The business logic behind Alibaba’s push is straightforward. The company gives away or cheaply licenses its models to pull developers into its cloud, the same pattern it used in earlier rounds of open-source software. A developer who builds on Qwen today is likely to buy compute from Alibaba Cloud tomorrow, and coding is the workload most likely to drive that conversion. The company has made AI its top strategic priority, and Qwen is the vehicle.
Alibaba’s timing is deliberate. The company has spent heavily on model research while keeping its cloud prices aggressive, and it has signaled that winning developers matters more than near-term AI margins. A top-two coding result is the kind of proof point that turns into enterprise contracts, particularly in Asia, where Alibaba Cloud is a default choice for many companies and Qwen is already embedded in their toolchains. The international push follows the same logic: the more widely Qwen is used, the more demand flows back to Alibaba’s infrastructure. The flagship’s performance also lifts the family’s open-weight models, which developers can run for free and which carry the same brand into private deployments.
Pricing is the other weapon. Chinese models are typically far cheaper per token than their Western counterparts, and Qwen’s coding ability at a fraction of the price changes the economics for teams that use AI heavily. Analysts said the gap between first and second on coding benchmarks is small and shrinking; what matters more is the ecosystem around a model — tooling, integrations, cost and trust. On trust, Chinese models face a handicap in Western enterprises, some of which are cautious about sending code to providers outside their own jurisdictions.
The benchmark result also carries a message for the broader model race. China’s labs have concentrated on engineering excellence and efficient architectures, and their models now trade at the frontier in the capabilities that pay for themselves — coding chief among them. The result does not make Alibaba the best model developer in the world, and the Claude series’ hold on the top spot shows how far ahead the leading agentic systems remain. But it makes the argument that the frontier is now a multinational affair.
There are caveats to any leaderboard reading, analysts noted. Benchmarks measure snapshots: scores depend on test selection, the order in which models were released, and the constant patching of problems the community finds in the tests themselves. A single release can jump ten places, and the same release can slip as new tests arrive. The durable signal, they said, is the trend line — Chinese models have climbed steadily against Western ones for two years, and the climb has not stopped.
For developers, the practical takeaway is more choice at lower prices. The reshuffle means the best coding model of the month can come from Hangzhou rather than San Francisco, and teams can switch when it does. Alibaba will press the advantage with aggressive cloud pricing and international availability; rivals will answer with their own releases. The one certainty, analysts said, is that the order will change again — and the pace of change is the point.


