Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

China's Open-Weight Models Take the Lead: New Data on Who Actually Gets Used

Chinese AI models took 57% to 67% of all tokens on OpenRouter in the week of 14 September 2026, up from 6% to 13% in February, according to usage data CNBC reported on 26 September.

AI & modelsExplainerRachel NwosuPublished: 28 September 20266 min readSources 5
China's Open-Weight Models Take the Lead: New Data on Who Actually Gets Used

Two benchmark tables now tell the story of the AI model race. One measures how good models are on tests. The other measures which ones developers actually pay for. In 2026 the second table moved first, and it moved toward Chinese open-weight releases.

Chinese models went from a small minority of traffic to a majority on two developer gateways, OpenRouter and Vercel. The usage data behind that shift was shared with CNBC and published on 26 September. On OpenRouter, Chinese models accounted for 57% to 67% of tokens used in the week of 14 September, up from 6% to 13% in February. On Vercel the share rose to 55% in August from 11% in January. OpenRouter's figure covers companies in the U.S., Europe and what it defines as the "Global South", 82 countries across Central and South America, Africa and Asia. Vercel did not specify a geographic breakdown.

That is not the same as a quality lead. CNBC reported that the most advanced U.S. models still lead most benchmarks, and that U.S. frontier models still attract more overall spending.

Price, then competence

Peter Walker, head of insights at OpenRouter, told CNBC that Chinese open source models released this year "can credibly perform in advanced agentic use cases, especially in regards to coding, in a way that was just not true in late 2025." He added that they are "incredibly cost-effective compared to most models from American labs." Harpreet Arora, head of agentic infrastructure at Vercel, put the mechanism plainly to CNBC: "Chinese models are becoming capable enough for more tasks at a much lower cost. Once a model meets the quality bar for the job, that price difference becomes compelling." Arora also said companies still want frontier U.S. models for some more complicated tasks. That suggests the shift is workload-specific rather than a wholesale switch.

Dianne Penn, Anthropic's head of product management for research and labs, told CNBC the company was trying to make its models answer "more efficient, so it uses less tokens depending on your effort setting." OpenAI and Anthropic both announced cheaper models in the week before CNBC's report. Read together with the usage numbers, that is a price response to a price problem.

Who is actually switching

The geography matters. Businesses in what OpenRouter defines as the Global South have been the biggest users of Chinese AI models on its system, with 67% of the tokens those companies use going to Chinese models. About half of all tokens on OpenRouter are used by companies in the U.S.

Daniel Remler, a senior fellow in the technology and national security program at the Center for a New American Security, told CNBC the Chinese AI buildout represents "real economic and security risks for the United States." His stated worry is not benchmark scores: "The ultimate concern is that the integration of Chinese AI models pulls countries into a Chinese technology sphere of influence that hardens into geopolitical alignment." Remler said Southeast Asia in particular may see significant uptake given economic and cultural links with China, and that "anywhere from Lagos to São Paulo to Jakarta where entrepreneurs and governments are looking for cheap, open models, will look first to Chinese AI."

Two U.S. House Committees are investigating the impact of rising adoption of Chinese models, according to CNBC, and the U.S. has restricted Chinese AI companies from buying the most advanced chips through export controls. Washington's concerns include remote access to Nvidia chips through overseas data centers and "distillation", where new models mimic older, more established ones. AI was a major focus as U.S. President Donald Trump and Chinese President Xi Jinping met in the week of the report, CNBC said.

The staleness problem underneath

There is a second, quieter problem with the way model releases get judged, and it has nothing to do with geopolitics. A release date is when a lab shipped a model. The training cutoff is when it stopped reading. The gap between the two is how far behind the model already was on launch day.

A page published on 16 September, "How stale is your AI?", lists release dates and training cutoffs for 20 models across 8 labs. GPT-6 Astra, which it dates to 3 September 2026, has a training cutoff of 30 April 2026. GPT-6 Sol has a cutoff of 20 April 2026, and GPT-6 Luna 18 May 2026. Claude Opus 5.5 and Claude Fable 5.1 both stop reading in June 2026. Ten of the 20 models on that page carry a cutoff their lab publishes. The rest are marked "not established", which the page says means the checked vendor sources did not establish one, not that no cutoff exists.

"A blank means the checked vendor sources did not establish a cutoff for that model; it is not proof that the lab has never published one."

The page also reports a test of whether models reach for a search tool when they should. Over 2,000 calls across 16 models, each given a web search tool, frontier models decided correctly almost every time. Weaker ones answered settled questions from memory after the answer had changed, and searched the web for things like the boiling point of water. In one prompt about the king of Norway, the page says five models named a dead man.

Search does not fix this, the page argues, because the model forgets what it read when the session ends, and because the decision to search runs on the same weights that hold the stale fact.

The small-model counterpoint

Not every release in the same week is chasing scale. Cactus published Needle 3 on 18 September, a foundation model for tiny devices with 9 to 29 MB CQ2-bit binaries and a 29M to 121M parameter laddered architecture. Its documentation claims that the 4-layer subnetwork can match DeepSeek V4 Flash when tuned on downstream tasks for one epoch, and reports 400 to 4,000 tokens/s decode on a Raspberry Pi 5.

On 21 September, a project called mini-AGI published a continual learning byte-level language model trained from scratch on a single 8 GB VRAM GPU, reading one stream of data. Its README calls it "a small toy-level model" and says the weights are not published yet because the run is still on its first pass over the corpus. Its lowest held-out loss so far is 0.7417 nats, or 1.0700 bits per character, at 562.9M of 7,879M characters, 7.14% through the data.

The same day, Cua released CUA-S1 as a source-only research release: small, specialized "System 1" models for computer-use decisions such as choosing which value belongs in a field. The company describes it as an engineering analogy for fast, bounded decisions, not a replacement for a general-purpose agent's planning.

None of these three will top a benchmark table. That is the point. The release cycle now runs on several tracks at once, and the one with the fastest adoption is not the one with the best scores.

Comments 0

Sources

5
  1. 01Chinese AI models surge in global popularity — and Washington is worriedEN
  2. 02How stale is your AI? Release age and training cutoff for 20 modelsEN
  3. 03Needle 3 - 8-29 MB foundation model for tiny devicesEN
  4. 04mini-AGI: Continual learning model trained from scratch on 8GB VRAMEN
  5. 05CUA-S1 – A System One Model for Computer UseEN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.