Chinese open models take the majority of tokens on two developer gateways
Chinese AI models went from a small minority to a majority of tokens on OpenRouter and Vercel in 2026, according to usage data shared with CNBC, while Washington opens investigations into the shift.

Two developer gateways route traffic to many model providers. Both now send most of their tokens to Chinese labs. On OpenRouter, Chinese models accounted for 57% to 67% of tokens used in the week of Sept. 14, up from 6% to 13% in February, according to usage data shared with CNBC. On Vercel, the share rose to 55% in August from 11% in January. CNBC reported the figures on 26 September.
The numbers describe routing, not benchmark scores. They are still the clearest adoption signal published this quarter.
OpenRouter's data covers companies in the U.S., Europe and what it defines as the "Global South", 82 countries across Central and South America, Africa and Asia. Vercel did not specify a geographic breakdown for its figure. About half the tokens on OpenRouter are used by companies in the U.S., so the platform's overall mix is not driven by one region alone.
Price, then capability
Peter Walker, head of insights at OpenRouter, told CNBC that Chinese open source models released this year "can credibly perform in advanced agentic use cases, especially in regards to coding, in a way that was just not true in late 2025." He added that they are "incredibly cost-effective compared to most models from American labs." Harpreet Arora, head of agentic infrastructure at Vercel, put the same point in ordering terms: "Chinese models are becoming capable enough for more tasks at a much lower cost. Once a model meets the quality bar for the job, that price difference becomes compelling." Arora also said companies still want frontier U.S. models for some harder tasks. That qualifier matters. The CNBC report notes that U.S. frontier models still attract more overall spending, even where token share has flipped. Token counts reward cheap, high-volume work such as code generation and agent loops. Spending rewards the expensive calls that customers reserve for tasks they cannot afford to get wrong.
In the "Global South" group that OpenRouter tracks, more than two-thirds of tokens, 67%, went to Chinese models in recent weeks. CNAS senior fellow Daniel Remler told CNBC that Southeast Asia in particular may see significant uptake because of economic and cultural links with China and growing digital infrastructure. "Anywhere from Lagos to São Paulo to Jakarta where entrepreneurs and governments are looking for cheap, open models, will look first to Chinese AI," he said.
Anthropic and OpenAI both announced cheaper models earlier that week. Dianne Penn, who leads product management for research and labs at Anthropic, told CNBC the company was trying to make answers "more efficient, so it uses less tokens depending on your effort setting."
Washington opens files
Two U.S. House committees are investigating the impact of rising adoption of Chinese models, according to the same report. The stated concerns are technology competition, security and Beijing's global influence. The U.S. has tried to protect its lead by restricting Chinese AI companies from buying the most advanced chips through export controls. Washington is also worried about remote access to Nvidia chips through overseas data centers, and about "distillation", where new models are trained to mimic older, established ones. Remler framed the risk in alignment terms rather than raw performance. Chinese AI represents "real economic and security risks for the United States," he told CNBC, and "the ultimate concern is that the integration of Chinese AI models pulls countries into a Chinese technology sphere of influence that hardens into geopolitical alignment." AI was a major focus when U.S. President Donald Trump and Chinese President Xi Jinping met this week, CNBC reported.
The adoption data sits alongside a separate set of claims about model freshness that is harder to verify and easier to get wrong. A Show HN page published on 16 September, stale.jock.pl, lists release dates and training cutoffs for 20 models across 8 labs, and counts upward from each date live. It says 10 of the 20 models have a cutoff their lab actually publishes, and that 5 of 8 labs, Anthropic, Google DeepMind, Meta, OpenAI and xAI, have published a cutoff for at least one model on the list. The page notes that a blank means the checked vendor sources did not establish a cutoff, not proof that none exists.
Its own example is the cleanest illustration of the gap. GPT-6 Astra, per that page, was released on 3 September 2026 with training data stopping at 30 April 2026, roughly four months behind on launch day. The same page reports a test in which 16 models were each given a web search tool across more than 2,000 calls: frontier models decided correctly almost every time, while weaker ones stated settled facts that had changed without checking, and searched the web for things like the boiling point of water.
That is a caveat for anyone reading adoption numbers as a quality verdict. Token share measures what developers route. It does not measure whether the model knows what day it is.
Three other releases this month point in different directions. Cactus Compute published Needle 3 on 18 September, a family of models it describes as producing 9 to 29 MB CQ2-bit binaries, with a 4-layer subnetwork claimed to match DeepSeek V4 Flash when tuned on downstream tasks for one epoch; the company says it trained on 360B tokens of a proprietary structured dataset and reports 400 to 4,000 tokens/s decode on a Raspberry Pi 5. Cua published CUA-S1 on 19 September, an MIT-licensed source-only research release of small models for computer-use decisions, with weights hosted separately on Hugging Face. And on 21 September, a project called mini-AGI showed a continual-learning byte-level model training on a single 8 GB VRAM GPU, with weights not yet published and the run 7.14% through its corpus.
None of those three moves the OpenRouter or Vercel numbers. They do show where the smaller end of the market is pushing: device-local inference, bounded decisions and training runs that fit on hardware a single developer owns. Whether that changes the token mix is a question for the next set of gateway figures, not this one.
Sources
5- 01Chinese AI models surge in global popularity, and Washington is worriedEN
- 02How stale is your AI? Release age and training cutoff for 20 modelsEN
- 03Needle 3: 8-29 MB foundation model for tiny devicesEN
- 04CUA-S1: A System One Model for Computer UseEN
- 05mini-AGI: continual learning model trained on 8GB VRAMEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.