Chinese open models: the price advantage is real, but it is not enough
Chinese open-weight models are closing the performance gap at lower cost. But the real cost per completed task can rise, and the American ecosystem remains strong on cloud and proprietary data.

Chinese open-weight models have reshaped the artificial intelligence market over the past two years. An analysis by Banque Lombard Odier, reported by CorCom, acknowledges that these models have narrowed the performance gap with Western alternatives at a markedly lower cost. It also warns that their advantages are partly overstated.
The cost per task, not per token
Price per token does not measure the real cost. Distillation and picking the model best suited to a single task can inflate the cost per completed activity, and the time it takes to finish it, cancelling out part of the initial saving. Falling prices tend to widen adoption rather than reduce overall spending. This is the Jevons paradox: a cheaper resource gets consumed in greater quantities.
In the analysis, the United States still holds a solid position thanks to cloud infrastructure and proprietary data, while frontier models continue to justify a price premium on complex tasks. On the hardware side, American dominance in advanced accelerators remains the main source of advantage. China compensates with an abundance of energy and with the speed at which it brings available computing power into service.
A concrete case
The most recent example comes from DeepSeek. The V4.1 Flash model went live on 10 September 2026 on China's national supercomputer network and is accessible via API from the model's service page. It uses a Mixture-of-Experts architecture with 552 billion parameters, an asymmetric Causal-Encoder-Decoder structure, 8 billion parameters activated on input and 16 billion on output. The official specification sheet reports a reduction in cache memory to a quarter of the HBM requirement and to an eighth of the SSD requirement compared with the previous generation. It lists a cache of 890 bytes per token and a context of up to one million tokens, released under the MIT licence and in the FP8, BF16, GPTQ, AWG and GGUF quantisation formats.
"The cheapest model is not the one that costs least per token, but the one that solves the problem on the first attempt" is the operating rule that emerges from the analysis.
What remains to be done
For companies the consequence is methodological: measure costs per result rather than per volume, evaluate the model on the specific workload rather than on general benchmarks, and keep portability between suppliers so you are not locked in when prices change. From this point of view, the availability of open-weight models is insurance against the risk of dependency.
Sources
3- 01CorCom — AI, la Cina sfida i big USA sui prezzi ma l'ecosistema americano tieneIT
- 02IT之家 — DeepSeek V4.1 Flash sulla rete nazionale dei supercomputerZH
- 03DeepSeek — V4.1 Flash release notes (architettura, cache, licenza MIT)EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.