Same score, different bill: how to read open-source model benchmarks
On the Artificial Analysis Intelligence Index, Xiaomi's MiMo-V2.6-Pro and the closed-source Grok 4.7 both score 46. But one task costs about 29 times more on Grok 4.7. Cost is turning into a benchmark dimension of its own.

Read an open-source model launch post and the easiest thing to miss is what sits behind the score. Xiaomi's MiMo-V2.6-Pro scores 46 on the Artificial Analysis Intelligence Index, the best result among open-source models. Grok 4.7, a closed-source model released the same day, also scores 46 on that index. On that one line, the two look evenly matched.
Change the dimension and the picture shifts. Artificial Analysis measured Grok 4.7's high-end configuration at an average of 3.74 US dollars per task. MiMo-V2.6-Pro averages 0.13 US dollars per task, a gap of about 29 times. The same score can come with very different usage costs.
The second thing easily missed is the exact test setup. MiMo-V2.6's results come from reinforcement learning training with only 30 steps. Each step holds 1 568 prompts, and the model tries 16 different routes per question. On DeepSWE v1.1, which tests coding agents, Pro climbed from 58.4 to 72.57 and Flash from 48.7 to 65.68. Those numbers only mean something if out-of-sample generalization holds. How many points the model racked up on the training set says little.
The third variable is the choice of evaluation itself. In the notes for V4.1 Flash, the team says that after internal and external testing the model beats V4 Pro on performance, cost, speed and total time. But "beats across the board" is the team's own wording. The test methods, sample sizes and evaluation tools decide how far those conclusions travel.
So the right way to read a benchmark is to look at three things together: the conditions the model ran under, the cost per completed task, and how it performs out of sample. Cite only the highest score and you swap a point for a surface.
For companies picking a model, the more practical move is a small comparison on your own tasks. Fix the evaluation framework, compare cost per task against success rate, then decide which model goes into production. The score is the entry point. The bill is the conclusion.
Sources
3All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.