Skip to content
World clockEU--:--UK--:--USA--:--CN--:--PLDEFRIT中文EN

portal about AI and technologyevents · analysis · interviews · technical background

Search
LIVE
›

Terminal-Bench and DeepSWE: V4.1 Flash beats the frontier on agentic tasks

In the model card's table, V4.1-Flash scores 90.6 on Terminal-Bench 2.1 and 74.2 on DeepSWE v1.1, matching or beating Opus-5.0, GPT-5.6 Sol, K3 and GLM-5.3.

AI & modelsAnalysisRachel NwosuPublished: 25 September 20267 min readSources 2
Terminal-Bench and DeepSWE: V4.1 Flash beats the frontier on agentic tasks

The DeepSeek-V4.1-Flash model card carries a comparison table. It lines the model up against Opus-5.0, GPT-5.6 Sol, K3 and GLM-5.3, plus two earlier members of the family, DeepSeek-V4-Pro and DeepSeek-V4-Flash. Every measurement was taken at maximum reasoning effort.

On reasoning, V4.1-Flash scores 90.9 on GPQA Diamond. The rivals sit a little higher: GPT-5.6 Sol at 94.1, Opus-5.0 at 93.4, K3 at 92.9. GLM-5.3 trails at 88.1. On Codeforces the model earns a rating of 3,471, well clear of DeepSeek-V4-Pro (3,348) and DeepSeek-V4-Flash (3,289).

Where agents matter

Agentic tasks flip the picture. On Terminal-Bench 2.1, V4.1-Flash takes 90.6 points, the highest in the table. Opus-5.0 follows at 89.1, then GPT-5.6 Sol at 88.8, K3 at 88.3 and GLM-5.3 at 88.2. On DeepSWE v1.1, which measures how well a model solves real programming tasks, V4.1-Flash reaches 74.2. Opus-5.0 scores 74.0, GPT-5.6 Sol 73.0, K3 67.5 and GLM-5.3 66.9.

Newer versions of the test tell a harsher story. On Terminal-Bench 3.0 the model scores 30.0, against 43.3 for Opus-5.0 and 34.4 for GPT-5.6 Sol. On Terminal-Bench 4.0 it scores 31.2, against 51.8 for Opus-5.0 and 39.9 for GPT-5.6 Sol. On these two harder sets the American models lead by more than a dozen points.

The harness moves the result by 8 points

DeepSeek also publishes V4.1-Flash results across different harnesses, the agent scaffolds that wrap the model. On DeepSWE v1.1 the same model scores 74.2 in the mini-SWE harness, 72.6 in DSH Minimal, 70.5 in DSH Standard and 67.6 in DSH PTC. External harnesses give 69.8 in Claude Code, 66.2 in Pi, 65.6 in Codex and 65.5 in OpenCode. That is a spread of nearly 9 points with identical weights. An agent's score today depends on the model and on the layer wrapping it.

In other categories V4.1-Flash scores 88.1 on CyberGym and 54.8 on AutomationBench, and 63.9 on HLE with tools. The same table also shows a generational jump inside the family. On Terminal-Bench 4.0, DeepSeek-V4-Pro had 12.4 and V4-Flash 7.0. V4.1-Flash raised that to 31.2, more than double, with 8 billion active parameters at input.

The independent service Artificial Analysis reports two modes for the same model: maximum reasoning with an intelligence index of 39, and a non-reasoning mode with an index of 25. Two entirely different methodologies side by side show how carefully any single ranking has to be read.

Comments 0

Sources

2
  1. 01DeepSeek-V4.1-Flash — wyniki oceny i tabele porównawczeEN
  2. 02DeepSeek V4.1 Flash — indeks inteligencji i prędkość (Artificial Analysis)EN

All figures and quotations in this text come from the sources listed below.

Content prepared by the editorial team with AI assistance.

Rachel Nwosu

Rachel Nwosu

AI, models and technology

Rachel Nwosu covers AI, models and technology for FLASH24, working from public model documentation, benchmark releases and repository histories rather than press summaries, and she skips announcements that arrive without reproducible numbers. She checks training-data claims against dataset cards and reruns reported metrics where code is available. She spends much of her week interviewing researchers and engineers, tracking model launch calendars, and comparing vendor benchmarks with independent evaluations. Outside the desk she runs 3D printers, restores old computers, and tests how models learn from internet junk. She does not publish benchmark figures she cannot trace to a source.

Newsroom →

Comments

0
  1. No comments yet — be the first.

Write a comment

Comments are public. We do not publish abuse, spam or advertising.