DSec: DeepSeek shows how it builds 3 million environments a day for agents
In a paper signed by Liang Wenfeng, DeepSeek describes DSec: 5,000 new sandboxes per second, 3 million a day and 380,000 running at once.

Training an agent is not about feeding data to a network. The agent has to write code inside an environment, compile it, open a browser, and sometimes install an operating system. Every step changes the state of that environment, and it can break at any moment. So each training round needs a fresh, clean sandbox that gets thrown away after use. The Chinese outlet IT之家 describes DeepSeek's work on a system called DSec (DeepSeek Elastic Compute), signed by the company's co-founder, Liang Wenfeng.
Scale is the whole problem here. DSec produces more than 5,000 sandboxes per second, which comes to 3 million a day, and peak concurrent operation reaches 380,000 environments. The single cluster behind it holds about 160 nodes, 30,000 CPU cores and 250 TB of RAM. Resource oversubscription and dense packing let one node hold 3,200 containers or 800 micro virtual machines at once.
The problem: a system for each of 3 million sandboxes
The bottleneck is not scheduling itself but building the environment. The classic Docker approach packs a base image, a working directory and a set of tools into one complete image, and that works at small scale. In DSec the container backend used 11,266 base images and 102,171 working directories in total. Of all sandboxes, 67.8 percent need at least one extra layer with a working directory or a tool package on top of the base image. With that much variety, updating a single package would force a rebuild of every image that contains it.
The fix is to split the environment into three independently versioned, read-only layers in the EROFS format: base image, working directory and tool package, joined at sandbox startup by overlayfs. A package update then touches only its own layer. The second piece is a component called Chronus, which mediates communication from inside the sandbox to the training framework. The framework then knows what stage the agent is at and what feedback signal to send back.
Why this concerns every agent model
The V4.1-Flash model card confirms that this was the point of the training run. The post-training recipe stayed standard, and all the significant changes were in the data pipeline, in mass automatic synthesis of tasks and agent environments. In other words, the quality of an agent now depends on the environment factory in the background, not on another trick in the learning algorithm. DSec is DeepSeek's answer to the question of where to get 3 million such environments a day without building a data centre that chokes on itself.
Sources
2- 01DeepSeek 新论文公开 Agent 训练,梁文锋署名 (IT之家)ZH
- 02DeepSeek-V4.1-Flash — sekcja Post-training (data pipeline)EN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.