5,000 sandboxes a second: DeepSeek publishes DSec, its agent training infrastructure
A new DeepSeek paper signed by Liang Wenfeng describes DSec, an agent training system that produces more than 5,000 sandboxes a second, reaches 3 million in a single day and peaks at 380,000 running at once.

Large model training competes on compute. Agent training competes on environments. A new DeepSeek paper signed by Liang Wenfeng describes a system called DSec (DeepSeek Elastic Compute), built to mass-produce sandboxes for agent training. The published technical details say DSec generates more than 5,000 sandboxes a second, reaches 3 million in a day and peaks at 380,000 running at once.
The single cluster behind that scale is large too: roughly 160 nodes, 30,000 CPU cores and 250TB of memory. Those numbers show where the bottleneck now sits. Agent training is no longer limited by the GPU alone. Scheduling and environment supply hold it back.
Why is training an agent so much work? Pretraining runs on the GPU cluster itself: feed in data, compute gradients. An agent is different. It writes code in a sandbox, runs compilations, opens a browser, even installs an operating system. Every step changes the state of the environment, and any step can break it. So each training round needs a fresh, clean sandbox. The sandbox is used once and thrown away when the run ends.
The problem keeps coming back to infrastructure. These systems have to install a full operating system and toolchain into every sandbox at a rate of 5,000 a second. At the same time, they cannot let several hundred thousand concurrent sandboxes blow out the cluster's memory and CPU.
The paper also notes that different kinds of agent tasks place very different demands on the environment. An agent grinding through programming problems needs only a stateless function-calling environment. An agent working on SWE-bench needs a complete Linux user space. In security offense and defense and in computer-use scenarios, container-level isolation is not enough, and virtual machines are required. Training an agent to operate commercial software needs a full Windows or macOS with a graphical interface and drivers.
That also explains why improvements in agent capability depend more and more on infrastructure engineering. How environments are built, how sandboxes are scheduled, how resources are isolated: the answers decide whether an agent can work reliably in the real world.
For the domestic compute ecosystem, DSec points to one trend. As model capabilities converge, the engineering infrastructure for training and inference is becoming the factor that separates them.
Sources
2All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.