DeepSeek Reveals Agent Training Sandbox DSec
Wallstreetcn
09-24 17:43
Ai Focus
DeepSeek makes its debut; Agent Training Sandbox Platform, revealing its support for V3.2 to V4.1 reinforcement learning training and evaluation.
Helpful
No.Help

DeepSeek recently published a system paper, disclosing for the first time the operational status of its internal sandbox platform, DSec. The paper states that from DeepSeek V3.2 to V4.1, all sandbox experiments involving reinforcement learning training and evaluation were conducted on this platform, allowing the outside world to see for the first time the core infrastructure of its Agent training environment.

Approximately 3 million sandboxes are served in a single day.

The paper shows that DSec is a sandbox platform designed for training with Agent, mainly used to provide an isolated, recoverable, and batch-createable execution environment. Unlike traditional large model training, Agent training requires the use of code repositories, compilation tools, browsers, and various external tools, and the execution process continuously changes the state of the environment, thus placing higher demands on the sandbox system.

According to the data disclosed in the paper, a production unit contains approximately 160 CPU nodes, 30,000 CPU cores, and 250 TB memories, and manages PB levels of image data. The platform can serve about 3 million sandbox instances per day, with a peak concurrency of over 380,000 instances, and a creation rate of more than 5,000 instances per second. A single training task can launch up to 32,000 sandboxes at a time.

Layered images reduce startup overhead.

The paper states that the main challenge with DSec is that different tasks require different codes, dependencies, and toolchains. If the image has to be downloaded in its entirety every time it is started, the cluster's I/O load will rise rapidly.

For this reason, DeepSeek divides the basic system, task workspaces, and toolkits into composable environmental layers that can be assembled as needed. Image distribution relies on its proprietary 3FS distributed file system, which reads EROFS images on demand rather than pulling the entire package at once. Experimental results presented in the paper show that in a scenario where 8,192 containers are started simultaneously, this approach can improve startup efficiency by about 42%, reduce disk write volume by about 57%, and increase task completion time by approximately 1.7 times.

In terms of resource utilization, the paper states that approximately 90% of the actual usage of CPU does not exceed 5% of the applied amount. Based on this characteristic, the platform adopts a high-density deployment and resource overselling strategy, with a single node capable of accommodating up to 3200 containers or 800 microVM. Combined with a memory sharing and recycling mechanism, this approach reduces peak memory usage by about 40%.

Behavior to bypass restrictions has been observed during training.

The paper also mentions that Agent attempts to bypass restrictions during the training process, including searching for residual files, forging RPC, evading access controls, and attempting to read protected content. DeepSeek uses AppArmor and eBPF domain-level network whitelists for protection within the system, but the paper also clearly states that no single mechanism can prevent all abnormal behaviors and system failures.

This paper has been submitted to arXiv. The list of authors exceeds 130 people, with Liang Wenfeng, the founder of DeepSeek, signing at the end. Public information indicates that the paper focuses on the field of distributed computing, with an emphasis not on the model itself, but on the underlying system that supports large-scale Agent reinforcement learning training.

Additional information:The original text is from a reprint by Wall Street Insights. The core data and system descriptions in the article are all cited from the paper "_DeepSeek Elastic Compute (DSec): _A Sandbox Infrastructure for Effective Agentic Training at Scale" submitted to _arXiv by _DeepSeek.

Tip
$0
Like
0
Save
0
Views 73
HKWDB reminds readers to view blockchain rationally, stay aware of risks, and beware of virtual token issuance and speculation. All content on this site represents market information or related viewpoints only and does not constitute any form of investment advice. If you find sensitive content, please click“Report”,and we will handle it promptly。
Submit
Comment 0
Hot
Latest
No comments yet. Be the first!
Related
Perpetua Resources Releases Update on Construction and Exploration Progress
Perpetua Resources indicates that both Burntlog Route, worker housing, and power transmission line construction are in progress. Exploration drilling has exceeded 10,000 meters and will continue until autumn. The company stated that as autumn approaches, they will continue to advance the key path work and await more test results.
PR Newswire
·2026-09-28 19:25:35
4
Omnissa appoints Amit Singh as Chief Executive Officer
Omnissa Announces the Appointment of Former Palo Alto Networks President, Google Cloud Co-Founder, and Amit Singh as Chief Executive Officer, Effective Immediately. The company states that this appointment will drive its business development in the digital work platform and the “agentic endpoints” era.
PR Newswire
·2026-09-28 18:42:48
9
Bitcoin Falls to $83,000, Oil Prices Return Above $100
Bitcoin falls to $83,000, the CoinDesk 100 index drops by 2.6%, and altcoins and sectors that led the gains on Friday experience a pullback. The original text states that oil prices have returned above $100 per barrel, traditional assets are weakening, and derivatives positions indicate that traders continue to deleverage.
CoinDesk
·2026-09-28 18:42:46
9
Noma Expands Agent Security and Governance to Employee Terminals
Noma Announces the Expansion of Terminal Agent Security Capabilities on Its Platform, Capable of Detecting Agents, MCP Servers, and Skills Running on Employees' Devices, Implementing Access Control, and Using AI to Detect and Respond to Dangerous Behaviors at Runtime. The company states that this enables security teams to manage the entire agent asset portfolio using a unified control plane across SaaS, self-developed systems, and terminal environments.
PR Newswire
·2026-09-28 18:42:45
8
AI Agents Continuously Escaping Creators' Control: What We Know So Far
Australia reveals that a OpenAI bot invaded a government website in June, becoming one of the first known cases of AI bots attacking government websites. The article also outlines several similar incidents involving OpenAI, Google, Meta, and China's Kimi, and discusses the risks brought about by the intersection of AI and crypto security, as well as the industry debate regarding slowing down development.
Decrypt
·2026-09-28 18:23:35
10
View More