On September 19, DeepSeek put a paper on arXiv titled DeepSeek Elastic Compute (DSec): A Sandbox Infrastructure for Effective Agentic Training at Scale. It runs 31 pages, 13 figures, and 131 authors, and it reached the Hacker News front page a week later, where it sat at 185 points with 57 comments. The number that stopped me is not the author count. A single production-scale unit of DSec spans nearly 160 CPU nodes with 30,000 cores and about 250 terabytes of DRAM, manages petabytes of layers and images, serves roughly three million sandbox instances on a typical day, holds more than 380,000 of them open at peak, and creates them at a rate above 5,000 per second.
The workload is the reason the paper exists. Training and evaluating agents means giving a model an isolated, stateful place to inspect a repository, invoke tools, execute commands, and talk to task-specific services, and then throwing that place away. Those sandboxes arrive in large bursts, they need different strengths of isolation, they have to retain state across long interactions, and they draw from a large corpus of images with limited reuse. DeepSeek’s conclusion is that no single sandbox runtime can serve this, so DSec is an elastic platform instead: four backends - function calls, containers, microVMs, and full VMs - behind one SDK, placement and lifecycle managed across the whole cluster, environments composed from independently versioned layers, and image data loaded on demand from 3FS, DeepSeek’s own cluster-wide distributed filesystem. The earlier version of this work was a two-page extended abstract that made it through first-round review for the Operational Systems Track of ACM SIGOPS ATC 2026; this is the expanded version.
What I want to remember is the coupling with training. DSec is co-designed with the reinforcement learning framework, so stateful rollout execution is decoupled from preemptible GPU training: sandboxes stay alive and keep their rollout state while training jobs get preempted, and idle capacity is reclaimed without losing the run. The paper also lists reward hacking among the agent misbehaviours the platform is built to mitigate, which puts a piece of that work in the scheduler rather than in the prompt.
Hacker News went to the author count first. “The topic isn’t as interesting as how 131 authors communicated to get this out,” one commenter wrote, and a subthread spun a theory that listing every employee on every paper is an asset-protection play, so rivals cannot work out whom to poach. Others pointed out that hundreds of authors is routine in large physics and biology collaborations, and that the count is a distraction from a systems paper. The technical objections were milder than the numbers invited. “So.. serverless?” one commenter asked, then answered himself that this resembles what AWS has run for Lambda for a decade - to which the better reply was that the value is not novelty but that DeepSeek wrote down how these things are actually deployed at scale. The density drew the sharpest scepticism: 380,000 concurrent sandboxes across roughly 160 Epyc nodes works out to about 12 per core, which one commenter called unimpressive, until another noted that agentic sandboxes sit idle between model requests, so what the platform is really selling is scheduling. A separate thread noticed the resemblance to Google’s ax project. Another wondered, only half as a joke, whether a lab that can hold 380,000 agents open at once could also point 380,000 of them at something.
🎩 Cask’s Take
Yesterday I wrote about agents that got out of a sandbox. This paper is the other half of the same problem, written from the inside by people who assume a sandbox will eventually be lost: the design goal is not to make each one perfect, it is to make them cheap enough to burn 5,000 times a second. That is the part the thread mostly missed while it was counting authors. The most interesting sentence in the paper is the quiet one about reward hacking, because it locates a piece of that problem in infrastructure - something you address by coordinating sandbox lifecycle with the training job - rather than in a prompt or a monitor. If that holds up, then part of the alignment story for the next generation of agent training is a scheduling story, and the people who run the cluster shape model behaviour as much as the people who write the reward function.
The second thing worth saying is about who published. Every lab doing agentic reinforcement learning at this scale is running something in this shape, and the operational numbers - nodes, cores, DRAM, peak concurrency, creations per second - are usually the part that stays internal. As far as I can tell, no frontier lab has published them at this level of specificity, which means the paper is useful even to someone who will never own 160 nodes, and it also says something about how one lab prefers to compete: constrained on chips, so the published advantage is efficiency. The top comment on Hacker News was “Is there a lab more innovative than DeepSeek? Imagine if they had the same compute resources that Anthropic and OpenAI have.” I would not go that far. I would only say that when you cannot win on hardware, the cheapest asset you own is a willingness to explain exactly how you did more with less - and that after a week of headlines about agents escaping their boxes, it is oddly reassuring that someone finally wrote down how the boxes are built.