DOCKSMITH: Scaling Reliable Coding Environments via an Agentic Docker Builder
Jiaran Zhang, Lu Ma, Yanhao Li, Fanqi Wan, DI QI, Xin Wu, Zhewei Huang, Liangyu Chen, YINGWEI MA, Qi Han, Xiangyu Zhang
Abstract
Reliable Docker-based environment construction is a dominant bottleneck for scaling execution-grounded training and evaluation of software engineering agents. We introduce DockSmith, a specialized agentic Docker builder designed to address this challenge. DockSmith treats environment construction not merely as a preprocessing step, but as a core agentic capability that exercises long-horizon tool use, dependency reasoning, and failure recovery, yielding supervision that transfers beyond Docker building itself. DockSmith is trained on large-scale, execution-grounded Docker-building trajectories produced by a SWE-Factory--style pipeline augmented with a loop-detection controller and a cross-task success memory. Training a 30B-A3B model on these trajectories achieves open-source state-of-the-art performance on Multi-Docker-Eval, with 39.72% Fail-to-Pass and 58.28% Commit Rate. Moreover, DockSmith improves out-of-distribution performance on SWE-bench Verified, SWE-bench Multilingual, and Terminal-Bench 2.0, demonstrating the broader agentic benefits of environment construction. Our model and Docker-building trajectories are publicly available at https://huggingface.co/JiaranZhang/DockSmith.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b325ffe6-28f9-4dce-9b2a-ca12f94168b9Builds on7
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Repo2Run: Automated Building Executable Environment for Code Repository at ScaleRuida Hu, Chao Peng, Xinchen Wang, Junjielong Xu et al.NeurIPS 2025 · 49 citations
- You Name It, I Run It: An LLM Agent to Execute Tests of Arbitrary ProjectsIslem Bouzenia, Michael PradelISSTA 2025 · 17 citations
- Less is More? An Empirical Study on Configuration Issues in Python PyPI EcosystemYun Peng, Ruida Hu, Ruoke Wang, Cuiyun Gao et al.ICSE 2024 · 5 citations
Related papers
- Large-Scale Terminal Agentic Trajectory Generation from Dockerized EnvironmentsSiwei Wu, Yizhi Li, Yuyang Song, Wei Zhang et al.ICML 2026 · 16 citations
- SWE-rebench V2: Language-Agnostic SWE Task Collection at ScaleIbragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov, Aleksandr GolubevICML 2026 · 13 citations
- MEnvAgent: Scalable Polyglot Environment Construction for Verifiable Software EngineeringChuanzhe Guo, Jingjing Wu, Sijun He, Yang Chen et al.ICML 2026 · 3 citations
- Process-Level Trajectory Evaluation for Environment Configuration in Software Engineering AgentsJiayi Kuang, Yinghui Li, Xin Zhang, Yangning Li et al.ICLR 2026 · 18 citations
- SWE Data Construction, Automatically!Lianghong Guo, Yanlin Wang, Caihua Li, Wei Tao et al.FSE 2026
