RoboFactory: Exploring Embodied Agent Collaboration with Compositional Constraints
Yiran Qin, Li Kang, Xiufeng Song, Zhenfei Yin, Xiaohong Liu, Xihui Liu, Ruimao Zhang, Lei Bai
Abstract
Designing effective embodied multi-agent systems is critical for solving complex real-world tasks across domains. Due to the complexity of multi-agent embodied systems, existing methods fail to automatically generate safe and efficient training data for such systems. To this end, we propose the concept of compositional constraints for embodied multi-agent systems, addressing the challenges arising from collaboration among embodied agents. We design various interfaces tailored to different types of constraints, enabling seamless interaction with the physical world. Leveraging compositional constraints and specifically designed interfaces, we develop an automated data collection framework for embodied multi-agent systems and introduce the first benchmark for embodied multi-agent manipulation, RoboFactory. Based on RoboFactory benchmark, we adapt and evaluate the method of imitation learning and analyzed its performance in different difficulty agent tasks. Furthermore, we explore the architectures and training strategies for multi-agent imitation learning, aiming to build safe and efficient embodied multi-agent systems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54dfc8a4-28c2-4b80-b019-7d6ba10e8d16Cited by top-tier papers4
- RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for RoboticsEnshen Zhou, Jingkun An, Cheng Chi, Yi Han et al.NeurIPS 2025 · 159 citations
- Heterogeneous Agent Q-weighted Policy OptimizationBor-Jiun Lin, Chun-Yi LeeICLR 2026 · 102 citations
- GauDP: Reinventing Multi-Agent Collaboration through Gaussian-Image Synergy in Diffusion PoliciesZiye Wang, Li Kang, Yiran Qin, Jiahua Ma et al.NeurIPS 2025 · 5 citations
- MangoBench: A Benchmark for Multi-Agent Goal-Conditioned Offline Reinforcement LearningYi Wang, Ningze Zhong, Zhiheng Fu, Longguang Wang et al.CVPR 2026
Builds on12
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- CAMEL: Communicative Agents for "Mind" Exploration of Large Language Model SocietyGuohao Li, Hasan Hammoud, Hani Itani, Dmitrii Khizbullin et al.NeurIPS 2023 · 1,975 citations
- WebArena: A Realistic Web Environment for Building Autonomous AgentsShuyan Zhou, Frank F. Xu, Hao Zhu, Xuhui Zhou et al.ICLR 2024 · 1,197 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- ExpeL: LLM Agents Are Experiential LearnersAndrew Zhao, Daniel Huang, Quentin Xu, Matthieu Lin et al.AAAI 2024 · 484 citations
Related papers
- Moving Out: Physically-grounded Human-AI CollaborationXuhui Kang, Sung-Wook Lee, Haolin Liu, Yuyan Wang et al.ICML 2026
- Rethinking Video Generation Model for the Embodied WorldYufan Deng, Zilin Pan, Hongyu Zhang, Xiaojie Li et al.ICML 2026 · 24 citations
- AutoCGP: Closed-Loop Concept-Guided Policies from Unlabeled DemonstrationsPei Zhou, Ruizhe Liu, Qian Luo, Fan Wang et al.ICLR 2025
- WebFactory: Automated Compression of Foundational Language Intelligence into Grounded Web AgentsSicheng Fan, Qingyun Shi, Shengze Xu, Shengbo Cai et al.ICLR 2026 · 7 citations
- CraftFactory: A Conditioned Control Policy Benchmark for Compositional GeneralizationJinbing Hou, Youpeng Zhao, Jian ZhaoAAAI 2025
