Debate2Create: Robot Co-design via Multi-Agent LLM Debate
Kevin Qiu, Marek Cygan
Abstract
We introduce Debate2Create (D2C), a multi-agent LLM framework that formulates robot co-design as structured, iterative debate grounded in physics-based evaluation. A design agent and control agent engage in a thesis-antithesis-synthesis loop, while criterion-specific LLM judges provide multi-objective feedback to steer exploration. Across five MuJoCo locomotion benchmarks, D2C achieves the highest default-normalized score among the evaluated LLM-based and black-box baselines, with gains up to 3.2x on Ant and nearly 9x on Swimmer. Iterative debate yields 18-35% gains over compute-matched zero-shot generation, and D2C-generated rewards transfer to default morphologies in 4/5 tasks. These results suggest that structured, simulator-grounded multi-agent interaction is a useful mechanism for joint morphology-reward optimization under a fixed-topology, per-candidate-RL protocol. Project page: debate2create.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Reflexion: language agents with verbal reinforcement learningNoah Shinn, Federico Cassano, Ashwin Gopinath, Karthik Narasimhan et al.NeurIPS 2023 · 5,828 citations
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan et al.NeurIPS 2023 · 4,972 citations
- Improving Factuality and Reasoning in Language Models through Multiagent DebateYilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum et al.ICML 2024 · 1,562 citations
- Teaching Large Language Models to Self-DebugXinyun Chen, Maxwell Lin, Nathanael Schärli, Denny ZhouICLR 2024 · 1,085 citations
- Eureka: Human-Level Reward Design via Coding Large Language ModelsYecheng Jason Ma, William Liang, Guanzhi Wang, De-An Huang et al.ICLR 2024 · 582 citations
Related papers
- ReLook: Vision-Grounded RL with a Multimodal LLM Critic for Agentic Web CodingYuhang Li, Chenchen Zhang, Ruilin Lv, Ao Liu et al.ACL 2026 · 8 citations
- MOTIF: Multi-strategy Optimization via Turn-based Interactive FrameworkNguyen Viet Tuan Kiet, Tung Dao, Cong Dao Tran, Huynh Thi Thanh BinhAAAI 2026 · 1 citation
- MAC-AMP: A Closed-Loop Multi-Agent Collaboration System for Multi-Objective Antimicrobial Peptide DesignGen Zhou, Sugitha Janarthanan, Lianghong Chen, Pingzhao HuICLR 2026 · 1 citation
- AutoMS: Multi-Agent Evolutionary Search for Cross-Physics Inverse Microstructure DesignZhenyuan Zhao, Yu Xing, Tianyang Xue, Lingxin Cao et al.ICML 2026
- RF-Agent: Automated Reward Function Design via Language Agent Tree SearchNing Gao, Xiuhui Zhang, Xingyu Jiang, Mukang You et al.NeurIPS 2025 · 8 citations
