InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction Graphs
Bin Li, Ruichi Zhang, Han Liang, Jingyan Zhang, Juze Zhang, Xin Chen, Lan Xu, Jingyi Yu, Jingya Wang
Abstract
Humanoid agents are expected to emulate the complex coordination inherent in human social behaviors. However, existing methods are largely confined to single-agent scenarios, overlooking the physically plausible interplay essential for multi-agent interactions. To bridge this gap, we propose InterAgent, the first end-to-end framework for textdriven physics-based multi-agent humanoid control. At its core, we introduce an autoregressive diffusion transformer equipped with multi-stream blocks, which decouples proprioception, exteroception, and action to mitigate crossmodal interference while enabling synergistic coordination. We further propose a novel interaction graph exteroception representation that explicitly captures fine-grained joint-tojoint spatial dependencies to facilitate network learning. Additionally, within it we devise a sparse edge-based attention mechanism that dynamically prunes redundant connections and emphasizes critical inter-agent spatial relations, thereby enhancing the robustness of interaction modeling. Extensive experiments demonstrate that InterAgent consistently outperforms multiple strong baselines, achieving state-of-the-art performance. It enables producing coherent, physically plausible, and semantically faithful multiagent behaviors from only text prompts. Our code and data will be released to facilitate future research. Project page: https://binlee26.github.io/InterAgent-Page.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7b2e3a25-200f-4892-9e4f-be2ada630415Cited by top-tier papers2
- TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team SizeStefan Lionar, Gim Hee LeeCVPR 2026 · 3 citations
- MoLingo: Motion-Language Alignment for Text-to-Human Motion GenerationYannan He, Garvita Tiwari, Xiaohan Zhang, Pankaj Bora et al.CVPR 2026 · 2 citations
Builds on50
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng et al.CVPR 2026 · 6 citations
- DiffE2E: Rethinking End-to-End Driving with a Hybrid Diffusion-Regression-Classification PolicyRui Zhao, Yuze Fan, Ziguo Chen, Fei Gao et al.NeurIPS 2025 · 7 citations
- Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction GenerationQingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang et al.ICLR 2026 · 10 citations
- SPREAD: Spatial-Physical REasoning via geometry Aware DiffusionMinzhang Li, Kuixiang Shao, Xuebing Li, Yuyang Jiao et al.CVPR 2026
- OneHOI: Unifying Human-Object Interaction Generation and EditingJiun Tian Hoe, Weipeng Hu, Xudong Jiang, Yap-Peng Tan et al.CVPR 2026
