Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow
Prerit Gupta, Shourya Verma, Ananth Grama, Aniket Bera
Abstract
Generating realistic, context-aware two-person motion conditioned on diverse modalities remains a fundamental challenge for graphics, animation and embodied AI systems. Real-world applications such as VR/AR companions, social robotics and game agents require models capable of producing coordinated interpersonal behavior while flexibly switching between interactive and reactive generation. We introduce DualFlow, the first unified and efficient framework for multi-modal two-person motion generation. DualFlow conditions 3D motion generation on diverse inputs, including text, music, and prior motion sequences. Leveraging rectified flow, it achieves deterministic straight-line sampling paths between noise and data, reducing inference time and mitigating error accumulation common in diffusion-based models. To enhance semantic grounding, DualFlow employs a novel Retrieval-Augmented Generation (RAG) module for two-person motion that retrieves motion exemplars using music features and LLM-based text decompositions of spatial relations, body movements, and rhythmic patterns. We use contrastive rectified flow objective to further sharpen alignment with conditioning signals and add synchronization loss to improve inter-person temporal coordination. Extensive evaluations across interactive, reactive, and multi-modal benchmarks demonstrate that DualFlow consistently improves motion quality, responsiveness, and semantic fidelity. DualFlow achieves state-of-the-art performance in two-person multi-modal motion generation, producing coherent, expressive, and rhythmically synchronized motion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5eb558f-2820-421c-a337-d562041288f3Cited by top-tier papers2
- Sketch2Colab: Sketch-Conditioned Multi-Human Animation via Controllable Flow DistillationDivyanshu Daiya, Aniket BeraCVPR 2026
- Unified Number-Free Text-to-Motion Generation Via Flow MatchingGuanhe Huang, Oya ÇeliktutanCVPR 2026
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Retrieval Augmented Language Model Pre-TrainingKelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat et al.ICML 2020 · 2,937 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- Human Motion Diffusion as a Generative PriorYoni Shafir, Guy Tevet, Roy Kapon, Amit Haim BermanoICLR 2024 · 371 citations
Related papers
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng et al.CVPR 2026 · 6 citations
- MDD: A Dataset for Text-and-Music Conditioned Duet Dance GenerationPrerit Gupta, Jason Alexander Fotso-Puepi, Zhengyuan Li, Jay Mehta et al.ICCV 2025 · 1 citation
- Multi-Person Interaction Generation from Two-Person Motion PriorsWenning Xu, Shiyu Fan, Paul Henderson, Edmond S. L. HoSIGGRAPH 2025 · 2 citations
- Retrieving Semantics from the Deep: an RAG Solution for Gesture SynthesisMuhammad Hamza Mughal, Rishabh Dabral, Merel C. J. Scholman, Vera Demberg et al.CVPR 2025
- EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsBohong Chen, Yumeng Li, Yinglin Xu, Youyi Zheng et al.SIGGRAPH 2026
