Text2Interact: High-Fidelity and Diverse Text-to-Two-Person Interaction Generation
Qingxuan Wu, Zhiyang Dou, chuan guo, Yiming Huang, Qiao Feng, Bing Zhou, Jian Wang, Lingjie Liu
Abstract
Modeling human–human interactions from text remains challenging because it requires not only realistic individual dynamics but also precise, text-consistent spatiotemporal coupling between agents. Currently, progress is hindered by 1) limited two-person training data, inadequate to capture the diverse intricacies of two-person interactions; and 2) insufficiently fine-grained text-to-interaction modeling, where language conditioning collapses rich, structured prompts into a single sentence embedding. To address these limitations, we propose our Text2Interact framework, designed to generate realistic, text-aligned human–human interactions through a scalable high-fidelity interaction data synthesizer and an effective spatiotemporal coordination pipeline. First, we present InterCompose, a scalable synthesis-by-composition pipeline that aligns LLM-generated interaction descriptions with strong single-person motion priors. Given a prompt and a motion for an agent, InterCompose retrieves candidate single-person motions, trains a conditional reaction generator for another agent, and uses a neural motion evaluator to filter weak or misaligned samples—expanding interaction coverage without extra capture. Second, we propose InterActor, a text-to-interaction model with word-level conditioning that preserves token-level cues (initiation, response, contact ordering) and an adaptive interaction loss that emphasizes contextually relevant inter-person joint pairs, improving coupling and physical plausibility for fine-grained interaction modeling. Extensive experiments show consistent gains in motion diversity, fidelity, and generalization, including out-of-distribution scenarios and user studies. Code will be released at github.com/Qingxuan-Wu/Text2Interact.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Next-Scale Autoregressive Models for Text-to-Motion GenerationZhiwei Zheng, Shibo Jin, Lingjie Liu, Mingmin ZhaoCVPR 2026 · 6 citations
- InterAgent: Physics-based Multi-agent Command Execution via Diffusion on Interaction GraphsBin Li, Ruichi Zhang, Han Liang, Jingyan Zhang et al.CVPR 2026 · 4 citations
- Stability-Driven Motion Generation for Object-Guided Human-Human Co-ManipulationJiahao Xu, Xiaohan Yuan, Xingchen Wu, Chongyang Xu et al.CVPR 2026
- Unified Number-Free Text-to-Motion Generation Via Flow MatchingGuanhe Huang, Oya ÇeliktutanCVPR 2026
- Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion GenerationKe Fan, Jiangning Zhang, Ran Yi, Jingyu Gong et al.CVPR 2026
Related papers
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng et al.CVPR 2026 · 6 citations
- InterMask: 3D Human Interaction Generation via Collaborative Masked ModelingMuhammad Gohar Javed, Chuan Guo, Li Cheng, Xingyu LiICLR 2025
- SINC: Spatial Composition of 3D Human Motions for Simultaneous Action GenerationNikos Athanasiou, Mathis Petrovich, Michael J. Black, Gül VarolICCV 2023 · 69 citations
- InterControl: Zero-shot Human Interaction Generation by Controlling Every JointZhenzhi Wang, Jingbo Wang, Yixuan Li, Dahua Lin et al.NeurIPS 2024 · 27 citations
- MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention GuidanceNathan Sala, Ofir Abramovich, Ariel Shamir, Daniel Cohen-Or et al.SIGGRAPH 2026
