Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
Zichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed Mian
Abstract
Text-driven Human-Object Interaction (Text-to-HOI) generation is an emerging field with applications in animation, video games, virtual reality, and robotics. A key challenge in HOI generation is maintaining interaction consistency in long sequences. Existing Text-to-Motion-based approaches, such as discrete motion tokenization, cannot be directly applied to HOI generation due to limited data in this domain and the complexity of the modality. To address the problem of interaction consistency in long sequences, we propose an autoregressive diffusion model (ARDHOI) that predicts the next continuous token. Specifically, we introduce a Contrastive Variational Autoencoder (cVAE) to learn a physically plausible space of continuous HOI tokens, thereby ensuring that generated human-object motions are realistic and natural. For generating sequences autoregressively, we develop a Mamba-based context encoder to capture and maintain consistent sequential actions. Additionally, we implement an MLP-based denoiser to generate the subsequent token conditioned on the encoded context. Our model has been evaluated on the OMOMO and BEHAVE datasets, where it outperforms existing state-of-the-art methods in terms of both performance and inference speed. This makes ARDHOI a robust and efficient solution for text-driven HOI tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He et al.CVPR 2026 · 14 citations
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion DiffusionLin Wu, Zhixiang Chen, Jianglin LanNeurIPS 2025 · 10 citations
- Pulp Motion: Framing-aware multimodal camera and human motion generationRobin Courant, Xi WANG, David Loiseaux, Marc Christie et al.ICLR 2026 · 8 citations
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction GenerationZichen Geng, Zeeshan Hayder, Wei Liu, Hesheng Wang et al.CVPR 2026 · 1 citation
- MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion GenerationWenfeng Song, Xuehan Wang, Shuai Li, Yi Chen et al.CVPR 2026
Builds on25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
Related papers
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 9 citations
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng et al.CVPR 2026 · 6 citations
- Disentangled Hierarchical VAE for 3D Human-Human Interaction GenerationZichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu et al.ICLR 2026 · 3 citations
- LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand DiffusionMuchen Li, Sammy Christen, Chengde Wan, Yujun Cai et al.CVPR 2025
- Executing your Commands via Motion Diffusion in Latent SpaceXin Chen, Biao Jiang, Wen Liu, Zilong Huang et al.CVPR 2023
