Auto-Regressive Diffusion for Generating 3D Human-Object Interactions
Zichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed Mian
摘要
Text-driven Human-Object Interaction (Text-to-HOI) generation is an emerging field with applications in animation, video games, virtual reality, and robotics. A key challenge in HOI generation is maintaining interaction consistency in long sequences. Existing Text-to-Motion-based approaches, such as discrete motion tokenization, cannot be directly applied to HOI generation due to limited data in this domain and the complexity of the modality. To address the problem of interaction consistency in long sequences, we propose an autoregressive diffusion model (ARDHOI) that predicts the next continuous token. Specifically, we introduce a Contrastive Variational Autoencoder (cVAE) to learn a physically plausible space of continuous HOI tokens, thereby ensuring that generated human-object motions are realistic and natural. For generating sequences autoregressively, we develop a Mamba-based context encoder to capture and maintain consistent sequential actions. Additionally, we implement an MLP-based denoiser to generate the subsequent token conditioned on the encoded context. Our model has been evaluated on the OMOMO and BEHAVE datasets, where it outperforms existing state-of-the-art methods in terms of both performance and inference speed. This makes ARDHOI a robust and efficient solution for text-driven HOI tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- InterPrior: Scaling Generative Control for Physics-Based Human-Object InteractionsSirui Xu, Samuel Schulter, Morteza Ziyadi, Xialin He 等CVPR 2026 · 被引用 14 次
- HOI-Dyn: Learning Interaction Dynamics for Human-Object Motion DiffusionLin Wu, Zhixiang Chen, Jianglin LanNeurIPS 2025 · 被引用 10 次
- Pulp Motion: Framing-aware multimodal camera and human motion generationRobin Courant, Xi WANG, David Loiseaux, Marc Christie 等ICLR 2026 · 被引用 8 次
- ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction GenerationZichen Geng, Zeeshan Hayder, Wei Liu, Hesheng Wang 等CVPR 2026 · 被引用 1 次
- MoCoDiff: A Controllable Autoregressive Diffusion Model for Expressive Motion GenerationWenfeng Song, Xuehan Wang, Shuai Li, Yi Chen 等CVPR 2026
它引用的顶会 Paper25
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang 等ICML 2024 · 被引用 1,725 次
相关 Paper
- Causal Motion Diffusion Models for Autoregressive Motion GenerationQing Yu, Akihisa Watanabe, Kent FujiwaraCVPR 2026 · 被引用 9 次
- Interact2Ar: Full-Body Human-Human Interaction Generation via Autoregressive Diffusion ModelsPablo Ruiz-Ponce, Sergio Escalera, José García Rodríguez, Jiankang Deng 等CVPR 2026 · 被引用 6 次
- Disentangled Hierarchical VAE for 3D Human-Human Interaction GenerationZichen Geng, Zeeshan Hayder, Bo Miao, Jian Liu 等ICLR 2026 · 被引用 3 次
- LatentHOI: On the Generalizable Hand Object Motion Generation with Latent Hand DiffusionMuchen Li, Sammy Christen, Chengde Wan, Yujun Cai 等CVPR 2025
- Executing your Commands via Motion Diffusion in Latent SpaceXin Chen, Biao Jiang, Wen Liu, Zilong Huang 等CVPR 2023
