ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
Zichen Geng, Zeeshan Hayder, Wei Liu, Hesheng Wang, Ajmal Saeed Mian
Abstract
3D human reaction generation faces three main challenges:
(1) high motion fidelity, (2) real-time inference, and (3) autoregressive adaptability for online scenarios. Existing methods fail to meet all three simultaneously. We propose ARMFlow, a MeanFlow-based autoregressive framework that models temporal dependencies between actor and reactor motions. It consists of a causal context encoder and an MLP-based velocity predictor. We introduce Bootstrap Contextual Encoding (BSCE) in training, encoding generated history instead of the ground-truth ones, to alleviate error accumulation in autoregressive generation. We further introduce the offline variant ReMFlow, achieving state-ofthe-art performance with the fastest inference among offline methods. Our ARMFlow addresses key limitations of online settings by: ( 1) enhancing semantic alignment via a global contextual encoder; (2) achieving high accuracy and low latency in a single-step inference; and (3) reducing accumulated errors through BSCE. Our single-step online generation surpasses existing online methods on InterHuman and InterX by about 30% in FID, while matching offline stateof-the-art performance despite using only partial sequence conditions. The official implementation is publicly available at: https://github.com/ZenGengChin/armflow official
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
Related papers
- ReMoGen: Real-time Human Interaction-to-Reaction Generation via Modular Learning from Diverse DataYaoqin Ye, Yiteng Xu, Qin Sun, Xinge Zhu et al.CVPR 2026 · 2 citations
- Unified Number-Free Text-to-Motion Generation Via Flow MatchingGuanhe Huang, Oya ÇeliktutanCVPR 2026
- Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion GenerationKaifeng Zhao, Mathis Petrovich, Haotian Zhang, Tingwu Wang et al.SIGGRAPH 2026
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion SynthesisKaiyang Ji, Ye Shi, Zichen Jin, Kangyi Chen et al.ICCV 2025 · 3 citations
- ReGenNet: Towards Human Action-Reaction SynthesisLiang Xu, Yizhou Zhou, Yichao Yan, Xin Jin et al.CVPR 2024 · 18 citations
