WaveAR: Wavelet-Aware Continuous Autoregressive Diffusion for Accurate Human Motion Prediction
Shengchuan Gao, Shuo Wang, Yabiao Wang, Ran Yi
Abstract
This work tackles a challenging problem: stochastic human motion prediction (SHMP), which aims to forecast diverse and physically plausible future pose sequences based on a short history of observed motion. While autoregressive sequence models have excelled in related generation tasks, their reliance on vectorquantized tokenization limits motion fidelity and training stability. To overcome these drawbacks, we introduce WaveAR, a novel AR based framework which is the first successful application of a continuous autoregressive generation paradigm to HMP to our best knowledge. WaveAR consists of two stages. In the first stage, a lightweight Spatio-Temporal VAE (ST-VAE) compresses the raw 3Djoint sequence into a downsampled latent token stream, providing a compact yet expressive foundation. In the second stage, we apply masked autoregressive prediction directly in this continuous latent space, conditioning on both unmasked latents and multi-scale spectral cues extracted via a 2D discrete wavelet transform. A fusion module consisting of alternating cross-attention and self-attention layers adaptively fuses temporal context with low-and high-frequency wavelet subbands, and a small MLP-based diffusion head predicts per-token noise residuals under a denoising loss. By avoiding vector quantization and integrating localized frequency information, WaveAR preserves fine-grained motion details while maintaining fast inference speed. Extensive experiments on standard benchmarks demonstrate that our approach delivers more accurate and computationally efficient predictions than prior state-of-the-art methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 661b70dd-25e2-450d-8986-9a69094a4a41Cited by top-tier papers2
- PoseAnything: General Pose-guided Video Generation with Part-aware Temporal CoherenceRuiyan Wang, Teng Hu, Kaihui Huang, Zihan Su et al.CVPR 2026
- Open the Motion Door: Atomic Motion Decomposition and Recomposition for Open-Vocabulary Motion GenerationKe Fan, Jiangning Zhang, Ran Yi, Jingyu Gong et al.CVPR 2026
Builds on25
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
- MSR-GCN: Multi-Scale Residual Graph Convolution Networks for Human Motion PredictionLingwei Dang, Yongwei Nie, Chengjiang Long, Qing Zhang et al.ICCV 2021 · 252 citations
Related papers
- Weakly-supervised Action Transition Learning for Stochastic Human Motion PredictionWei Mao, Miaomiao Liu, Mathieu SalzmannCVPR 2022 · 29 citations
- MotionStreamer: Streaming Motion Generation via Diffusion-Based Autoregressive Model in Causal Latent SpaceLixing Xiao, Shunlin Lu, Huaijin Pi, Ke Fan et al.ICCV 2025 · 11 citations
- A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗Yujun Cai, Yiwei Wang, Yiheng Zhu, Tat-Jen Cham et al.ICCV 2021 · 83 citations
- Contextually Plausible and Diverse 3D Human Motion PredictionSadegh Aliakbarian, Fatemeh Sadat Saleh, Lars Petersson, Stephen Gould et al.ICCV 2021 · 44 citations
- Rethinking Diffusion for Text-Driven Human Motion Generation: Redundant Representations, Evaluation, and Masked AutoregressionZichong Meng, Yiming Xie, Xiaogang Peng, Zeyu Han et al.CVPR 2025
