Deep Compositional Phase Diffusion for Long Motion Sequence Generation
Ho Yin Au, Jie Chen, Junkun Jiang, Jingyu Xiang
Abstract
Recent research on motion generation has shown significant progress in generating semantically aligned motion with singular semantics. However, when employing these models to create composite sequences containing multiple semantically generated motion clips, they often struggle to preserve the continuity of motion dynamics at the transition boundaries between clips, resulting in awkward transitions and abrupt artifacts. To address these challenges, we present Compositional Phase Diffusion, which leverages the Semantic Phase Diffusion Module (SPDM) and Transitional Phase Diffusion Module (TPDM) to progressively incorporate semantic guidance and phase details from adjacent motion clips into the diffusion process. Specifically, SPDM and TPDM operate within the latent motion frequency domain established by the pre-trained Action-Centric Motion Phase Autoencoder (ACT-PAE). This allows them to learn semantically important and transition-aware phase information from variable-length motion clips during training. Experimental results demonstrate the competitive performance of our proposed framework in generating compositional motion sequences that align semantically with the input conditions, while preserving phase transitional continuity between preceding and succeeding motion clips. Additionally, motion inbetweening task is made possible by keeping the phase parameter of the input motion sequences fixed throughout the diffusion process, showcasing the potential for extending the proposed framework to accommodate various application scenarios. Codes are available at https://github.com/asdryau/TransPhase.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e8ea74b7-c211-4643-a59c-f389dbd08f64Cited by top-tier papers3
- LaMoGen: Language to Motion Generation Through LLM-Guided Symbolic InferenceJunkun JIANG, Ho Yin Au, Jingyu Xiang, Jie ChenCVPR 2026 · 2 citations
- SOSControl: Enhancing Human Motion Generation Through Saliency-Aware Symbolic Orientation and Timing ControlHo Yin Au, Junkun Jiang, Jie ChenAAAI 2026 · 1 citation
- FunPhase: A Periodic Functional Autoencoder for Motion Generation via Phase ManifoldsMarco Pegoraro, Evan Atherton, Bruno Roy, Aliasghar Khani et al.ICML 2026 · 1 citation
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Action-Conditioned 3D Human Motion Synthesis with Transformer VAEMathis Petrovich, Michael J. Black, Gül VarolICCV 2021 · 672 citations
- Generating Diverse and Natural 3D Human Motions from TextChuan Guo, Shihao Zou, Xinxin Zuo, Sen Wang et al.CVPR 2022 · 462 citations
Related papers
- A Unified 3D Human Motion Synthesis Model via Conditional Variational Auto-Encoder∗Yujun Cai, Yiwei Wang, Yiheng Zhu, Tat-Jen Cham et al.ICCV 2021 · 83 citations
- Synthesizing Long-Term Human Motions with Diffusion Models via Coherent SamplingZhao Yang, Bing Su, Ji-Rong WenACM MM 2023 · 17 citations
- WalkTheDog: Cross-Morphology Motion Alignment via Phase ManifoldsPeizhuo Li, Sebastian Starke, Yuting Ye, Olga Sorkine-HornungSIGGRAPH 2024 · 10 citations
- MoAlign: Motion-Centric Representation Alignment for Video Diffusion ModelsAritra Bhowmik, Denis Korzhenkov, Cees G. M. Snoek, Amir Habibian et al.ICLR 2026 · 15 citations
- TriC-Motion: Tri-Domain Causal Modeling Grounded Text-to-Motion GenerationYiyang Cao, Yunze Deng, Ziyu Lin, Bin Feng et al.ICLR 2026
