A Unified Diffusion Framework for Scene-aware Human Motion Estimation from Sparse Signals
Jiangnan Tang, Jingya Wang, Kaiyang Ji, Lan Xu, Jingyi Yu, Ye Shi
Abstract
Estimating full-body human motion via sparse tracking signals from head-mounted displays and hand controllers in 3D scenes is crucial to applications in AR/VR. One of the biggest challenges to this task is the one-to-many mapping from sparse observations to dense full-body motions, which endowed inherent ambiguities. To help resolve this ambiguous problem, we introduce a new framework to combine rich contextual information provided by scenes to benefit fullbody motion tracking from sparse observations. To estimate plausible human motions given sparse tracking signals and 3D scenes, we develop S<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> Fusion, a unified framework fusing Scene and sparse Signals with a conditional difFusion model. S<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> Fusion first extracts the spatial-temporal relations residing in the sparse signals via a periodic autoencoder, and then produces time-alignment feature embedding as additional inputs. Subsequently, by drawing initial noisy motion from a pre-trained prior, S<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> Fusion utilizes conditional diffusion to fuse scene geometry and sparse tracking signals to generate full-body scene-aware motions. The sampling procedure of S<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> Fusion is further guided by a specially designed scene-penetration loss and phase-matching loss, which effectively regularizes the motion of the lower body even in the absence of any tracking signals, making the generated motion much more plausible and coherent. Extensive experimental results have demonstrated that our S<sup xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">2</sup> Fusion outperforms the state-of-the-art in terms of estimation quality and smoothness. Code is available at https://github.com/jntang/S2Fusion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cebb664-5a70-4e2b-93be-02114944ab00Cited by top-tier papers11
- InterDreamer: Zero-Shot Text to 3D Dynamic Human-Object InteractionSirui Xu, Ziyin Wang, Yu-Xiong Wang, Liangyan GuiNeurIPS 2024 · 78 citations
- Human-Object Interaction via Automatically Designed VLM-Guided Motion PolicyZekai Deng, Ye Shi, Kaiyang Ji, Lan Xu et al.ICLR 2026 · 11 citations
- Towards Immersive Human-X Interaction: A Real-Time Framework for Physically Plausible Motion SynthesisKaiyang Ji, Ye Shi, Zichen Jin, Kangyi Chen et al.ICCV 2025 · 3 citations
- UltraPoser: Pushing the Limits of IMU-based Full-Body Pose Estimation with Ultrasound Sensing on Consumer WearablesYadong Li, Shuning Wang, Yongjian Fu, Justin Chen et al.UIST 2025 · 2 citations
- EgoPoseVR: Spatiotemporal Multi-Modal Reasoning for Egocentric Full-Body Pose in Virtual RealityHaojie Cheng, Shaun Jing Heng Ong, Shaoyu Cai, Aiden Tat Yang Koh et al.IEEE VR 2026 · 1 citation
Builds on47
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- DiffWave: A Versatile Diffusion Model for Audio SynthesisZhifeng Kong, Wei Ping, Jiaji Huang, Kexin Zhao et al.ICLR 2021 · 1,902 citations
- AMASS: Archive of Motion Capture As Surface ShapesNaureen Mahmood, Nima Ghorbani, Nikolaus F. Troje, Gerard Pons-Moll et al.ICCV 2019 · 1,784 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
Related papers
- EnvPoser: Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty ModelingSongpengcheng Xia, Yu Zhang, Zhuo Su, Xiaozheng Zheng et al.CVPR 2025
- Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion ModelYuming Du, Robin Kips, Albert Pumarola, Sebastian Starke et al.CVPR 2023
- Estimating Ego-Body Pose from Doubly Sparse Egocentric Video DataSeunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma et al.NeurIPS 2024 · 9 citations
- Realistic Full-Body Tracking from Sparse Observations via Joint-Level ModelingXiaozheng Zheng, Zhuo Su, Chao Wen, Zhou Xue et al.ICCV 2023 · 57 citations
- FLAG: Flow-based 3D Avatar Generation from Sparse ObservationsSadegh Aliakbarian, Pashmina Cameron, Federica Bogo, Andrew W. Fitzgibbon et al.CVPR 2022
