SurMo: Surface-based 4D Motion Modeling for Dynamic Human Rendering
Tao Hu, Fangzhou Hong, Ziwei Liu
Abstract
Dynamic human rendering from video sequences has achieved remarkable progress by formulating the rendering as a mapping from static poses to human images. However, existing methods focus on the human appearance reconstruction of every single frame while the temporal motion relations are not fully explored. In this paper, we propose a new 4D motion modeling paradigm, SurMo, that jointly models the temporal dynamics and human appearances in a unified framework with three key designs: 1) Surface-based motion encoding that models 4D human motions with an efficient compact surface-based triplane. It encodes both spatial and temporal motion relations on the dense surface manifold of a statistical body template, which inherits body topology priors for generalizable novel view synthesis with sparse training observations. 2) Physical motion decoding that is designed to encourage physical motion learning by decoding the motion triplane features at timestep t to predict both spatial derivatives and temporal derivatives at the next timestep t + 1 in the training stage. 3) 4D appearance decoding that renders the motion triplanes into images by an efficient volumetric surface-conditioned renderer that focuses on the rendering of body surfaces with motion learning conditioning. Extensive experiments validate the stateof-the-art performance of our new paradigm and illustrate the expressiveness of surface-based motion triplanes for rendering high-fidelity view-consistent humans with fast motions and even motion-dependent shadows. Our project page is at: https://taohuumd.github.io/projects/SurMo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human AvatarsYifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma et al.ICCV 2025 · 3 citations
- Sequential Gaussian Avatars with Hierarchical Motion ContextWangze Xu, Yifan Zhan, Zhihang Zhong, Xiao SunICCV 2025 · 1 citation
- Free-viewpoint Human Animation with Pose-correlated Reference SelectionFa-Ting Hong, Zhan Xu, Haiyang Liu, Qinjie Lin et al.CVPR 2025
Builds on25
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano et al.CVPR 2022 · 984 citations
- StyleNeRF: A Style-based 3D Aware Generator for High-resolution Image SynthesisJiatao Gu, Lingjie Liu, Peng Wang, Christian TheobaltICLR 2022 · 622 citations
- Animatable Neural Radiance Fields for Modeling Dynamic Human BodiesSida Peng, Junting Dong, Qianqian Wang, Shangzhan Zhang et al.ICCV 2021 · 461 citations
- HumanNeRF: Free-viewpoint Rendering of Moving People from Monocular VideoChung-Yi Weng, Brian Curless, Pratul P. Srinivasan, Jonathan T. Barron et al.CVPR 2022 · 411 citations
Related papers
- Learning Motion-Dependent Appearance for High-Fidelity Rendering of Dynamic Humans from a Single CameraJae Shin Yoon, Duygu Ceylan, Tuanfeng Y. Wang, Jingwan Lu et al.CVPR 2022 · 10 citations
- H4D: Human 4D Modeling by Learning Neural Compositional RepresentationBoyan Jiang, Yinda Zhang, Xingkui Wei, Xiangyang Xue et al.CVPR 2022 · 20 citations
- ST-4DGS: Spatial-Temporally Consistent 4D Gaussian Splatting for Efficient Dynamic Scene RenderingDeqi Li, Shi-Sheng Huang, Zhiyuan Lu, Xinran Duan et al.SIGGRAPH 2024 · 33 citations
- Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion ModelsMarc Benedí San Millán, Angela Dai, Matthias NießnerICLR 2026 · 3 citations
- Holoported Characters: Real-Time Free-Viewpoint Rendering of Humans from Sparse RGB CamerasAshwath Shetty, Marc Habermann, Guoxing Sun, Diogo C. Luvizon et al.CVPR 2024 · 9 citations
