An End-to-End, Low-Cost, and High-Fidelity 3D Video Pipeline for Mobile Devices
Tiancheng Fang, Chaoyue Niu, Yujie Sun, Chengfei Lv, Xiaotang Jiang, Ben Xue, Fan Wu, Guihai Chen
Abstract
To provide full-body 3D videos of performers showcasing diverse clothes with dynamic movements for e-commerce platforms, we develop an end-to-end, low-cost, and high-fidelity production and deployment pipeline. We first set up a low-cost capture studio with only 24 RGB cameras and embrace fast neural surface reconstruction to produce high-quality meshes without depth information. We then quickly group all the frames with local motion priors, select a keyframe for each group, and accurately register any other frame to the keyframe under the guidance of semantic labels, thereby avoiding transmitting all the frames to mobile devices and loading them into memory. For real-time rendering, we propose an on-device sparse computation method for efficient deformation from keyframes to the other frames. Evaluation over 2 self-captured performances and 8 public performances reveals that the pipeline achieves the reconstruction time of 28 minutes per frame, the average PSNR of 30.4, the average bandwidth requirement of 4.2MB/s, and the on-device frame rate of 60 fps, demonstrating superiority over existing baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Related papers
- Holoported Characters: Real-Time Free-Viewpoint Rendering of Humans from Sparse RGB CamerasAshwath Shetty, Marc Habermann, Guoxing Sun, Diogo C. Luvizon et al.CVPR 2024 · 9 citations
- Clothed Human Performance Capture with a Double-layer Neural Radiance FieldsKangkan Wang, Guofeng Zhang, Suxu Cong, Jian YangCVPR 2023
- Link to the Past: Temporal Propagation for Fast 3D Human Reconstruction from Monocular VideoMatthew Marchellus, Nadhira Noor, In Kyu ParkCVPR 2025
- Novel View Synthesis of Human Interactions from Sparse Multi-view VideosQing Shuai, Chen Geng, Qi Fang, Sida Peng et al.SIGGRAPH 2022 · 43 citations
- Gaussian Head & Shoulders: High Fidelity Neural Upper Body Avatars with Anchor Gaussian Guided Texture WarpingTianhao (Walter) Wu, Jing Yang, Zhilin Guo, Jingyi Wan et al.ICLR 2025
