UAV4D: Dynamic Neural Rendering of Human-Centric UAV Imagery Using Gaussian Splatting
Jaehoon Choi, Dongki Jung, Chris Maxey, Sungmin Eum, Yonghan Lee, Dinesh Manocha, Heesung Kwon
摘要
Despite significant advancements in dynamic neural rendering, existing methods fail to address the unique challenges posed by UAV-captured scenarios, particularly those involving monocular camera setups, top-down perspective, and multiple small, moving humans, which are not adequately represented in existing datasets. In this work, we introduce UAV4D, a framework for enabling photorealistic rendering for dynamic real-world scenes captured by UAVs. Specifically, we address the challenge of reconstructing dynamic scenes with multiple moving pedestrians from monocular video data without the need for additional sensors. We use a combination of a 3D foundation model and a human mesh reconstruction model to reconstruct both the scene background and humans. We propose a novel approach to resolve the scene scale ambiguity and place both humans and the scene in world coordinates by identifying human-scene contact points. Additionally, we exploit the SMPL model and background mesh to initialize Gaussian splats, enabling holistic scene rendering. We evaluated our method on three complex UAV-captured datasets: VisDrone, Manipal-UAV, and Okutama-Action, each with distinct characteristics and 10 -50 humans. Our results demonstrate the benefits of our approach over existing methods in novel view synthesis, achieving a 1.5 dB PSNR improvement and superior visual sharpness. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D ReconstructionHanyang Liu, Rongjun QinCVPR 2026 · 被引用 2 次
- AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene ReconstructionTingyun Li, Xinyi Liu, Yongjun Zhang, Yi Wan 等CVPR 2026
它引用的顶会 Paper38
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman 等ICCV 2021 · 被引用 2,700 次
- Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Dor Verbin, Pratul P. Srinivasan 等CVPR 2022 · 被引用 1,603 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
相关 Paper
- Uncertainty Matters in Dynamic Gaussian Splatting for Monocular 4D ReconstructionFengzhi Guo, Chih-Chuan Hsu, Sihao Ding, Cheng ZhangICLR 2026 · 被引用 6 次
- Motion4D: Learning 3D-Consistent Motion and Semantics for 4D Scene UnderstandingHaoran Zhou, Gim Hee LeeNeurIPS 2025 · 被引用 3 次
- 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes from Monocular VideosMengqi Guo, Bo Xu, Yanyan Li, Gim Hee LeeNeurIPS 2025 · 被引用 2 次
- MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian SplattingHaoran Zhou, Gim Hee LeeCVPR 2026 · 被引用 1 次
- MoSca: Dynamic Gaussian Fusion from Casual Videos via 4D Motion ScaffoldsJiahui Lei, Yijia Weng, Adam W. Harley, Leonidas J. Guibas 等CVPR 2025
