FlowCam: Training Generalizable 3D Radiance Fields without Camera Poses via Pixel-Aligned Scene Flow
Cameron Smith, Yilun Du, Ayush Tewari, Vincent Sitzmann
摘要
Reconstruction of 3D neural fields from posed images has emerged as a promising method for self-supervised representation learning. The key challenge preventing the deployment of these 3D scene learners on large-scale video data is their dependence on precise camera poses from structure-from-motion, which is prohibitively expensive to run at scale. We propose a method that jointly reconstructs camera poses and 3D neural scene representations online and in a single forward pass. We estimate poses by first lifting frame-to-frame optical flow to 3D scene flow via differentiable rendering, preserving locality and shift-equivariance of the image processing backbone. SE(3) camera pose estimation is then performed via a weighted least-squares fit to the scene flow field. This formulation enables us to jointly supervise pose estimation and a generalizable neural scene representation via re-rendering the input video, and thus, train end-to-end and fully self-supervised on real-world video datasets. We demonstrate that our method performs robustly on diverse, real-world video, notably on sequences traditionally challenging to optimization-based pose estimation techniques. Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape PredictionPeng Wang, Hao Tan, Sai Bi, Yinghao Xu 等ICLR 2024 · 被引用 170 次
- LU-NeRF: Scene and Pose Estimation by Synchronizing Local Unposed NeRFsZezhou Cheng, Carlos Esteves, Varun Jampani, Abhishek Kar 等ICCV 2023 · 被引用 46 次
- True Self-Supervised Novel View Synthesis is TransferableThomas W. Mitchel, Hyunwoo Ryu, Vincent SitzmannICLR 2026 · 被引用 13 次
- TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free ReconstructionYihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang 等CVPR 2026 · 被引用 13 次
- No Pose at All: Self-Supervised Pose-Free 3D Gaussian Splatting from Sparse ViewsRanran Huang, Krystian MikolajczykICCV 2025 · 被引用 12 次
它引用的顶会 Paper27
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil 等NeurIPS 2020 · 被引用 4,036 次
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell 等NeurIPS 2020 · 被引用 4,008 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- BARF: Bundle-Adjusting Neural Radiance FieldsChen-Hsuan Lin, Wei-Chiu Ma, Antonio Torralba, Simon LuceyICCV 2021 · 被引用 867 次
相关 Paper
- Flow-NeRF: Joint Learning of Geometry, Poses, and Dense Flow within Unified Neural RepresentationsXunzhi Zheng, Dan XuCVPR 2025
- Neural Scene Flow Fields for Space-Time View Synthesis of Dynamic ScenesZhengqi Li, Simon Niklaus, Noah Snavely, Oliver WangCVPR 2021
- Self-Supervised Representation Learning from Flow EquivarianceYuwen Xiong, Mengye Ren, Wenyuan Zeng, Raquel Urtasun WaabiICCV 2021 · 被引用 32 次
- Self-Supervised Monocular Scene Flow EstimationJunhwa Hur, Stefan RothCVPR 2020
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 被引用 37 次
