AnyCam: Learning to Recover Camera Poses and Intrinsics from Casual Videos
Felix Wimbauer, Weirong Chen, Dominik Muhle, Christian Rupprecht, Daniel Cremers
摘要
Estimating camera motion and intrinsics from casual videos is a core challenge in computer vision. Traditional bundle-adjustment based methods, such as SfM and SLAM, struggle to perform reliably on arbitrary data. Although specialized SfM approaches have been developed for handling dynamic scenes, they either require intrinsics or computationally expensive test-time optimization and often fall short in performance. Recently, methods like Dust3r have reformulated the SfM problem in a more data-driven way. While such techniques show promising results, they are still 1) not robust towards dynamic objects and 2) require labeled data for supervised training. As an alternative, we propose AnyCam, a fast transformer model that directly estimates camera poses and intrinsics from a dynamic video sequence in feed-forward fashion. Our intuition is that such a network can learn strong priors over realistic camera poses. To scale up our training, we rely on an uncertainty-based loss formulation and pre-trained depth and flow networks instead of motion or trajectory supervision. This allows us to use diverse, unlabelled video datasets obtained mostly from YouTube. Additionally, we ensure that the predicted trajectory does not accumulate drift over time through a lightweight trajectory refinement step. We test AnyCam on established datasets, where it delivers accurate camera poses and intrinsics both qualitatively and quantitatively. Furthermore, even with trajectory refinement, Any-Cam is significantly faster than existing works for SfM in dynamic settings. Finally, by combining camera information, uncertainty, and depth, our model can produce high-quality 4D pointclouds. For more details and code, please check out our project page: fwmb.github.io/anycam
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OccAny: Generalized Unconstrained Urban 3D OccupancyAnh-Quan Cao, Tuan-Hung VuCVPR 2026 · 被引用 6 次
- Back on Track: Bundle Adjustment for Dynamic Scene ReconstructionWeirong Chen, Ganlin Zhang, Felix Wimbauer, Rui Wang 等ICCV 2025 · 被引用 1 次
- CogniMap3D: Cognitive 3D Mapping and Rapid RetrievalFeiran Wang, Junyi Wu, Dawen Cai, Yuan Hong 等ICLR 2026 · 被引用 1 次
- Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric AnnotationsYouyu Chen, Junjun Jiang, Yueru Luo, Kui Jiang 等CVPR 2026
它引用的顶会 Paper20
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai 等ICCV 2023 · 被引用 388 次
相关 Paper
- Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICCV 2025 · 被引用 11 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionJunyi Zhang, Charles Herrmann, Junhwa Hur, Varun Jampani 等ICLR 2025 · 被引用 3 次
- Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeChuhan Zhang, Guillaume Le Moing, Skanda Koppula, Ignacio Rocco 等CVPR 2026 · 被引用 52 次
- AnyMap: Learning a General Camera Model for Structure-from-Motion with Unknown Distortion in Dynamic ScenesAndrea Porfiri Dal Cin, Georgi Dikov, Jihong Ju, Mohsen GhafoorianCVPR 2025
