Back on Track: Bundle Adjustment for Dynamic Scene Reconstruction
Weirong Chen, Ganlin Zhang, Felix Wimbauer, Rui Wang, Nikita Araslanov, Andrea Vedaldi, Daniel Cremers
Abstract
Traditional SLAM systems, which rely on bundle adjustment, struggle with the highly dynamic scenes commonly found in casual videos. Such videos entangle the motion of dynamic elements, undermining the assumption of static environments required by traditional systems. Existing techniques either filter out dynamic elements or model their motion independently. However, the former often results in incomplete reconstructions, while the latter can lead to inconsistent motion estimates. Taking a novel approach, this work leverages a 3D point tracker to separate camera-induced motion from the observed motion of dynamic objects. By considering only the camera-induced component, bundle adjustment can operate reliably on all scene elements. We further ensure depth consistency across video frames with lightweight post-processing based on scale maps. Our framework combines the core of traditional SLAM—bundle adjustment—with a robust learning-based 3D tracker. Integrating motion decomposition, bundle adjustment, and depth refinement, our unified framework, BA-Track, accurately tracks camera motion and produces temporally coherent and scale-consistent dense reconstructions, accommodating both static and dynamic elements. Our experiments on challenging datasets reveal significant improvements in camera pose estimation and 3D reconstruction accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext efb10d08-d942-4401-9b89-1d0938d9dc89Cited by top-tier papers11
- LaVR: Scene Latent Conditioned Generative Video Trajectory Re-Rendering using Large 4D Reconstruction ModelsMingyang Xie, Numair Khan, Tianfu Wang, Naina Dhingra et al.CVPR 2026 · 10 citations
- TrackingWorld: World-centric Monocular 3D Tracking of Almost All PixelsJiahao Lu, Weitao Xiong, Jiacheng Deng, Peng Li et al.NeurIPS 2025 · 7 citations
- AnomalyVFM - Transforming Vision Foundation Models into Zero-Shot Anomaly DetectorsMatic Fucka, Vitjan Zavrtanik, Danijel SkocajCVPR 2026 · 4 citations
- MotionCrafter: Dense Geometry and Motion Reconstruction with a 4D VAERuijie Zhu, Jiahao Lu, Wenbo Hu, Xiaoguang Han et al.CVPR 2026 · 3 citations
- WildPose: A Unified Framework for Robust Pose Estimation in the WildJianhao Zheng, Liyuan Zhu, Zihan Zhu, Iro ArmeniCVPR 2026 · 1 citation
Builds on20
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
Related papers
- HumanBA: Human-Aware Bundle Adjustment via Global Human-Camera DecouplingFengyuan Yang, Tanuj Sur, Tze Ho Elden Tse, Angela YaoCVPR 2026
- Augmenting TV Shows via Uncalibrated Camera Small Motion Tracking in Dynamic SceneYizhen Lao, Jie Yang, Xinying Wang, Jianxin Lin et al.ACM MM 2021 · 1 citation
- Dynamic Visual SLAM using a General 3D PriorXingguang Zhong, Liren Jin, Marija Popovic, Jens Behley et al.CVPR 2026 · 1 citation
- DROID-SLAM in the WildMoyang Li, Zihan Zhu, Marc Pollefeys, Daniel BarathCVPR 2026 · 10 citations
- DG-SLAM: Robust Dynamic Gaussian Splatting SLAM with Hybrid Pose OptimizationYueming Xu, Haochen Jiang, Zhongyang Xiao, Jianfeng Feng et al.NeurIPS 2024 · 65 citations
