Shakes on a Plane: Unsupervised Depth Estimation from Unstabilized Photography
Ilya Chugunov, Yuxuan Zhang, Felix Heide
Abstract
Modern mobile burst photography pipelines capture and merge a short sequence of frames to recover an enhanced image, but often disregard the 3D nature of the scene they capture, treating pixel motion between images as a 2D aggregation problem. We show that in a "long-burst", fortytwo 12-megapixel RAW frames captured in a two-second sequence, there is enough parallax information from natural hand tremor alone to recover high-quality scene depth. To this end, we devise a test-time optimization approach that fits a neural RGB-D representation to long-burst data and simultaneously estimates scene depth and camera motion. Our plane plus depth model is trained end-to-end, and performs coarse-to-fine refinement by controlling which multiresolution volume features the network has access to at what time during training. We validate the method experimentally, and demonstrate geometrically accurate depth reconstructions with no additional hardware or separate data pre-processing and pose-estimation steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 49f3d23b-d973-440b-abe8-656118bf62e9Cited by top-tier papers5
- AONeuS: A Neural Rendering Framework for Acoustic-Optical Sensor FusionMohamad Qadri, Kevin Zhang, Akshay Hinduja, Michael Kaess et al.SIGGRAPH 2024 · 23 citations
- TurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured LightParsa Mirdehghan, Maxx Wu, Wenzheng Chen, David B. Lindell et al.CVPR 2024 · 5 citations
- Dark3R: Learning Structure from Motion in the DarkAndrew Y. Guo, Anagh Malik, SaiKiran Kumar Tedla, Yutong Dai et al.CVPR 2026 · 3 citations
- Neural Spline Fields for Burst Image Fusion and Layer SeparationIlya Chugunov, David Shustin, Ruyu Yan, Chenyang Lei et al.CVPR 2024
- Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation ModelsYu Jiang, Hanwen Jiang, Ahmed Abdelkader, Wen-Sheng Chu et al.CVPR 2026
Builds on22
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Block-NeRF: Scalable Large Scene Neural View SynthesisMatthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan et al.CVPR 2022 · 702 citations
Related papers
- The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth RefinementIlya Chugunov, Yuxuan Zhang, Zhihao Xia, Xuaner Zhang et al.CVPR 2022 · 11 citations
- Consistent depth of moving objects in videoZhoutong Zhang, Forrester Cole, Richard Tucker, William T. Freeman et al.SIGGRAPH 2021 · 26 citations
- Mesoscopic Photogrammetry With an Unstabilized Phone CameraKevin C. Zhou, Colin L. V. Cooke, Jaehee Park, Ruobing Qian et al.CVPR 2021
- Digital Gimbal: End-to-End Deep Image Stabilization With Learnable Exposure TimesOmer Dahary, Matan Jacoby, Alex M. BronsteinCVPR 2021
- Image as an Imu: Estimating Camera Motion From a Single Motion-Blurred ImageJerred Chen, Ronald ClarkICCV 2025
