The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement
Ilya Chugunov, Yuxuan Zhang, Zhihao Xia, Xuaner Zhang, Jiawen Chen, Felix Heide
Abstract
Modern smartphones can continuously stream multi-megapixel RGB images at 60 Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in camera pose, which we can capture along with RGB and depth in a circular buffer. In this work we explore how, from a bundle of these measurements acquired during viewfinding, we can combine dense micro-baseline parallax cues with kilopixel LiDAR depth to distill a high-fidelity depth map. We take a test-time optimization approach and train a coordinate MLP to output photometrically and geometrically consistent depth estimates at the continuous coordinates along the path traced by the photographer's natural hand shake. With no additional hardware, artificial hand motion, or user interaction beyond the press of a button, our proposed method brings high-resolution depth estimates to point-and-shoot “table-top” photography – textured objects at close range.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d581d5c5-c09b-4263-852d-2fb607bc3d36Cited by top-tier papers10
- GenSDF: Two-Stage Learning of Generalizable Signed Distance FunctionsGene Chou, Ilya Chugunov, Felix HeideNeurIPS 2022 · 46 citations
- Transient Neural Radiance Fields for Lidar View Synthesis and 3D ReconstructionAnagh Malik, Parsa Mirdehghan, Sotiris Nousias, Kyros Kutulakos et al.NeurIPS 2023 · 40 citations
- AONeuS: A Neural Rendering Framework for Acoustic-Optical Sensor FusionMohamad Qadri, Kevin Zhang, Akshay Hinduja, Michael Kaess et al.SIGGRAPH 2024 · 23 citations
- Neural Fields for Structured LightingAarrushi Shandilya, Benjamin Attal, Christian Richardt, James Tompkin et al.ICCV 2023 · 16 citations
- Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time FeedbackTeresa Yeo, Oguzhan Fatih Kar, Zahra Sodagar, Amir ZamirICCV 2023 · 10 citations
Builds on10
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz et al.ICCV 2021 · 1,442 citations
- Depth-supervised NeRF: Fewer Views and Faster Training for FreeKangle Deng, Andrew Liu, Jun-Yan Zhu, Deva RamananCVPR 2022 · 756 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
Related papers
- Shakes on a Plane: Unsupervised Depth Estimation from Unstabilized PhotographyIlya Chugunov, Yuxuan Zhang, Felix HeideCVPR 2023
- Mesoscopic Photogrammetry With an Unstabilized Phone CameraKevin C. Zhou, Colin L. V. Cooke, Jaehee Park, Ruobing Qian et al.CVPR 2021
- One shot 3D photographyJohannes Kopf, Kevin Matzen, Suhib Alsisan, Ocean Quigley et al.SIGGRAPH 2020 · 65 citations
- Learning Single Camera Depth Estimation Using Dual-PixelsRahul Garg, Neal Wadhwa, Sameer Ansari, Jonathan T. BarronICCV 2019 · 123 citations
- Image as an Imu: Estimating Camera Motion From a Single Motion-Blurred ImageJerred Chen, Ronald ClarkICCV 2025
