The Implicit Values of A Good Hand Shake: Handheld Multi-Frame Neural Depth Refinement
Ilya Chugunov, Yuxuan Zhang, Zhihao Xia, Xuaner Zhang, Jiawen Chen, Felix Heide
摘要
Modern smartphones can continuously stream multi-megapixel RGB images at 60 Hz, synchronized with high-quality 3D pose information and low-resolution LiDAR-driven depth estimates. During a snapshot photograph, the natural unsteadiness of the photographer's hands offers millimeter-scale variation in camera pose, which we can capture along with RGB and depth in a circular buffer. In this work we explore how, from a bundle of these measurements acquired during viewfinding, we can combine dense micro-baseline parallax cues with kilopixel LiDAR depth to distill a high-fidelity depth map. We take a test-time optimization approach and train a coordinate MLP to output photometrically and geometrically consistent depth estimates at the continuous coordinates along the path traced by the photographer's natural hand shake. With no additional hardware, artificial hand motion, or user interaction beyond the press of a button, our proposed method brings high-resolution depth estimates to point-and-shoot “table-top” photography – textured objects at close range.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- GenSDF: Two-Stage Learning of Generalizable Signed Distance FunctionsGene Chou, Ilya Chugunov, Felix HeideNeurIPS 2022 · 被引用 46 次
- Transient Neural Radiance Fields for Lidar View Synthesis and 3D ReconstructionAnagh Malik, Parsa Mirdehghan, Sotiris Nousias, Kyros Kutulakos 等NeurIPS 2023 · 被引用 40 次
- AONeuS: A Neural Rendering Framework for Acoustic-Optical Sensor FusionMohamad Qadri, Kevin Zhang, Akshay Hinduja, Michael Kaess 等SIGGRAPH 2024 · 被引用 23 次
- Neural Fields for Structured LightingAarrushi Shandilya, Benjamin Attal, Christian Richardt, James Tompkin 等ICCV 2023 · 被引用 16 次
- Rapid Network Adaptation: Learning to Adapt Neural Networks Using Test-Time FeedbackTeresa Yeo, Oguzhan Fatih Kar, Zahra Sodagar, Amir ZamirICCV 2023 · 被引用 10 次
它引用的顶会 Paper10
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Nerfies: Deformable Neural Radiance FieldsKeunhong Park, Utkarsh Sinha, Jonathan T. Barron, Sofien Bouaziz 等ICCV 2021 · 被引用 1,442 次
- Depth-supervised NeRF: Fewer Views and Faster Training for FreeKangle Deng, Andrew Liu, Jun-Yan Zhu, Deva RamananCVPR 2022 · 被引用 756 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
相关 Paper
- Shakes on a Plane: Unsupervised Depth Estimation from Unstabilized PhotographyIlya Chugunov, Yuxuan Zhang, Felix HeideCVPR 2023
- Mesoscopic Photogrammetry With an Unstabilized Phone CameraKevin C. Zhou, Colin L. V. Cooke, Jaehee Park, Ruobing Qian 等CVPR 2021
- One shot 3D photographyJohannes Kopf, Kevin Matzen, Suhib Alsisan, Ocean Quigley 等SIGGRAPH 2020 · 被引用 65 次
- Learning Single Camera Depth Estimation Using Dual-PixelsRahul Garg, Neal Wadhwa, Sameer Ansari, Jonathan T. BarronICCV 2019 · 被引用 123 次
- Image as an Imu: Estimating Camera Motion From a Single Motion-Blurred ImageJerred Chen, Ronald ClarkICCV 2025
