Learning to Zoom and Unzoom
Chittesh Thavamani, Mengtian Li, Francesco Ferroni, Deva Ramanan
摘要
Many perception systems in mobile computing, autonomous navigation, and AR/VR face strict compute constraints that are particularly challenging for high-resolution input images. Previous works propose nonuniform downsamplers that "learn to zoom" on salient image regions, reducing compute while retaining task-relevant image information. However, for tasks with spatial labels (such as 2D/3D object detection and semantic segmentation), such distortions may harm performance. In this work (LZU), we "learn to zoom" in on the input image, compute spatial features, and then "unzoom" to revert any deformations. To enable efficient and differentiable unzooming, we approximate the zooming warp with a piecewise bilinear mapping that is invertible. LZU can be applied to any task with 2D spatial input and any model with 2D spatial features, and we demonstrate this versatility by evaluating on a variety of tasks and datasets: object detection on Argoverse-HD, semantic segmentation on Cityscapes, and monocular 3D object detection on nuScenes. Interestingly, we observe boosts in performance even when high-resolution sensor data is unavailable, implying that LZU can be used to "learn to upsample" as well. Code and additional visuals are available at https://tchittesh.github.io/lzu/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingYutong Kou, Jin Gao, Bing Li, Gang Wang 等NeurIPS 2023 · 被引用 74 次
- FouriDown: Factoring Down-Sampling into Shuffling and SuperposingQi Zhu, Man Zhou, Jie Huang, Naishan Zheng 等NeurIPS 2023 · 被引用 12 次
- Variance-Insensitive and Target-Preserving Mask Refinement for Interactive Image SegmentationChaowei Fang, Ziyin Zhou, Junye Chen, Hanjing Su 等AAAI 2024 · 被引用 7 次
- Foveated Instance SegmentationHongyi Zeng, Wenxuan Liu, Tianhua Xia, Jinhui Chen 等CVPR 2025
它引用的顶会 Paper6
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee 等ICCV 2019 · 被引用 107 次
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 被引用 58 次
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 被引用 45 次
- Learning to Downsample for Segmentation of Ultra-High Resolution ImagesChen Jin, Ryutaro Tanno, Thomy Mertzanidou, Eleftheria Panagiotaki 等ICLR 2022 · 被引用 40 次
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora 等CVPR 2020
相关 Paper
- Parallax-Tolerant Unsupervised Deep Image StitchingLang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu 等ICCV 2023 · 被引用 111 次
- RAW-Domain Degradation Models for Realistic Smartphone Super-ResolutionAli Mosleh, Faraz Ali, Fengjia Zhang, Stavros Tsogkas 等CVPR 2026
- FeatUp: A Model-Agnostic Framework for Features at Any ResolutionStephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann 等ICLR 2024 · 被引用 117 次
- LinK: Linear Kernel for LiDAR-based 3D PerceptionTao Lu, Xiang Ding, Haisong Liu, Gangshan Wu 等CVPR 2023
- LDA-AQU: Adaptive Query-guided Upsampling via Local Deformable AttentionZewen Du, Zhenjiang Hu, Guiyu Zhao, Ying Jin 等ACM MM 2024 · 被引用 4 次
