Learning to Zoom and Unzoom
Chittesh Thavamani, Mengtian Li, Francesco Ferroni, Deva Ramanan
Abstract
Many perception systems in mobile computing, autonomous navigation, and AR/VR face strict compute constraints that are particularly challenging for high-resolution input images. Previous works propose nonuniform downsamplers that "learn to zoom" on salient image regions, reducing compute while retaining task-relevant image information. However, for tasks with spatial labels (such as 2D/3D object detection and semantic segmentation), such distortions may harm performance. In this work (LZU), we "learn to zoom" in on the input image, compute spatial features, and then "unzoom" to revert any deformations. To enable efficient and differentiable unzooming, we approximate the zooming warp with a piecewise bilinear mapping that is invertible. LZU can be applied to any task with 2D spatial input and any model with 2D spatial features, and we demonstrate this versatility by evaluating on a variety of tasks and datasets: object detection on Argoverse-HD, semantic segmentation on Cityscapes, and monocular 3D object detection on nuScenes. Interestingly, we observe boosts in performance even when high-resolution sensor data is unavailable, implying that LZU can be used to "learn to upsample" as well. Code and additional visuals are available at https://tchittesh.github.io/lzu/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a84b23e8-aa77-4eef-8565-f87e238cd9f5Cited by top-tier papers4
- ZoomTrack: Target-aware Non-uniform Resizing for Efficient Visual TrackingYutong Kou, Jin Gao, Bing Li, Gang Wang et al.NeurIPS 2023 · 74 citations
- FouriDown: Factoring Down-Sampling into Shuffling and SuperposingQi Zhu, Man Zhou, Jie Huang, Naishan Zheng et al.NeurIPS 2023 · 12 citations
- Variance-Insensitive and Target-Preserving Mask Refinement for Interactive Image SegmentationChaowei Fang, Ziyin Zhou, Junye Chen, Hanjing Su et al.AAAI 2024 · 7 citations
- Foveated Instance SegmentationHongyi Zeng, Wenxuan Liu, Tianhua Xia, Jinhui Chen et al.CVPR 2025
Builds on6
- Efficient Segmentation: Learning Downsampling Near Semantic BoundariesDmitrii Marin, Zijian He, Peter Vajda, Priyam Chatterjee et al.ICCV 2019 · 107 citations
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 58 citations
- FOVEA: Foveated Image Magnification for Autonomous NavigationChittesh Thavamani, Mengtian Li, Nicolas Cebron, Deva RamananICCV 2021 · 45 citations
- Learning to Downsample for Segmentation of Ultra-High Resolution ImagesChen Jin, Ryutaro Tanno, Thomy Mertzanidou, Eleftheria Panagiotaki et al.ICLR 2022 · 40 citations
- nuScenes: A Multimodal Dataset for Autonomous DrivingHolger Caesar, Varun Bankiti, Alex H. Lang, Sourabh Vora et al.CVPR 2020
Related papers
- Parallax-Tolerant Unsupervised Deep Image StitchingLang Nie, Chunyu Lin, Kang Liao, Shuaicheng Liu et al.ICCV 2023 · 111 citations
- RAW-Domain Degradation Models for Realistic Smartphone Super-ResolutionAli Mosleh, Faraz Ali, Fengjia Zhang, Stavros Tsogkas et al.CVPR 2026
- FeatUp: A Model-Agnostic Framework for Features at Any ResolutionStephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann et al.ICLR 2024 · 117 citations
- LinK: Linear Kernel for LiDAR-based 3D PerceptionTao Lu, Xiang Ding, Haisong Liu, Gangshan Wu et al.CVPR 2023
- LDA-AQU: Adaptive Query-guided Upsampling via Local Deformable AttentionZewen Du, Zhenjiang Hu, Guiyu Zhao, Ying Jin et al.ACM MM 2024 · 4 citations
