Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision Recognition
Kai Zhao, Liting Ruan, Haoran Jiang, Xiaoqiang Zhu, Xianchao Zhang, Dan Zeng
Abstract
Images are typically sampled on a uniform grid, despite their non-uniform information distribution-some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsampling methods that adaptively downsample images based on predicted per-pixel importance, allocating more pixels to informative areas. However, these methods require high-resolution processing to accurately estimate importance, which undermines their efficiency: the prediction itself must process the full-resolution image, consuming most of the computational budget. This high-resolution importance prediction is necessary because each input may differ significantly in structure and content. In this paper, we take a different approach and introduce a learn-to-downsample paradigm tailored for aligned vision recognition tasks, such as face recognition and palmprint recognition, where input alignment ensures consistent spatial structure across images. This alignment ensures structural consistency across images, allowing a shared, input-agnostic downsampling template applicable to all inputs. Furthermore, instead of relying on implicit importance maps, we introduce a flow-based representation that explicitly models the spatial warping from the original image to the downsampled version. The flow representation is not only more efficient but also more controllable: we regularize the flow using its Jacobian determinant to precisely control the sampling density and coverage, enabling interpretable and tunable sampling patterns. Extensive experiments on two aligned recognition tasks, face and palmprint recognition, demonstrate that our method substantially reduces computational cost with minimal accuracy degradation, achieving a significantly better performance-efficiency trade-off than existing predictive downsampling methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 80e24d64-32ab-48c7-a48b-32a34aef3737Builds on9
- AdaFace: Quality Adaptive Margin for Face RecognitionMinchul Kim, Anil K. Jain, Xiaoming LiuCVPR 2022 · 509 citations
- Learning to Downsample for Segmentation of Ultra-High Resolution ImagesChen Jin, Ryutaro Tanno, Thomy Mertzanidou, Eleftheria Panagiotaki et al.ICLR 2022 · 40 citations
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao et al.NeurIPS 2024 · 39 citations
- RPG-Palm: Realistic Pseudo-data Generation for Palmprint RecognitionLei Shen, Jianlong Jin, Ruixin Zhang, Huaen Li et al.ICCV 2023 · 18 citations
- PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint GenerationJianlong Jin, Lei Shen, Ruixin Zhang, Chenglong Zhao et al.AAAI 2024 · 18 citations
Related papers
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local AttendersMatthew Walmer, Saksham Suri, Anirud Aggarwal, Abhinav ShrivastavaCVPR 2026 · 2 citations
- Learning Optical Flow From a Few MatchesShihao Jiang, Yao Lu, Hongdong Li, Richard HartleyCVPR 2021
- Adapting Dense Matching for Homography Estimation with Grid-based AccelerationKaining Zhang, Yuxin Deng, Jiayi Ma, Paolo FavaroCVPR 2025
- Superpixel Segmentation With Fully Convolutional NetworksFengting Yang, Qian Sun, Hailin Jin, Zihan ZhouCVPR 2020
- Resolution Adaptive Networks for Efficient InferenceLe Yang, Yizeng Han, Xi Chen, Shiji Song et al.CVPR 2020
