Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision Recognition
Kai Zhao, Liting Ruan, Haoran Jiang, Xiaoqiang Zhu, Xianchao Zhang, Dan Zeng
摘要
Images are typically sampled on a uniform grid, despite their non-uniform information distribution-some regions are rich in content while others are not. The mismatch leads to inefficient computation allocation in deep learning models. To address this, recent studies have proposed predictive downsampling methods that adaptively downsample images based on predicted per-pixel importance, allocating more pixels to informative areas. However, these methods require high-resolution processing to accurately estimate importance, which undermines their efficiency: the prediction itself must process the full-resolution image, consuming most of the computational budget. This high-resolution importance prediction is necessary because each input may differ significantly in structure and content. In this paper, we take a different approach and introduce a learn-to-downsample paradigm tailored for aligned vision recognition tasks, such as face recognition and palmprint recognition, where input alignment ensures consistent spatial structure across images. This alignment ensures structural consistency across images, allowing a shared, input-agnostic downsampling template applicable to all inputs. Furthermore, instead of relying on implicit importance maps, we introduce a flow-based representation that explicitly models the spatial warping from the original image to the downsampled version. The flow representation is not only more efficient but also more controllable: we regularize the flow using its Jacobian determinant to precisely control the sampling density and coverage, enabling interpretable and tunable sampling patterns. Extensive experiments on two aligned recognition tasks, face and palmprint recognition, demonstrate that our method substantially reduces computational cost with minimal accuracy degradation, achieving a significantly better performance-efficiency trade-off than existing predictive downsampling methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- AdaFace: Quality Adaptive Margin for Face RecognitionMinchul Kim, Anil K. Jain, Xiaoming LiuCVPR 2022 · 被引用 509 次
- Learning to Downsample for Segmentation of Ultra-High Resolution ImagesChen Jin, Ryutaro Tanno, Thomy Mertzanidou, Eleftheria Panagiotaki 等ICLR 2022 · 被引用 40 次
- D-LLM: A Token Adaptive Computing Resource Allocation Strategy for Large Language ModelsYikun Jiang, Huanyu Wang, Lei Xie, Hanbin Zhao 等NeurIPS 2024 · 被引用 39 次
- RPG-Palm: Realistic Pseudo-data Generation for Palmprint RecognitionLei Shen, Jianlong Jin, Ruixin Zhang, Huaen Li 等ICCV 2023 · 被引用 18 次
- PCE-Palm: Palm Crease Energy Based Two-Stage Realistic Pseudo-Palmprint GenerationJianlong Jin, Lei Shen, Ruixin Zhang, Chenglong Zhao 等AAAI 2024 · 被引用 18 次
相关 Paper
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local AttendersMatthew Walmer, Saksham Suri, Anirud Aggarwal, Abhinav ShrivastavaCVPR 2026 · 被引用 2 次
- Learning Optical Flow From a Few MatchesShihao Jiang, Yao Lu, Hongdong Li, Richard HartleyCVPR 2021
- Adapting Dense Matching for Homography Estimation with Grid-based AccelerationKaining Zhang, Yuxin Deng, Jiayi Ma, Paolo FavaroCVPR 2025
- Superpixel Segmentation With Fully Convolutional NetworksFengting Yang, Qian Sun, Hailin Jin, Zihan ZhouCVPR 2020
- Resolution Adaptive Networks for Efficient InferenceLe Yang, Yizeng Han, Xi Chen, Shiji Song 等CVPR 2020
