Learning to Resize Images for Computer Vision Tasks
Hossein Talebi, Peyman Milanfar
摘要
For all the ways convolutional neural nets have revolutionized computer vision in recent years, one important aspect has received surprisingly little attention: the effect of image size on the accuracy of tasks being trained for. Typically, to be efficient, the input images are resized to a relatively small spatial resolution (e.g. 224 × 224), and both training and inference are carried out at this resolution. The actual mechanism for this re-scaling has been an afterthought: Namely, off-the-shelf image re sizers such as bilinear and bicubic are commonly used in most machine learning software frameworks. But do these re sizers limit the on-task performance of the trained networks? The answer is yes. Indeed, we show that the typical linear re sizer can be replaced with learned resizers that can substantially improve performance. Importantly, while the classical re-sizers typically result in better perceptual quality of the downscaled images, our proposed learned resizers do not necessarily give better visual quality, but instead improve task performance.Our learned image resizer is jointly trained with a base-line vision model. This learned CNN-based resizer creates machine friendly visual manipulations that lead to a consistent improvement of the end task metric over the baseline model. Specifically, here we focus on the classification task with the ImageNet dataset [26], and experiment with four different models to learn resizers adapted to each model. Moreover, we show that the proposed resizer can also be useful for fine-tuning the classification baselines for other vision tasks. To this end, we experiment with three different baselines to develop image quality assessment (IQA) models on the AVA dataset [24].
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Learning Strides in Convolutional Neural NetworksRachid Riad, Olivier Teboul, David Grangier, Neil ZeghidourICLR 2022 · 被引用 54 次
- FastFlow: Accelerating Deep Learning Model Training with Smart Offloading of Input Data PipelineTaegeon Um, Byungsoo Oh, Byeongchan Seo, Minhyeok Kweun 等VLDB 2023 · 被引用 45 次
- Learning to Downsample for Segmentation of Ultra-High Resolution ImagesChen Jin, Ryutaro Tanno, Thomy Mertzanidou, Eleftheria Panagiotaki 等ICLR 2022 · 被引用 40 次
- Constructive Distortion: Improving MLLMs with Attention-Guided Image WarpingDwip Dalal, Gautam Vashishtha, Utkarsh Mishra, Jeonghwan Kim 等ICLR 2026 · 被引用 17 次
- MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any ResolutionWenzhuo Liu, Fei Zhu, Shijie Ma, Cheng-Lin LiuNeurIPS 2024 · 被引用 17 次
它引用的顶会 Paper4
- Toward Real-World Single Image Super-Resolution: A New Benchmark and a New ModelJianrui Cai, Hui Zeng, Hongwei Yong, Zisheng Cao 等ICCV 2019 · 被引用 713 次
- Kernel Modeling Super-Resolution on Real Low-Resolution ImagesRuofan Zhou, Sabine SüsstrunkICCV 2019 · 被引用 149 次
- Dual Directed Capsule Network for Very Low Resolution Image RecognitionManeet Singh, Shruti Nagpal, Richa Singh, Mayank VatsaICCV 2019 · 被引用 56 次
- ThumbNet: One Thumbnail Image Contains All You Need for RecognitionChen Zhao, Bernard GhanemACM MM 2020 · 被引用 14 次
相关 Paper
- MULLER: Multilayer Laplacian Resizer for VisionZhengzhong Tu, Peyman Milanfar, Hossein TalebiICCV 2023 · 被引用 8 次
- MUSIQ: Multi-scale Image Quality TransformerJunjie Ke, Qifei Wang, Yilin Wang, Peyman Milanfar 等ICCV 2021 · 被引用 1,325 次
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 被引用 135 次
- Image Intrinsic Scale Assessment: Bridging the Gap Between Quality and ResolutionVlad Hosu, Lorenzo Agnolucci, Daisuke Iso, Dietmar SaupeICCV 2025 · 被引用 1 次
- Learning Filter Basis for Convolutional Neural Network CompressionYawei Li, Shuhang Gu, Luc Van Gool, Radu TimofteICCV 2019 · 被引用 106 次
