Faster Vision Transformers with Adaptive Patches
Rohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang, László A. Jeni, Kris Kitani
摘要
Vision Transformers (ViTs) partition input images into uniformly sized patches regardless of their content, resulting in long input sequence lengths for high-resolution images. We present Adaptive Patch Transformers (APT), which addresses this by using multiple different patch sizes within the same image. APT reduces the total number of input tokens by allocating larger patch sizes in more homogeneous areas and smaller patches in more complex ones. APT achieves a drastic speedup in ViT inference and training, increasing throughput by 40% on ViT-L and 50% on ViT-H while maintaining downstream performance. It can be applied to a previously fine-tuned ViT and converges in as little as 1 epoch. It also significantly reduces training and inference time without loss of performance in high-resolution dense visual tasks, achieving up to 30% faster training and inference in visual QA, object detection, and semantic segmentation. We will release all code and trained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian SplattingArthur Moreau, Richard Shaw, Michal Nazarczuk, Jisu Shin 等CVPR 2026 · 被引用 10 次
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion TransformersDahye Kim, Deepti Ghadiyaram, Raghudeep GaddeCVPR 2026 · 被引用 3 次
- Token Warping Helps MLLMs Look from Nearby ViewpointsPhillip Y. Lee, Chanho Park, Mingue Park, Seungwoo Yoo 等CVPR 2026 · 被引用 1 次
- Content-Aware Dynamic Patchification for Efficient Video DiffusionSheng Li, Connelly Barnes, Mamshad Nayeem Rizve, Hongwu Peng 等CVPR 2026
- Adaptive Volumetric Mechanical Property Fields Invariant to ResolutionRishit Dagli, Donglai Xiang, Vismay Modi, Xuning Yang 等ICML 2026
它引用的顶会 Paper47
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
相关 Paper
- MSPE: Multi-Scale Patch Embedding Prompts Vision Transformers to Any ResolutionWenzhuo Liu, Fei Zhu, Shijie Ma, Cheng-Lin LiuNeurIPS 2024 · 被引用 17 次
- AiluRus: A Scalable ViT Framework for Dense PredictionJin Li, Yaoming Wang, Xiaopeng Zhang, Bowen Shi 等NeurIPS 2023 · 被引用 18 次
- FlexiViT: One Model for All Patch SizesLucas Beyer, Pavel Izmailov, Alexander Kolesnikov, Mathilde Caron 等CVPR 2023
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang 等NeurIPS 2021 · 被引用 283 次
