Memory Efficient Transformer Adapter for Dense Predictions
Dong Zhang, Rui Yan, Pingcheng Dong, Kwang-Ting Cheng
Abstract
While current Vision Transformer (ViT) adapter methods have shown promising accuracy, their inference speed is implicitly hindered by inefficient memory access operations, e.g., standard normalization and frequent reshaping. In this work, we propose META, a simple and fast ViT adapter that can improve the model's memory efficiency and decrease memory time consumption by reducing the inefficient memory access operations. Our method features a memory-efficient adapter block that enables the common sharing of layer normalization between the self-attention and feed-forward network layers, thereby reducing the model's reliance on normalization operations. Within the proposed block, the cross-shaped self-attention is employed to reduce the model's frequent reshaping operations. Moreover, we augment the adapter block with a lightweight convolutional branch that can enhance local inductive biases, particularly beneficial for the dense prediction tasks, e.g., object detection, instance segmentation, and semantic segmentation. The adapter block is finally formulated in a cascaded manner to compute diverse head features, thereby enriching the variety of feature representations. Empirically, extensive evaluations on multiple representative datasets validate that META substantially enhances the predicted quality, while achieving a new state-of-the-art accuracy-efficiency trade-off. Theoretically, we demonstrate that META exhibits superior generalization capability and stronger adaptability.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5c578521-4cff-4e9c-a4e2-28bcb009cc82Cited by top-tier papers3
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 13 citations
- Svit-Split: Unleashing the Power of Vision Foundation Models Via Efficient Splitting HeadsYifan Li, Xin Li, Tianqin Li, Wenbin He et al.ICCV 2025 · 2 citations
- Vision-MoR: Scaling Vision Transformer via Patch-Level Mixture-of-RecursionsYunhong He, Zhengqing Yuan, Weixiang Sun, Yiyang Li et al.AAAI 2026
Builds on32
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
Related papers
- Vision Transformer Adapter for Dense PredictionsZhe Chen, Yuchen Duan, Wenhai Wang, Junjun He et al.ICLR 2023 · 204 citations
- EfficientViT: Memory Efficient Vision Transformer with Cascaded Group AttentionXinyu Liu, Houwen Peng, Ningxin Zheng, Yuqing Yang et al.CVPR 2023
- Dynamic Tuning Towards Parameter and Inference Efficiency for ViT AdaptationWangbo Zhao, Jiasheng Tang, Yizeng Han, Yibing Song et al.NeurIPS 2024 · 41 citations
- Unleashing Vanilla Vision Transformer with Masked Image Modeling for Object DetectionYuxin Fang, Shusheng Yang, Shijie Wang, Yixiao Ge et al.ICCV 2023 · 67 citations
- Parameter-, Memory-, Time-Efficient Multi-Task Dense Vision AdaptationHaiming Yao, Wei Luo, Qiyu Chen, Jianxing Liao et al.AAAI 2026
