Dynamic Inference with Grounding Based Vision and Language Models
Burak Uzkent, Amanmeet Garg, Wentao Zhu, Keval Doshi, Jingru Yi, Xiaolong Wang, Mohamed Omar
Abstract
Transformers have been recently utilized for vision and language tasks successfully. For example, recent image and language models with more than 200M parameters have been proposed to learn visual grounding in the pre-training step and show impressive results on downstream vision and language tasks. On the other hand, there exists a large amount of computational redundancy in these large models which skips their run-time efficiency. To address this problem, we propose dynamic inference for grounding based vision and language models conditioned on the input imagetext pair. We first design an approach to dynamically skip multihead self-attention and feed forward network layers across two backbones and multimodal network. Additionally, we propose dynamic token pruning and fusion for two backbones. In particular, we remove redundant tokens at different levels of the backbones and fuse the image tokens with the language tokens in an adaptive manner. To learn policies for dynamic inference, we train agents using reinforcement learning. In this direction, we replace the CNN backbone in a recent grounding-based vision and language model, MDETR, with a vision transformer and call it ViT-MDETR. Then, we apply our dynamic inference method to ViTMDETR, called D-ViTDMETR, and perform experiments on image-language tasks. Our results show that we can improve the run-time efficiency of the state-of-the-art models MDETR and GLIP by up to ∼ 50% on Referring Expression Comprehension and Segmentation, and VQA with only maximum ∼ 0.3% accuracy drop.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Layer-wise Alignment: Examining Safety Alignment Across Image Encoder Layers in Vision Language ModelsSaketh Bachu, Erfan Shayegani, Rohit Lal, Trishna Chakraborty et al.ICML 2025
- AdaSpark: Adaptive Sparsity for Efficient Long-Video UnderstandingHandong Li, Zikang Liu, Longteng Guo, Tongtian Yue et al.CVPR 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
Related papers
- Learning to Jointly Share and Prune Weights for Grounding Based Vision and Language ModelsShangqian Gao, Burak Uzkent, Yilin Shen, Heng Huang et al.ICLR 2023
- Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained ModelsQiong Wu, Wei Yu, Yiyi Zhou, Shubin Huang et al.NeurIPS 2023 · 16 citations
- Not All Images are Worth 16x16 Words: Dynamic Transformers for Efficient Image RecognitionYulin Wang, Rui Huang, Shiji Song, Zeyi Huang et al.NeurIPS 2021 · 283 citations
- ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative PruningWen Luo, Peng Chen, Xiaotao Huang, LiQun HuangAAAI 2026
- Skip-It? Theoretical Conditions for Layer Skipping in Vision–Language ModelsMax Hartman, Vidhata Jayaraman, Moulik Choraria, Akhil Bhimaraju et al.ICML 2026 · 1 citation
