Learning to Jointly Share and Prune Weights for Grounding Based Vision and Language Models
Shangqian Gao, Burak Uzkent, Yilin Shen, Heng Huang, Hongxia Jin
Abstract
Transformers have seen growing interest in processing different modalities, including language and image data. As a result, we can process vision and language data using transformers that are architecturally similar. Leveraging this feature of transformers, we propose weight sharing across two transformer backbones and within the same transformer backbone and pruning across two backbones in a unified framework. More specifically, we investigate weight sharing and pruning for two components of the transformers: (1) Multi-Head Attention (MSA) and (2) Feed-Forward Network (FFN) layers. To jointly perform weight sharing and pruning, we propose to use a regularization term to align model weights and the desired structure during the multimodal pre-training step. The structure vectors of sharing and pruning are generated by using a hypernetwork, which can capture complex interactions between pruning and sharing across layers and modalities. We train the hypernetwork and model weights iteratively so that the learned structure evolves along with model weights. After minimizing the proposed objective in the pre-training step, we perform weight sharing and pruning and fine-tune the compressed model on downstream tasks. Finally, we perform experiments on vision and language tasks, including Referring Expression Comprehension (REC), Visual Question Answering (VQA), and Object Detection using the state-of-the-art grounding based models: MDETR and GLIP. Our experiments show that we can compress these models by by sharing and pruning MSA and FFN weights without almost any loss in accuracy.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers8
- PolicyCleanse: Backdoor Detection and Mitigation for Competitive Reinforcement LearningJunfeng Guo, Ang Li, Lixu Wang, Cong LiuICCV 2023 · 27 citations
- Structural Alignment for Network Pruning through Partial RegularizationShangqian Gao, Zeyu Zhang, Yanfu Zhang, Feihu Huang et al.ICCV 2023 · 26 citations
- InstructDET: Diversifying Referring Object Detection with Generalized InstructionsRonghao Dang, Jiangyan Feng, Haodong Zhang, Chongjian Ge et al.ICLR 2024 · 16 citations
- BilevelPruning: Unified Dynamic and Static Channel Pruning for Convolutional Neural NetworksShangqian Gao, Yanfu Zhang, Feihu Huang, Heng HuangCVPR 2024
- Jointly Training and Pruning CNNs via Learnable Agent Guidance and AlignmentAlireza Ganjdanesh, Shangqian Gao, Heng HuangCVPR 2024
Related papers
- Dynamic Inference with Grounding Based Vision and Language ModelsBurak Uzkent, Amanmeet Garg, Wentao Zhu, Keval Doshi et al.CVPR 2023
- UPop: Unified and Progressive Pruning for Compressing Vision-Language TransformersDachuan Shi, Chaofan Tao, Ying Jin, Zhendong Yang et al.ICML 2023 · 64 citations
- Coarse-to-Fine Vision-Language Pre-training with Fusion in the BackboneZi-Yi Dou, Aishwarya Kamath, Zhe Gan, Pengchuan Zhang et al.NeurIPS 2022 · 173 citations
- MADTP: Multimodal Alignment-Guided Dynamic Token Pruning for Accelerating Vision-Language TransformerJianjian Cao, Peng Ye, Shengze Li, Chong Yu et al.CVPR 2024
- EVE: Efficient Vision-Language Pre-training with Masked Prediction and Modality-Aware MoEJunyi Chen, Longteng Guo, Jia Sun, Shuai Shao et al.AAAI 2024 · 25 citations
