Module-wise Adaptive Distillation for Multimodality Foundation Models
Chen Liang, Jiahui Yu, Ming-Hsuan Yang, Matthew Brown, Yin Cui, Tuo Zhao, Boqing Gong, Tianyi Zhou
Abstract
Pre-trained multimodal foundation models have demonstrated remarkable generalizability but pose challenges for deployment due to their large sizes. One effective approach to reducing their sizes is layerwise distillation, wherein small student models are trained to match the hidden representations of large teacher models at each layer. Motivated by our observation that certain architecture components, referred to as modules, contribute more significantly to the student's performance than others, we propose to track the contributions of individual modules by recording the loss decrement after distillation each module and choose the module with a greater contribution to distill more frequently. Such an approach can be naturally formulated as a multi-armed bandit (MAB) problem, where modules and loss decrements are considered as arms and rewards, respectively. We then develop a modified-Thompson sampling algorithm named OPTIMA to address the nonstationarity of module contributions resulting from model updating. Specifically, we leverage the observed contributions in recent history to estimate the changing contribution of each module and select modules based on these estimations to maximize the cumulative contribution. We evaluate the effectiveness of OPTIMA through distillation experiments on various multimodal understanding and image captioning tasks, using the CoCa-Large model [48] as the teacher model. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a146c8a2-fec1-4f10-91b1-ecb0cb2d72e1Cited by top-tier papers6
- CLIP-KD: An Empirical Study of CLIP Model DistillationChuanguang Yang, Zhulin An, Libo Huang, Junyu Bi et al.CVPR 2024 · 50 citations
- Embodied Multi-Modal Agent trained by an LLM from a Parallel TextWorldYijun Yang, Tianyi Zhou, Kanxue Li, Dapeng Tao et al.CVPR 2024 · 23 citations
- TRACER: Persistent Regularization for Robust Multimodal FinetuningHesam Asadollahzadeh, Feng Liu, Christopher Leckie, Sarah ErfaniICML 2026
- Active Data Curation Effectively Distills Large-Scale Multimodal ModelsVishaal Udandarao, Nikhil Parthasarathy, Muhammad Ferjad Naeem, Talfan Evans et al.CVPR 2025
- Hallusionbench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language ModelsTianrui Guan, Fuxiao Liu, Xiyang Wu, Ruiqi Xian et al.CVPR 2024
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
Related papers
- LLaVA-MoD: Making LLaVA Tiny via MoE-Knowledge DistillationFangxun Shu, Yue Liao, Lei Zhang, Le Zhuo et al.ICLR 2025
- Masking Teacher and Reinforcing Student for Distilling Vision-Language ModelsByung-Kwan Lee, Yu-Chiang Frank Wang, Ryo HachiumaCVPR 2026 · 7 citations
- Increasing Model Capacity for Free: A Simple Strategy for Parameter Efficient Fine-tuningHaobo Song, Hao Zhao, Soumajit Majumder, Tao LinICLR 2024 · 11 citations
- Progressive Ensemble Distillation: Building Ensembles for Efficient InferenceDon Kurian Dennis, Abhishek Shetty, Anish Prasad Sevekari, Kazuhito Koishida et al.NeurIPS 2023
- Self-Evolutionary Reinforced Knowledge Distillation for Multi-Modal Tool-Use AgentsLei Shen, Chengyu Wang, Yuanjie Lyu, Yuanhao Yue et al.KDD 2026
