Distilled Dual-Encoder Model for Vision-Language Understanding
Zekun Wang, Wenhui Wang, Haichao Zhu, Ming Liu, Bing Qin, Furu Wei
摘要
On vision-language understanding (VLU) tasks, fusion-encoder vision-language models achieve superior results but sacrifice efficiency because of the simultaneous encoding of images and text. On the contrary, the dual encoder model that separately encodes images and text has the advantage in efficiency, while failing on VLU tasks due to the lack of deep cross-modal interactions. To get the best of both worlds, we propose DiDE, a framework that distills the knowledge of the fusion-encoder teacher model into the dual-encoder student model. Since the cross-modal interaction is the key to the superior performance of teacher model but is absent in the student model, we encourage the student not only to mimic the predictions of teacher, but also to calculate the cross-modal attention distributions and align with the teacher. Experimental results demonstrate that DiDE is competitive with the fusion-encoder teacher model in performance (only a 1% drop) while enjoying 4 times faster inference. Further analyses reveal that the proposed cross-modal attention distillation is crucial to the success of our framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- BridgeTower: Building Bridges between Encoders in Vision-Language Representation LearningXiao Xu, Chenfei Wu, Shachar Rosenman, Vasudev Lal 等AAAI 2023 · 被引用 99 次
- MixReorg: Cross-Modal Mixed Patch Reorganization is a Good Mask Learner for Open-World Semantic SegmentationKaixin Cai, Pengzhen Ren, Yi Zhu, Hang Xu 等ICCV 2023 · 被引用 22 次
- Module-wise Adaptive Distillation for Multimodality Foundation ModelsChen Liang, Jiahui Yu, Ming-Hsuan Yang, Matthew Brown 等NeurIPS 2023 · 被引用 17 次
- ViewCo: Discovering Text-Supervised Segmentation Masks via Multi-View Semantic ConsistencyPengzhen Ren, Changlin Li, Hang Xu, Yi Zhu 等ICLR 2023 · 被引用 16 次
- mCLIP: Multilingual CLIP via Cross-lingual TransferGuanhua Chen, Lu Hou, Yun Chen, Wenliang Dai 等ACL 2023 · 被引用 13 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen 等ICML 2021 · 被引用 5,401 次
相关 Paper
- KAID: Knowledge-Aware Interactive Distillation for Vision-Language ModelsDa Zhang, Feiyu Wang, Bingyu Li, Zhiyuan Zhao 等ACM MM 2025 · 被引用 10 次
- HieRD: Hierarchical Relational Distillation for Vision-Language Embedding ModelsVinh Le, Nguyen Dang, Tu Vu, Linh Van 等ICML 2026
- Align-KD: Distilling Cross-Modal Alignment Knowledge for Mobile Vision-Language Large Model EnhancementQianhan Feng, Wenshuo Li, Tong Lin, Xinghao ChenCVPR 2025
- Expediting Contrastive Language-Image Pretraining via Self-Distilled EncodersBumsoo Kim, Jinhyung Kim, Yeonsik Jo, Seung Hwan KimAAAI 2024 · 被引用 5 次
- No Head Left Behind - Multi-Head Alignment Distillation for TransformersTianyang Zhao, Kunwar Yashraj Singh, Srikar Appalaraju, Peng Tang 等AAAI 2024 · 被引用 5 次
