Expanding Large Pre-trained Unimodal Models with Multimodal Information Injection for Image-Text Multimodal Classification
Tao Liang, Guosheng Lin, Mingyang Wan, Tianrui Li, Guojun Ma, Fengmao Lv
Abstract
Fine-tuning pre-trained models for downstream tasks is mainstream in deep learning. However, the pre-trained models are limited to be fine-tuned by data from a specific modality. For example, as a visual model, DenseNet cannot directly take the textual data as its input. Hence, although the large pre-trained models such as DenseNet or BERT have a great potential for the downstream recognition tasks, they have weaknesses in leveraging multimodal information, which is a new trend of deep learning. This work focuses on fine-tuning pre-trained unimodal models with multimodal inputs of image-text pairs and expanding them for image-text multimodal recognition. To this end, we propose the Multimodal Information Injection Plug-in (MI2P) which is attached to different layers of the unimodal models (e.g., DenseNet and BERT). The proposed MI2P unit provides the path to integrate the information of other modalities into the unimodal models. Specifically, MI2P performs cross-modal feature transformation by learning the fine-grained correlations between the visual and textual features. Through the proposed MI2P unit, we can inject the language information into the vision backbone by attending the word-wise textual features to different visual channels, as well as inject the visual information into the language backbone by attending the channel-wise visual features to different textual words. Armed with the MI2P attachments, the pre-trained unimodal models can be expanded to process multimodal data without the need to change the network structures.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- ECoFLaP: Efficient Coarse-to-Fine Layer-Wise Pruning for Vision-Language ModelsYi-Lin Sung, Jaehong Yoon, Mohit BansalICLR 2024 · 22 citations
- SimMLM: A Simple Framework for Multi-Modal Learning with Missing ModalitySijie Li, Chen Chen, Jungong HanICCV 2025 · 14 citations
- C3Net: Compound Conditioned ControlNet for Multimodal Content GenerationJuntao Zhang, Yuehuai Liu, Yu-Wing Tai, Chi-Keung TangCVPR 2024 · 4 citations
- Semi-Supervised Multimodal Classification Through Learning from Modal and Strategic ComplementaritiesJunchi Chen, Richong Zhang, Junfan ChenAAAI 2025 · 1 citation
- Seek Common Ground While Reserving Differences: Semi-Supervised Image-Text Sentiment RecognitionWuyou Xia, Guoli Jia, Sicheng Zhao, Jufeng YangCVPR 2025
Builds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- Unicoder-VL: A Universal Encoder for Vision and Language by Cross-Modal Pre-TrainingGen Li, Nan Duan, Yuejian Fang, Ming Gong et al.AAAI 2020 · 966 citations
Related papers
- mPLUG: Effective and Efficient Vision-Language Learning by Cross-modal Skip-connectionsChenliang Li, Haiyang Xu, Junfeng Tian, Wei Wang et al.EMNLP 2022 · 159 citations
- CM-BERT: Cross-Modal BERT for Text-Audio Sentiment AnalysisKaicheng Yang, Hua Xu, Kai GaoACM MM 2020 · 129 citations
- mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and VideoHaiyang Xu, Qinghao Ye, Ming Yan, Yaya Shi et al.ICML 2023 · 237 citations
- E2E-VLP: End-to-End Vision-Language Pre-training Enhanced by Visual LearningHaiyang Xu, Ming Yan, Chenliang Li, Bin Bi et al.ACL 2021
- CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained KnowledgeLinli Yao, Weijing Chen, Qin JinWWW 2023 · 11 citations
