Exploring the Trade-Off within Visual Information for MultiModal Sentence Summarization
Minghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang, Hongbo Xu, Tingwen Liu
Abstract
MultiModal Sentence Summarization (MMSS) aims to generate a brief summary based on the given source sentence and its associated image. Previous studies on MMSS have achieved success by either selecting the task-relevant visual information or filtering out the task-irrelevant visual information to help the textual modality to generate the summary. However, enhancing from a single perspective usually introduces over-preservation or over-compression problems. To tackle these issues, we resort to Information Bottleneck (IB), which seeks to find a maximally compressed mapping of the input information that preserves as much information about the target as possible. Specifically, we propose a novel method, T(3), which adopts IB to balance the Trade-off between Task-relevant and Task-irrelevant visual information through the variational inference framework. In this way, the task-irrelevant visual information is compressed to the utmost while the task-relevant visual information is maximally retained. With the holistic perspective, the generated summary could maintain as many key elements as possible while discarding the unnecessary ones as far as possible. Extensive experiments on the representative MMSS dataset demonstrate the superiority of our proposed method. Our code is available at https://github.com/YuanMinghuan/T3.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Calibrate Before Use: Improving Few-shot Performance of Language ModelsZihao Zhao, Eric Wallace, Shi Feng, Dan Klein et al.ICML 2021 · 1,843 citations
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
Related papers
- Adapting Generative Pretrained Language Model for Open-domain Multimodal Sentence SummarizationDengtian Lin, Liqiang Jing, Xuemeng Song, Meng Liu et al.SIGIR 2023 · 15 citations
- Learning Optimal Multimodal Information Bottleneck RepresentationsQilong Wu, Yiyang Shao, Jun Wang, Xiaobo SunICML 2025
- DIUSum: Dynamic Image Utilization for Multimodal SummarizationMin Xiao, Junnan Zhu, Feifei Zhai, Yu Zhou et al.AAAI 2024 · 10 citations
- Information Screening whilst Exploiting! Multimodal Relation Extraction with Feature Denoising and Multimodal Topic ModelingShengqiong Wu, Hao Fei, Yixin Cao, Lidong Bing et al.ACL 2023 · 60 citations
- IBMA: Information Bottleneck-Based Multimodal AlignmentYancheng Wang, Zeyu Dong, Dongfang Sun, Alvin Silva et al.ICML 2026
