Adapting Generative Pretrained Language Model for Open-domain Multimodal Sentence Summarization
Dengtian Lin, Liqiang Jing, Xuemeng Song, Meng Liu, Teng Sun, Liqiang Nie
Abstract
Multimodal sentence summarization, aiming to generate a brief summary of the source sentence and image, is a new yet challenging task. Although existing methods have achieved compelling success, they still suffer from two key limitations: 1) lacking the adaptation of generative pre-trained language models for open-domain MMSS, and 2) lacking the explicit critical information modeling. To address these limitations, we propose a BART-MMSS framework, where BART is adopted as the backbone. To be specific, we propose a prompt-guided image encoding module to extract the source image feature. It leverages several soft to-be-learned prompts for image patch embedding, which facilitates the visual content injection to BART for open-domain MMSS tasks. Thereafter, we devise an explicit source critical token learning module to directly capture the critical tokens of the source sentence with the reference of the source image, where we incorporate explicit supervision to improve performance. Extensive experiments on a public dataset fully validate the superiority of our proposed method. In addition, the predicted tokens by the vision-guided key-token highlighting module can be easily understood by humans and hence improve the interpretability of our model.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers3
- CFIR: Fast and Effective Long-Text To Image Retrieval for Large CorporaZijun Long, Xuri Ge, Richard McCreadie, Joemon M. JoseSIGIR 2024 · 10 citations
- Exploring the Trade-Off within Visual Information for MultiModal Sentence SummarizationMinghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang et al.SIGIR 2024 · 3 citations
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang et al.ACL 2025
Related papers
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang et al.ACM MM 2022 · 16 citations
- From Sights to Insights: Towards Summarization of Multimodal Clinical DocumentsAkash Ghosh, Mohit Tomar, Abhisek Tiwari, Sriparna Saha et al.ACL 2024 · 5 citations
- V2P: Vision-to-Prompt based Multi-Modal Product Summary GenerationXuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao et al.SIGIR 2022 · 25 citations
- KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense GenerationYiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer et al.ACL 2021
