Adapting Generative Pretrained Language Model for Open-domain Multimodal Sentence Summarization
Dengtian Lin, Liqiang Jing, Xuemeng Song, Meng Liu, Teng Sun, Liqiang Nie
摘要
Multimodal sentence summarization, aiming to generate a brief summary of the source sentence and image, is a new yet challenging task. Although existing methods have achieved compelling success, they still suffer from two key limitations: 1) lacking the adaptation of generative pre-trained language models for open-domain MMSS, and 2) lacking the explicit critical information modeling. To address these limitations, we propose a BART-MMSS framework, where BART is adopted as the backbone. To be specific, we propose a prompt-guided image encoding module to extract the source image feature. It leverages several soft to-be-learned prompts for image patch embedding, which facilitates the visual content injection to BART for open-domain MMSS tasks. Thereafter, we devise an explicit source critical token learning module to directly capture the critical tokens of the source sentence with the reference of the source image, where we incorporate explicit supervision to improve performance. Extensive experiments on a public dataset fully validate the superiority of our proposed method. In addition, the predicted tokens by the vision-guided key-token highlighting module can be easily understood by humans and hence improve the interpretability of our model.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper3
- CFIR: Fast and Effective Long-Text To Image Retrieval for Large CorporaZijun Long, Xuri Ge, Richard McCreadie, Joemon M. JoseSIGIR 2024 · 被引用 10 次
- Exploring the Trade-Off within Visual Information for MultiModal Sentence SummarizationMinghuan Yuan, Shiyao Cui, Xinghua Zhang, Shicheng Wang 等SIGIR 2024 · 被引用 3 次
- Language Constrained Multimodal Hyper Adapter For Many-to-Many Multimodal SummarizationNayu Liu, Fanglong Yao, Haoran Luo, Yong Yang 等ACL 2025
相关 Paper
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 等AAAI 2022 · 被引用 61 次
- Fine-tuning with Multi-modal Entity Prompts for News Image CaptioningJingjing Zhang, Shancheng Fang, Zhendong Mao, Zhiwei Zhang 等ACM MM 2022 · 被引用 16 次
- From Sights to Insights: Towards Summarization of Multimodal Clinical DocumentsAkash Ghosh, Mohit Tomar, Abhisek Tiwari, Sriparna Saha 等ACL 2024 · 被引用 5 次
- V2P: Vision-to-Prompt based Multi-Modal Product Summary GenerationXuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao 等SIGIR 2022 · 被引用 25 次
- KM-BART: Knowledge Enhanced Multimodal BART for Visual Commonsense GenerationYiran Xing, Zai Shi, Zhao Meng, Gerhard Lakemeyer 等ACL 2021
