Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
Tiezheng Yu, Wenliang Dai, Zihan Liu, Pascale Fung
Abstract
Multimodal abstractive summarization (MAS) models that summarize videos (vision modality) and their corresponding transcripts (text modality) are able to extract the essential information from massive multimodal data on the Internet. Recently, large-scale generative pretrained language models (GPLMs) have been shown to be effective in text generation tasks. However, existing MAS models cannot leverage GPLMs' powerful generation ability. To fill this research gap, we aim to study two research questions: 1) how to inject visual information into GPLMs without hurting their generation ability; and 2) where is the optimal place in GPLMs to inject the visual information? In this paper, we present a simple yet effective method to construct vision guided (VG) GPLMs for the MAS task using attention-based add-on layers to incorporate visual information while maintaining their original text generation ability. Results show that our best model significantly surpasses the prior state-of-the-art model by 5.7 ROUGE-1, 5.3 ROUGE-2, and 5.1 ROUGE-L scores on the How2 dataset (Sanabria et al., 2018) , and our visual guidance method contributes 83.6% of the overall improvement. Furthermore, we conduct thorough ablation studies to analyze the effectiveness of various modality fusion methods and fusion locations. * * The two authors contribute equally. The code is available at: https://github.com/ HLTCHKUST/VG-GPLMs Video Frames Transcript: so now we are going to go over some basics sheet music readings for the key of g flat major. so you noticed the key of g flat, when you are reading real books, there is going to be a treble cleft here. it is going to have 6 flats 1, b flat, e flat, a flat, d flat, g flat and c flat. so 6 flats equals key of g flat. [...] so if you have a flat and there is a natural sign, play the a. so go through the scale and you've got g flat, a flat, d flat, c, flat, d flat, e flat and f, so f is your only 9 flat note in the scale. (No mention of the piano) Reference Summary: learn how to read and write music intervals for improving your playing and improvisational skills on the piano in this free video clip series. Summary from Transcript (BART): learn tips on how to read and write intervals on sheet music in this free video clip on music theory and music lessons. Summary from Transcript+Video (VG-BART): learn how to sight read in the key of g flat for improving your playing and improvisational skills on the piano in this free video clip series.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ac32aa20-bf20-4959-8256-8fccd77661a2Cited by top-tier papers12
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang et al.ACL 2023 · 17 citations
- TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment AnalysisXianbing Zhao, Yixin Chen, Sicen Liu, Xuan Zang et al.WWW 2023 · 17 citations
- CFSum Coarse-to-Fine Contribution Network for Multimodal SummarizationMin Xiao, Junnan Zhu, Haitao Lin, Yu Zhou et al.ACL 2023 · 15 citations
- Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosNayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu et al.EMNLP 2022 · 10 citations
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar et al.CVPR 2024 · 7 citations
Builds on9
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 2,453 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy et al.ICCV 2019 · 1,396 citations
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu et al.AAAI 2020 · 1,047 citations
Related papers
- V2P: Vision-to-Prompt based Multi-Modal Product Summary GenerationXuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao et al.SIGIR 2022 · 25 citations
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination MitigationZheng Qi, Chao Shang, Evangelia Spiliopoulou, Nikolaos PappasICML 2026
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsYingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li et al.ICCV 2025 · 2 citations
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang et al.AAAI 2022 · 61 citations
- Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language ModelsZhihang Liu, Chen-Wei Xie, Pandeng Li, Liming Zhao et al.CVPR 2025
