Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive Summarization
Tiezheng Yu, Wenliang Dai, Zihan Liu, Pascale Fung
摘要
Multimodal abstractive summarization (MAS) models that summarize videos (vision modality) and their corresponding transcripts (text modality) are able to extract the essential information from massive multimodal data on the Internet. Recently, large-scale generative pretrained language models (GPLMs) have been shown to be effective in text generation tasks. However, existing MAS models cannot leverage GPLMs' powerful generation ability. To fill this research gap, we aim to study two research questions: 1) how to inject visual information into GPLMs without hurting their generation ability; and 2) where is the optimal place in GPLMs to inject the visual information? In this paper, we present a simple yet effective method to construct vision guided (VG) GPLMs for the MAS task using attention-based add-on layers to incorporate visual information while maintaining their original text generation ability. Results show that our best model significantly surpasses the prior state-of-the-art model by 5.7 ROUGE-1, 5.3 ROUGE-2, and 5.1 ROUGE-L scores on the How2 dataset (Sanabria et al., 2018) , and our visual guidance method contributes 83.6% of the overall improvement. Furthermore, we conduct thorough ablation studies to analyze the effectiveness of various modality fusion methods and fusion locations. * * The two authors contribute equally. The code is available at: https://github.com/ HLTCHKUST/VG-GPLMs Video Frames Transcript: so now we are going to go over some basics sheet music readings for the key of g flat major. so you noticed the key of g flat, when you are reading real books, there is going to be a treble cleft here. it is going to have 6 flats 1, b flat, e flat, a flat, d flat, g flat and c flat. so 6 flats equals key of g flat. [...] so if you have a flat and there is a natural sign, play the a. so go through the scale and you've got g flat, a flat, d flat, c, flat, d flat, e flat and f, so f is your only 9 flat note in the scale. (No mention of the piano) Reference Summary: learn how to read and write music intervals for improving your playing and improvisational skills on the piano in this free video clip series. Summary from Transcript (BART): learn tips on how to read and write intervals on sheet music in this free video clip on music theory and music lessons. Summary from Transcript+Video (VG-BART): learn how to sight read in the key of g flat for improving your playing and improvisational skills on the piano in this free video clip series.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Summary-Oriented Vision Modeling for Multimodal Abstractive SummarizationYunlong Liang, Fandong Meng, Jinan Xu, Jiaan Wang 等ACL 2023 · 被引用 17 次
- TMMDA: A New Token Mixup Multimodal Data Augmentation for Multimodal Sentiment AnalysisXianbing Zhao, Yixin Chen, Sicen Liu, Xuan Zang 等WWW 2023 · 被引用 17 次
- CFSum Coarse-to-Fine Contribution Network for Multimodal SummarizationMin Xiao, Junnan Zhu, Haitao Lin, Yu Zhou 等ACL 2023 · 被引用 15 次
- Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 VideosNayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu 等EMNLP 2022 · 被引用 10 次
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar 等CVPR 2024 · 被引用 7 次
它引用的顶会 Paper9
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 被引用 2,258 次
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li 等ICLR 2020 · 被引用 1,825 次
- VideoBERT: A Joint Model for Video and Language Representation LearningChen Sun, Austin Myers, Carl Vondrick, Kevin Murphy 等ICCV 2019 · 被引用 1,396 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
相关 Paper
- V2P: Vision-to-Prompt based Multi-Modal Product Summary GenerationXuemeng Song, Liqiang Jing, Dengtian Lin, Zhongzhou Zhao 等SIGIR 2022 · 被引用 25 次
- Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination MitigationZheng Qi, Chao Shang, Evangelia Spiliopoulou, Nikolaos PappasICML 2026
- Aligning Effective Tokens with Video Anomaly in Large Language ModelsYingxian Chen, Jiahui Liu, Ruidi Fan, Yanwei Li 等ICCV 2025 · 被引用 2 次
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 等AAAI 2022 · 被引用 61 次
- Hybrid-Level Instruction Injection for Video Token Compression in Multi-modal Large Language ModelsZhihang Liu, Chen-Wei Xie, Pandeng Li, Liming Zhao 等CVPR 2025
