Align and Attend: Multimodal Summarization with Dual Contrastive Losses
Bo He, Jun Wang, Jielin Qiu, Trung Bui, Abhinav Shrivastava, Zhaowen Wang
摘要
The goal of multimodal summarization is to extract the most important information from different modalities to form summaries. Unlike unimodal summarization, the multimodal summarization task explicitly leverages crossmodal information to help generate more reliable and highquality summaries. However, existing methods fail to leverage the temporal correspondence between different modalities and ignore the intrinsic correlation between different samples. To address this issue, we introduce Align and Attend Multimodal Summarization (A2Summ), a unified multimodal transformer-based model which can effectively align and attend the multimodal input. In addition, we propose two novel contrastive losses to model both inter-sample and intra-sample correlations. Extensive experiments on two standard video summarization datasets (TVSum and SumMe) and two multimodal summarization datasets (Daily Mail and CNN) demonstrate the superiority of A2Summ, achieving state-of-the-art performances on all datasets. Moreover, we collected a large-scale multimodal summarization dataset BLiSS, which contains livestream videos and transcribed texts with annotated summaries. Our code and dataset are publicly available at https://boheumd.github.io/A2Summ/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction TuningHang Hua, Yunlong Tang, Chenliang Xu, Jiebo LuoAAAI 2025 · 被引用 61 次
- Chop & Learn: Recognizing and Generating Object-State CompositionsNirat Saini, Hanyu Wang, Archana Swaminathan, Vinoj Jayasundara 等ICCV 2023 · 被引用 20 次
- CSTA: CNN-based Spatiotemporal Attention for Video SummarizationJaewon Son, Jaehun Park, Kwangsu KimCVPR 2024 · 被引用 17 次
- TutoAI: a cross-domain framework for AI-assisted mixed-media tutorial creation on physical tasksYuexi Chen, Vlad I. Morariu, Anh Truong, Zhicheng LiuCHI 2024 · 被引用 16 次
- MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of VideosJielin Qiu, Jiacheng Zhu, William Han, Aditesh Kumar 等CVPR 2024 · 被引用 7 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive SummarizationJingqing Zhang, Yao Zhao, Mohammad Saleh, Peter J. LiuICML 2020 · 被引用 2,453 次
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 被引用 2,340 次
- Unified Vision-Language Pre-Training for Image Captioning and VQALuowei Zhou, Hamid Palangi, Lei Zhang, Houdong Hu 等AAAI 2020 · 被引用 1,047 次
- Contrastive Learning with Hard Negative SamplesJoshua David Robinson, Ching-Yao Chuang, Suvrit Sra, Stefanie JegelkaICLR 2021 · 被引用 999 次
相关 Paper
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 被引用 1 次
- UniMS: A Unified Framework for Multimodal Summarization with Knowledge DistillationZhengkun Zhang, Xiaojun Meng, Yasheng Wang, Xin Jiang 等AAAI 2022 · 被引用 61 次
- Multimodal Video Summarization via Time-Aware TransformersXindi Shang, Zehuan Yuan, Anran Wang, Changhu WangACM MM 2021 · 被引用 31 次
- TopicCAT: Unsupervised Topic-Guided Co-Attention Transformer for Extreme Multimodal SummarisationPeggy Tang, Kun Hu, Lei Zhang, Junbin Gao 等ACM MM 2023 · 被引用 5 次
- VMSMO: Learning to Generate Multimodal Summary for Video-based News ArticlesMingzhe Li, Xiuying Chen, Shen Gao, Zhangming Chan 等EMNLP 2020 · 被引用 65 次
