Assist Non-native Viewers: Multimodal Cross-Lingual Summarization for How2 Videos
Nayu Liu, Kaiwen Wei, Xian Sun, Hongfeng Yu, Fanglong Yao, Li Jin, Zhi Guo, Guangluan Xu
摘要
Multimodal summarization for videos aims to generate summaries from multi-source information (videos, audio transcripts), which has achieved promising progress. However, existing works are restricted to monolingual video scenarios, ignoring the demands of non-native video viewers to understand the cross-language videos in practical applications. It stimulates us to propose a new task, named Multimodal Cross-Lingual Summarization for videos (MCLS), which aims to generate cross-lingual summaries from multimodal inputs of videos. First, to make it applicable to MCLS scenarios, we conduct a Video-guided Dual Fusion network (VDF) that integrates multimodal and cross-lingual information via diverse fusion strategies at both encoder and decoder. Moreover, to alleviate the problem of high annotation costs and limited resources in MCLS, we propose a triple-stage training framework to assist MCLS by transferring the knowledge from monolingual multimodal summarization data, which includes: 1) multimodal summarization on sufficient prevalent language videos with a VDF model; 2) knowledge distillation (KD) guided adjustment on bilingual transcripts; 3) multimodal summarization for cross-lingual videos with a KD induced VDF model. Experiment results on the reorganized How2 dataset show that the VDF model alone outperforms previous methods for multimodal summarization, and the performance further improves by a large margin via the proposed triple-stage training framework.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- De-Biased Court's View Generation with CausalityYiquan Wu, Kun Kuang, Yating Zhang, Xiaozhong Liu 等EMNLP 2020 · 被引用 71 次
- Vision Guided Generative Pre-trained Language Models for Multimodal Abstractive SummarizationTiezheng Yu, Wenliang Dai, Zihan Liu, Pascale FungEMNLP 2021 · 被引用 64 次
- Multistage Fusion with Forget Gate for Multimodal Summarization in Open-Domain VideosNayu Liu, Xian Sun, Hongfeng Yu, Wenkai Zhang 等EMNLP 2020 · 被引用 54 次
- Jointly Learning to Align and Summarize for Neural Cross-Lingual SummarizationYue Cao, Hui Liu, Xiaojun WanACL 2020 · 被引用 52 次
相关 Paper
- V2Xum-LLM: Cross-Modal Video Summarization with Temporal Prompt Instruction TuningHang Hua, Yunlong Tang, Chenliang Xu, Jiebo LuoAAAI 2025 · 被引用 61 次
- Cross-Lingual Abstractive Summarization with Limited Parallel ResourcesYu Bai, Yang Gao, Heyan HuangACL 2021
- A Knowledge Augmented and Multimodal-Based Framework for Video SummarizationJiehang Xie, Xuanbai Chen, Shao-Ping Lu, Yulu YangACM MM 2022 · 被引用 11 次
- Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-ExpertsJiajun Han, Xuran Yang, Hui ZhangACM MM 2025 · 被引用 1 次
- Align and Attend: Multimodal Summarization with Dual Contrastive LossesBo He, Jun Wang, Jielin Qiu, Trung Bui 等CVPR 2023
