Improving Neural Cross-Lingual Abstractive Summarization via Employing Optimal Transport Distance for Knowledge Distillation
Thong Thanh Nguyen, Anh Tuan Luu
Abstract
Current state-of-the-art cross-lingual summarization models employ multi-task learning paradigm, which works on a shared vocabulary module and relies on the self-attention mechanism to attend among tokens in two languages. However, correlation learned by self-attention is often loose and implicit, inefficient in capturing crucial cross-lingual representations between languages. The matter worsens when performing on languages with separate morphological or structural features, making the cross-lingual alignment more challenging, resulting in the performance drop. To overcome this problem, we propose a novel Knowledge-Distillation-based framework for Cross-Lingual Summarization, seeking to explicitly construct cross-lingual correlation by distilling the knowledge of the monolingual summarization teacher into the cross-lingual summarization student. Since the representations of the teacher and the student lie on two different vector spaces, we further propose a Knowledge Distillation loss using Sinkhorn Divergence, an Optimal-Transport distance, to estimate the discrepancy between those teacher and student representations. Due to the intuitively geometric nature of Sinkhorn Divergence, the student model can productively learn to align its produced cross-lingual hidden states with monolingual hidden states, hence leading to a strong correlation between distant languages. Experiments on cross-lingual summarization datasets in pairs of distant languages demonstrate that our method outperforms state-of-the-art models under both high and low-resourced settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1219b6b8-684f-4521-8acc-b6f4fcd3a1a7Cited by top-tier papers11
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsShuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao et al.EMNLP 2023 · 39 citations
- Multi-Level Optimal Transport for Universal Cross-Tokenizer Knowledge Distillation on Language ModelsXiao Cui, Mo Zhu, Yulei Qin, Liang Xie et al.AAAI 2025 · 31 citations
- Universal Vulnerabilities in Large Language Models: Backdoor Attacks for In-context LearningShuai Zhao, Meihuizi Jia, Anh Tuan Luu, Fengjun Pan et al.EMNLP 2024 · 28 citations
- Exploring the Potential of Large Language Models in Computational ArgumentationGuizhen Chen, Liying Cheng, Anh Tuan Luu, Lidong BingACL 2024 · 8 citations
- Adaptive Contrastive Learning on Multimodal Transformer for Review Helpfulness PredictionThong Nguyen, Xiaobao Wu, Anh Tuan Luu, Zhen Hai et al.EMNLP 2022 · 8 citations
Builds on11
- Contrastive Learning for Neural Topic ModelThong Nguyen, Anh Tuan LuuNeurIPS 2021 · 82 citations
- Capturing Greater Context for Question GenerationLuu Anh Tuan, Darsh J. Shah, Regina BarzilayAAAI 2020 · 77 citations
- Jointly Learning to Align and Summarize for Neural Cross-Lingual SummarizationYue Cao, Hui Liu, Xiaojun WanACL 2020 · 52 citations
- Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual SummarizationJunnan Zhu, Yu Zhou, Jiajun Zhang, Chengqing ZongACL 2020 · 51 citations
- Knowledge Distillation for Multilingual Unsupervised Neural Machine TranslationHaipeng Sun, Rui Wang, Kehai Chen, Masao Utiyama et al.ACL 2020 · 37 citations
Related papers
- MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language ModelsHoang Tran Vuong, Tue Le, Quyen Tran, Linh Ngo Van et al.AAAI 2026
- EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsMinh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep et al.EMNLP 2025
- Cross-Lingual Abstractive Summarization with Limited Parallel ResourcesYu Bai, Yang Gao, Heyan HuangACL 2021
- Implicit Word Reordering with Knowledge Distillation for Cross-Lingual Dependency ParsingZhuoran Li, Chunming Hu, Junfan Chen, Zhijun Chen et al.AAAI 2025 · 1 citation
- Dual-Space Knowledge Distillation for Large Language ModelsSongming Zhang, Xue Zhang, Zengkui Sun, Yufeng Chen et al.EMNLP 2024 · 3 citations
