Unifying Cross-lingual Summarization and Machine Translation with Compression Rate
Yu Bai, Heyan Huang, Kai Fan, Yang Gao, Yiming Zhu, Jiaao Zhan, Zewen Chi, Boxing Chen
摘要
Cross-Lingual Summarization (CLS) is a task that extracts important information from a source document and summarizes it into a summary in another language. It is a challenging task that requires a system to understand, summarize, and translate at the same time, making it highly related to Monolingual Summarization (MS) and Machine Translation (MT). In practice, the training resources for Machine Translation are far more than that for crosslingual and monolingual summarization. Thus incorporating the Machine Translation corpus into CLS would be beneficial for its performance. However, the present work only leverages a simple multi-task framework to bring Machine Translation in, lacking deeper exploration.
In this paper, we propose a novel task, Cross-lingual Summarization with Compression rate (CSC), to benefit Cross-Lingual Summarization by large-scale Machine Translation corpus. Through introducing compression rate, the information ratio between the source and the target text, we regard the MT task as a special CLS task with a compression rate of 100%. Hence they can be trained as a unified task, sharing knowledge more effectively. However, a huge gap exists between the MT task and the CLS task, where samples with compression rates between 30% and 90% are extremely rare. Hence, to bridge these two tasks smoothly, we propose an effective data augmentation method to produce document-summary pairs with different compression rates. The proposed method not only improves the performance of the CLS task, but also provides controllability to generate summaries in desired lengths. Experiments demonstrate that our method outperforms various strong baselines in three cross-lingual summarization datasets. We released our code and data at https://github.com/ybai-nlp/CLS_CR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- SoLoPO: Unlocking Long-Context Capabilities in LLMs via Short-to-Long Preference OptimizationHuashan Sun, Shengyi Liao, Yansen Han, Yu Bai 等ICLR 2026 · 被引用 9 次
- An Empirical Study of Many-to-Many Summarization with Large Language ModelsJiaan Wang, Fandong Meng, Zengkui Sun, Yunlong Liang 等ACL 2025
它引用的顶会 Paper4
- Jointly Learning to Align and Summarize for Neural Cross-Lingual SummarizationYue Cao, Hui Liu, Xiaojun WanACL 2020 · 被引用 52 次
- Attend, Translate and Summarize: An Efficient Method for Neural Cross-Lingual SummarizationJunnan Zhu, Yu Zhou, Jiajun Zhang, Chengqing ZongACL 2020 · 被引用 51 次
- Models and Datasets for Cross-Lingual SummarisationLaura Perez-Beltrachini, Mirella LapataEMNLP 2021 · 被引用 1 次
- Cross-Lingual Abstractive Summarization with Limited Parallel ResourcesYu Bai, Yang Gao, Heyan HuangACL 2021
相关 Paper
- A Variational Hierarchical Model for Neural Cross-Lingual SummarizationYunlong Liang, Fandong Meng, Chulun Zhou, Jinan Xu 等ACL 2022 · 被引用 36 次
- MultiSumm: Towards a Unified Model for Multi-Lingual Abstractive SummarizationYue Cao, Xiaojun Wan, Jin-ge Yao, Dian YuAAAI 2020 · 被引用 28 次
- CAR-Transformer: Cross-Attention Reinforcement Transformer for Cross-Lingual SummarizationYuang Cai, Yuyu YuanAAAI 2024 · 被引用 7 次
- Towards Unifying Multi-Lingual and Cross-Lingual SummarizationJiaan Wang, Fandong Meng, Duo Zheng, Yunlong Liang 等ACL 2023 · 被引用 24 次
- CrossSum: Beyond English-Centric Cross-Lingual Summarization for 1, 500+ Language PairsAbhik Bhattacharjee, Tahmid Hasan, Wasi Uddin Ahmad, Yuan-Fang Li 等ACL 2023 · 被引用 23 次
