Cross-language Sentence Selection via Data Augmentation and Rationale Training
Yanda Chen, Chris Kedzie, Suraj Nair, Petra Galuscáková, Rui Zhang, Douglas W. Oard, Kathleen R. McKeown
摘要
This paper proposes an approach to crosslanguage sentence selection in a low-resource setting. It uses data augmentation and negative sampling techniques on noisy parallel sentence data to directly learn a cross-lingual embedding-based query relevance model. Results show that this approach performs as well as or better than multiple state-of-theart machine translation + monolingual retrieval systems trained on the same parallel data. Moreover, when a rationale training secondary objective is applied to encourage the model to match word alignment hints from a phrase-based statistical machine translation model, consistent improvements are seen across three language pairs (English-Somali, English-Swahili and English-Tagalog) over a variety of state-of-the-art baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan 等ACL 2023 · 被引用 3 次
- ERNIE-M: Enhanced Multilingual Representation by Aligning Cross-lingual Semantics with Monolingual CorporaXuan Ouyang, Shuohuan Wang, Chao Pang, Yu Sun 等EMNLP 2021 · 被引用 68 次
- Dual-Alignment Pre-training for Cross-lingual Sentence EmbeddingZiheng Li, Shaohan Huang, Zihan Zhang, Zhi-Hong Deng 等ACL 2023 · 被引用 5 次
- Neural Machine Translation with Monolingual Translation MemoryDeng Cai, Yan Wang, Huayang Li, Wai Lam 等ACL 2021
- CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-trainingSeungyoon Lee, Minhyuk Kim, Seongtae Hong, Youngjoon Jang 等ACL 2026
