Detecting Fine-Grained Cross-Lingual Semantic Divergences without Supervision by Learning to Rank
Eleftheria Briakou, Marine Carpuat
Abstract
Detecting fine-grained differences in content conveyed in different languages matters for cross-lingual NLP and multilingual corpora analysis, but it is a challenging machine learning problem since annotation is expensive and hard to scale. This work improves the prediction and annotation of finegrained semantic divergences. We introduce a training strategy for multilingual BERT models by learning to rank synthetic divergent examples of varying granularity. We evaluate our models on the Rationalized English-French Semantic Divergences, a new dataset released with this work, consisting of English-French sentence-pairs annotated with semantic divergence classes and token-level rationales. Learning to rank helps detect finegrained sentence-level divergences more accurately than a strong sentence-level similarity model, while token-level predictions have the potential of further distinguishing between coarse and fine-grained divergences. ADV VERB ADJ NOUN how weak they are. BERT predictions permission, attention, hand, mercy, story WORDNET hypernyms communication, forgiveness, mercy absolutely fighting his policy mercy Seed Equivalent Sample Now, however, one of them is suddenly asking your help, and you can see from this how weak they are. Maintenant, cependant, l'un d'eux vient soudainement demander votre aide et vous pouvez voir à quel point ils sont faibles. Subtree Deletion Now, however, one of them is suddenly asking your help, and you can see from this. Maintenant, cependant, l'un d'eux vient soudainement demander votre aide et vous pouvez voir à quel point ils sont faibles . Phrase Replacement Now, however, one of them is absolutely fighting his policy , and you can see from this how weak they are. Maintenant, cependant, l'un d'eux vient soudainement demander votre aide et vous pouvez voir à quel point ils sont faibles. Lexical Substitution Now, however, one of them is suddenly asking your mercy , and you can see from this. Maintenant, cependant, l'un d'eux vient soudainement demander votre aide et vous pouvez voir à quel point ils sont faibles.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dd7e39a-8138-4cf7-930f-9dd5c582731aCited by top-tier papers8
- RankDNN: Learning to Rank for Few-Shot LearningQianyu Guo, Haotong Gong, Xujun Wei, Yanwei Fu et al.AAAI 2023 · 27 citations
- Bridging Background Knowledge Gaps in Translation with Automatic ExplicitationHyoJung Han, Jordan L. Boyd-Graber, Marine CarpuatEMNLP 2023 · 3 citations
- Characterizing and Measuring Linguistic Dataset DriftTyler A. Chang, Kishaloy Halder, Neha Anna John, Yogarshi Vyas et al.ACL 2023 · 2 citations
- Improving Translation Quality Estimation with Bias MitigationHui Huang, Shuangzhi Wu, Kehai Chen, Hui Di et al.ACL 2023 · 2 citations
- Not all Fake News is Written: A Dataset and Analysis of Misleading Video HeadlinesYoo Yeon Sung, Jordan L. Boyd-Graber, Naeemul HassanEMNLP 2023 · 1 citation
Builds on2
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- ERASER: A Benchmark to Evaluate Rationalized NLP ModelsJay DeYoung, Sarthak Jain, Nazneen Fatema Rajani, Eric P. Lehman et al.ACL 2020 · 36 citations
Related papers
- LAReQA: Language-Agnostic Answer Retrieval from a Multilingual PoolUma Roy, Noah Constant, Rami Al-Rfou, Aditya Barua et al.EMNLP 2020 · 39 citations
- Measure and Evaluation of Semantic Divergence across Two LanguagesSyrielle Montariol, Alexandre AllauzenACL 2021
- When is Wall a Pared and when a Muro?: Extracting Rules Governing Lexical SelectionAditi Chaudhary, Kayo Yin, Antonios Anastasopoulos, Graham NeubigEMNLP 2021 · 2 citations
- Beyond Noise: Mitigating the Impact of Fine-grained Semantic Divergences on Neural Machine TranslationEleftheria Briakou, Marine CarpuatACL 2021
- SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related DocumentsMichelle Wastl, Jannis Vamvas, Rico SennrichACL 2026
