Domain Adaptation in Multilingual and Multi-Domain Monolingual Settings for Complex Word Identification
George-Eduard Zaharia, Razvan-Alexandru Smadu, Dumitru-Clementin Cercel, Mihai Dascalu
Abstract
Complex word identification (CWI) is a cornerstone process towards proper text simplification. CWI is highly dependent on context, whereas its difficulty is augmented by the scarcity of available datasets which vary greatly in terms of domains and languages. As such, it becomes increasingly more difficult to develop a robust model that generalizes across a wide array of input examples. In this paper, we propose a novel training technique for the CWI task based on domain adaptation to improve the target character and context representations. This technique addresses the problem of working with multiple domains, inasmuch as it creates a way of smoothing the differences between the explored datasets. Moreover, we also propose a similar auxiliary task, namely text simplification, that can be used to complement lexical complexity prediction. Our model obtains a boost of up to 2.42% in terms of Pearson Correlation Coefficients in contrast to vanilla training techniques, when considering the CompLex from the Lexical Complexity Prediction 2021 dataset. At the same time, we obtain an increase of 3% in Pearson scores, while considering a cross-lingual setup relying on the Complex Word Identification 2018 dataset. In addition, our model yields state-of-the-art results in terms of Mean Absolute Error.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3575a5fd-524c-44b4-8891-061efe88fb7fCited by top-tier papers1
Ask how each one uses itBuilds on2
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Adversarial and Domain-Aware BERT for Cross-Domain Sentiment AnalysisChunning Du, Haifeng Sun, Jingyu Wang, Qi Qi et al.ACL 2020 · 165 citations
Related papers
- Label Confidence Weighted Learning for Target-level Sentence SimplificationXin Ying Qiu, Jingshen ZhangEMNLP 2024 · 1 citation
- SWiPE: A Dataset for Document-Level Simplification of Wikipedia PagesPhilippe Laban, Jesse Vig, Wojciech Kryscinski, Shafiq Joty et al.ACL 2023 · 5 citations
- Lexical Simplification with Pretrained EncodersJipeng Qiang, Yun Li, Yi Zhu, Yunhao Yuan et al.AAAI 2020 · 86 citations
- Explainable Prediction of Text Complexity: The Missing Preliminaries for Text SimplificationCristina Garbacea, Mengtian Guo, Samuel Carton, Qiaozhu MeiACL 2021
- DEplain: A German Parallel Corpus with Intralingual Translations into Plain Language for Sentence and Document SimplificationRegina Stodden, Omar Momen, Laura KallmeyerACL 2023 · 3 citations
