Translating away Translationese without Parallel Data
Rricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van Genabith
Abstract
Translated texts exhibit systematic linguistic differences compared to original texts in the same language, and these differences are referred to as translationese. Translationese has effects on various cross-lingual natural language processing tasks, potentially leading to biased results. In this paper, we explore a novel approach to reduce translationese in translated texts: translation-based style transfer. As there are no parallel human-translated and original data in the same language, we use a self-supervised approach that can learn from comparable (rather than parallel) mono-lingual original and translated data. However, even this self-supervised approach requires some parallel data for validation. We show how we can eliminate the need for parallel validation data by combining the self-supervised loss with an unsupervised loss. This unsupervised loss leverages the original language model loss over the style-transferred output and a semantic similarity loss between the input and style-transferred output. We evaluate our approach in terms of original vs. translationese binary classification in addition to measuring content preservation and target-style fluency. The results show that our approach is able to reduce translationese classifier accuracy to a level of a random classifier after style transfer while adequately preserving the content and fluency in the target original style.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b418a88d-3635-4a1e-a500-7a4c74facd2aCited by top-tier papers4
- Lost in Literalism: How Supervised Training Shapes Translationese in LLMsYafu Li, Ronghao Zhang, Zhilin Wang, Huajian Zhang et al.ACL 2025 · 12 citations
- Multi-perspective Alignment for Increasing Naturalness in Neural Machine TranslationHuiyuan Lai, Esther Ploeger, Rik van Noord, Antonio ToralACL 2025
- Lost in Translation, and Found: Detecting and Interpreting Translation EffectsShira Wein, Anna Serbina, Jiyuan Ji, Nathan Wolf et al.ACL 2026
- Robust Estimation of Population-Level Effects in Repeated-Measures NLP Experimental DesignsAlejandro Benito-Santos, Adrián Ghajari, Víctor FresnoACL 2025
Builds on6
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- Translation Artifacts in Cross-lingual Transfer LearningMikel Artetxe, Gorka Labaka, Eneko AgirreEMNLP 2020 · 68 citations
- Non-Parallel Text Style Transfer with Self-Parallel SupervisionRuibo Liu, Chongyang Gao, Chenyan Jia, Guangxuan Xu et al.ICLR 2022 · 19 citations
- A Call for More Rigor in Unsupervised Cross-lingual LearningMikel Artetxe, Sebastian Ruder, Dani Yogatama, Gorka Labaka et al.ACL 2020 · 5 citations
Related papers
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
- Exploring Contextual Word-level Style Relevance for Unsupervised Style TransferChulun Zhou, Liangyu Chen, Jiachen Liu, Xinyan Xiao et al.ACL 2020 · 34 citations
- On The Evaluation of Machine Translation SystemsTrained With Back-TranslationSergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael AuliACL 2020 · 15 citations
- Collaborative Learning of Bidirectional Decoders for Unsupervised Text Style TransferYun Ma, Yangbin Chen, Xudong Mao, Qing LiEMNLP 2021 · 6 citations
- Rethinking Style Transformer with Energy-based Interpretation: Adversarial Unsupervised Style Transfer using a Pretrained ModelHojun Cho, Dohee Kim, Seungwoo Ryu, ChaeHun Park et al.EMNLP 2022
