Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese Classification
Daria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith, Cristina España-Bonet
Abstract
Traditional hand-crafted linguisticallyinformed features have often been used for distinguishing between translated and original non-translated texts. By contrast, to date, neural architectures without manual feature engineering have been less explored for this task. In this work, we (i) compare the traditional feature-engineering-based approach to the feature-learning-based one and (ii) analyse the neural architectures in order to investigate how well the hand-crafted features explain the variance in the neural models' predictions. We use pre-trained neural word embeddings, as well as several end-to-end neural architectures in both monolingual and multilingual settings and compare them to feature-engineering-based SVM classifiers. We show that (i) neural architectures outperform other approaches by more than 20 accuracy points, with the BERT-based model performing the best in both the monolingual and multilingual settings; (ii) while many individual hand-crafted translationese features correlate with neural model predictions, feature importance analysis shows that the most important features for neural and classical architectures differ; and (iii) our multilingual experiments provide empirical evidence for translationese universals across languages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 688f3e85-835d-4cc7-8a4b-b8e08575ddcaCited by top-tier papers2
- Multi-perspective Alignment for Increasing Naturalness in Neural Machine TranslationHuiyuan Lai, Esther Ploeger, Rik van Noord, Antonio ToralACL 2025
- Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang et al.EMNLP 2025
Builds on3
- Statistical Power and Translationese in Machine Translation EvaluationYvette Graham, Barry Haddow, Philipp KoehnEMNLP 2020 · 82 citations
- On The Evaluation of Machine Translation SystemsTrained With Back-TranslationSergey Edunov, Myle Ott, Marc'Aurelio Ranzato, Michael AuliACL 2020 · 15 citations
- Translationese as a Language in "Multilingual" NMTParker Riley, Isaac Caswell, Markus Freitag, David GrangierACL 2020
Related papers
- Lost in Translation, and Found: Detecting and Interpreting Translation EffectsShira Wein, Anna Serbina, Jiyuan Ji, Nathan Wolf et al.ACL 2026
- Translating away Translationese without Parallel DataRricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van GenabithEMNLP 2023
- Revisiting Machine Translation for Cross-lingual ClassificationMikel Artetxe, Vedanuj Goswami, Shruti Bhosale, Angela Fan et al.EMNLP 2023 · 10 citations
- Language-agnostic BERT Sentence EmbeddingFangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan et al.ACL 2022
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
