Pretraining Language Models Using Translationese
Meet Doshi, Raj Dabre, Pushpak Bhattacharyya
摘要
In this paper, we explore the utility of Translationese as synthetic data created using machine translation for pre-training language models (LMs) for low-resource languages (LRLs). Our simple methodology consists of translating large amounts of web-crawled monolingual documents (clean) into the LRLs, followed by filtering the translated documents using tiny LMs trained on small but clean LRL data. Taking the case of Indian languages, we pre-train LMs from scratch with 28M and 85M parameters, and then fine-tune them for 5 downstream natural language understanding (NLU) and 4 generative (NLG) tasks. We observe that pretraining on filtered synthetic data leads to relative performance drops of only 0.87% for NLU and 2.35% for NLG, compared to pre-training on clean data, and this gap further diminishes upon the inclusion of a small amount of clean data. We also study the impact of synthetic data filtering and the choice of source language for synthetic data generation. Furthermore, evaluating continually pre-trained larger models like Gemma-2B and Llama-3-8B in few-shot settings, we observe that using synthetic data is competitive with using clean data. Our findings suggest that synthetic data shows promise for bridging the pre-training gap between English and LRLs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Lost in Literalism: How Supervised Training Shapes Translationese in LLMsYafu Li, Ronghao Zhang, Zhilin Wang, Huajian Zhang 等ACL 2025 · 被引用 12 次
- Modeling Annotator Disagreement with Demographic-Aware Experts and Synthetic PerspectivesYinuo Xu, Veronica Derricks, Allison Earl, David JurgensACL 2026 · 被引用 8 次
- UPDESH: Synthesizing Grounded Instruction Tuning Data for 13 Indic LanguagesPranjal A. Chitale, Varun Gumma, Sanchit Ahuja, Prashant Kodali 等ACL 2026 · 被引用 1 次
- Multilingual Language Model Pretraining using Machine-translated DataJiayi Wang, Yao Lu, Maurice Weber, Max Ryabinin 等EMNLP 2025
- Lost in Translation, and Found: Detecting and Interpreting Translation EffectsShira Wein, Anna Serbina, Jiyuan Ji, Nathan Wolf 等ACL 2026
它引用的顶会 Paper8
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- BARTScore: Evaluating Generated Text as Text GenerationWeizhe Yuan, Graham Neubig, Pengfei LiuNeurIPS 2021 · 被引用 1,143 次
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang 等EMNLP 2022 · 被引用 113 次
- On Faithfulness and Factuality in Abstractive SummarizationJoshua Maynez, Shashi Narayan, Bernd Bohnet, Ryan T. McDonaldACL 2020 · 被引用 54 次
相关 Paper
- BhashaKritika: Building Synthetic Pretraining Data at Scale for Indic LanguagesGuduru Manoj, Neel Prabhanjan Rachamalla, Ashish Kulkarni, Gautam Rajeev 等AAAI 2026 · 被引用 2 次
- Scaling Low-Resource MT via Synthetic Data Generation with LLMsOna de Gibert, Joseph Attieh, Teemu Vahtola, Mikko Aulamo 等EMNLP 2025 · 被引用 2 次
- Paramanu: Compact and Competitive Monolingual Language Models for Low-Resource Morphologically Rich Indian LanguagesMitodru Niyogi, Éric Gaussier, Arnab BhattacharyaACL 2026 · 被引用 4 次
- PromDA: Prompt-based Data Augmentation for Low-Resource NLU TasksYufei Wang, Can Xu, Qingfeng Sun, Huang Hu 等ACL 2022
- Small Data, Big Impact: Leveraging Minimal Data for Effective Machine TranslationJean Maillard, Cynthia Gao, Elahe Kalbassi, Kaushik Ram Sadagopan 等ACL 2023 · 被引用 3 次
