Losing our Tail, Again: (Un)Natural Selection & Multilingual LLMs
Eva Vanmassenhove
Abstract
Multilingual Large Language Models considerably changed how technologies influence language. While previous technologies could mediate or assist humans, there is now a tendency to offload the task of writing itself to these technologies, enabling models to change our languages more directly. While they provide us quick access to information and impressively fluent output, beneath their (apparent) sophistication lies a subtle, insidious threat: the gradual decline and loss of linguistic diversity. In this position paper, I explore how model collapse, with a particular focus on translation technology, can lead to the loss of linguistic forms, grammatical features, and cultural nuance. Model collapse refers to the consequences of self-consuming training loops, where automatically generated data (re-)enters the training data, leading to a gradual distortion of the data distribution and the underrepresentation of low-probability linguistic phenomena. Drawing on recent work in Computer Vision, Natural Language Processing and Machine Translation, I argue that the many tails of our linguistic distributions might be vanishing, and with them, the narratives and identities they carry. This paper is a call to resist linguistic flattening and to reimagine Natural Language Processing as a field that encourages, values and protects expressive multilingual diversity and creativity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54ba49e2-bb86-4c0b-a203-8cb69240ad79Builds on7
- Faith and Fate: Limits of Transformers on CompositionalityNouha Dziri, Ximing Lu, Melanie Sclar, Xiang Lorraine Li et al.NeurIPS 2023 · 728 citations
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina et al.ICLR 2024 · 332 citations
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun et al.ICLR 2024 · 279 citations
- The State and Fate of Linguistic Diversity and Inclusion in the NLP WorldPratik Joshi, Sebastin Santy, Amar Budhiraja, Kalika Bali et al.ACL 2020 · 40 citations
- Bias Amplification in Language Model Evolution: An Iterated Learning PerspectiveYi Ren, Shangmin Guo, Linlu Qiu, Bailin Wang et al.NeurIPS 2024 · 23 citations
Related papers
- Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey et al.EMNLP 2025
- Computation Mechanism Behind LLM Position GeneralizationChi Han, Heng JiACL 2025
- Grammar as Control: Modular Language Generation for the Long TailNdapa NakasholeACL 2026 · 1 citation
- Authorship Attribution in Multilingual Machine-Generated TextsLucio La Cava, Dominik Macko, Róbert Móro, Ivan Srba et al.ACL 2026 · 7 citations
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 4 citations
