Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons
Frederick Riemenschneider, Anette Frank
Abstract
Multilingual language models (MLLMs) have demonstrated remarkable abilities to transfer knowledge across languages, despite being trained without explicit cross-lingual supervision. We analyze the parameter spaces of three MLLMs to study how their representations evolve during pre-training, observing patterns consistent with compression: models initially form language-specific representations, which gradually converge into cross-lingual abstractions as training progresses. Through probing experiments, we observe a clear transition from uniform language identification capabilities across layers to more specialized layer functions. For deeper analysis, we focus on neurons that encode distinct semantic concepts. By tracing their development during pre-training, we show how they gradually align across languages. Notably, we identify specific neurons that emerge as increasingly reliable predictors for the same concepts across languages. This alignment manifests concretely in generation: once an MLLM exhibits cross-lingual generalization according to our measures, we can select concept-specific neurons identified from, e.g., Spanish text and manipulate them to guide token predictions. Remarkably, rather than generating Spanish text, the model produces semantically coherent English text. This demonstrates that cross-lingually aligned neurons encode generalized semantic representations, independent of the original language encoding. 2 Two checkpoints each from BLOOM-560M and BLOOM-7B1 were excluded from our analysis, as they appear to be corrupted (cf. https://huggingface.co/ bigscience/bloom-560m-intermediate/discussions/2 and https://huggingface.co/bigscience/ bloom-7b1-intermediate/discussions ). (a) Early training stage (step 1000). (b) Late training stage (step 400 000).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on11
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
- Cross-Lingual Ability of Multilingual BERT: An Empirical StudyKarthikeyan K, Zihan Wang, Stephen Mayhew, Dan RothICLR 2020 · 378 citations
- From Zero to Hero: On the Limitations of Zero-Shot Language Transfer with Multilingual TransformersAnne Lauscher, Vinit Ravishankar, Ivan Vulic, Goran GlavasEMNLP 2020 · 235 citations
- Few-shot Learning with Multilingual Generative Language ModelsXi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang et al.EMNLP 2022 · 113 citations
- On the Cross-lingual Transferability of Monolingual RepresentationsMikel Artetxe, Sebastian Ruder, Dani YogatamaACL 2020 · 57 citations
Related papers
- The Emergence of Abstract Thought in Large Language Models Beyond Any LanguageYuxin Chen, Yiran Zhao, Yang Zhang, An Zhang et al.NeurIPS 2025 · 27 citations
- The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language ModelJiawei Chen, Wentao Chen, Jing Su, Jingjing Xu et al.ICLR 2025
- Analyzing the Mono- and Cross-Lingual Pretraining Dynamics of Multilingual Language ModelsTerra Blevins, Hila Gonen, Luke ZettlemoyerEMNLP 2022 · 13 citations
- Discovering Language-neutral Sub-networks in Multilingual Language ModelsNegar Foroutan, Mohammadreza Banaei, Rémi Lebret, Antoine Bosselut et al.EMNLP 2022 · 9 citations
- Separating Tongue from Thought: Activation Patching Reveals Language-Agnostic Concept Representations in TransformersClément Dumas, Chris Wendler, Veniamin Veselovsky, Giovanni Monea et al.ACL 2025
