A Theory of Unsupervised Translation Motivated by Understanding Animal Communication
Shafi Goldwasser, David F. Gruber, Adam Tauman Kalai, Orr Paradise
Abstract
Neural networks are capable of translating between languages-in some cases even between two languages where there is little or no access to parallel translations, in what is known as Unsupervised Machine Translation (UMT). Given this progress, it is intriguing to ask whether machine learning tools can ultimately enable understanding animal communication, particularly that of highly intelligent animals. We propose a theoretical framework for analyzing UMT when no parallel translations are available and when it cannot be assumed that the source and target corpora address related subject domains or posses similar linguistic structure. We exemplify this theory with two stylized models of language, for which our framework provides bounds on necessary sample complexity; the bounds are formally proven and experimentally verified on synthetic data. These bounds show that the error rates are inversely related to the language complexity and amount of common ground. This suggests that unsupervised translation of animal communication may be feasible if the communication system is sufficiently complex. Recent interest in translating animal communication [2, 3, 9] has been motivated by breakthrough performance of Language Models (LMs). Empirical work has succeeded in unsupervised translation between human-language pairs such as English-French [23, 5] and programming languages such as Python-Java [33] . Key to this feasibility seems to be the fact that language statistics, captured by a LM (a probability distribution over text), encapsulate more than just grammar. For example, even though both are grammatically correct, The calf nursed from its mother is more than 1,000 times more likely than The calf nursed from its father . 2 Given this remarkable progress, it is natural to ask whether it is possible to collect and analyze animal communication data, aiming towards translating animal communication to a human language description. This is particularly interesting when the source language may be of highly social and intelligent animals, such as whales, and the target language is a human language, such as English. Challenges. The first and most basic challenge is understanding the goal, a question with a rich history of philosophical debate [38] . To define the goal, we consider a hypothetical ground-truth translator. As a thought experiment, consider a "mermaid" fluent in English and the source language * Authors listed alphabetically. 2 Probabilities computed using the GPT-3 API https://openai.com/api/ text-davinci-02 model. 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 003a2157-0d9d-4ea7-9840-211d975492eaCited by top-tier papers3
- WhAM: Towards A Translative Model of Sperm Whale VocalizationOrr Paradise, Liangyuan Chen, Pranav Muralikrishnan, Hugo Flores García et al.NeurIPS 2025 · 5 citations
- Unsupervised Translation of Emergent CommunicationIdo Levy, Orr Paradise, Boaz Carmeli, Ron Meir et al.AAAI 2025 · 3 citations
- LLMs Can Hide Text in Other Text of the Same LengthAntonio Norelli, Michael M. BronsteinICLR 2026 · 3 citations
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated PromptsTaylor Shin, Yasaman Razeghi, Robert L. Logan IV, Eric Wallace et al.EMNLP 2020 · 1,162 citations
- Leveraging Automated Unit Tests for Unsupervised Code TranslationBaptiste Rozière, Jie Zhang, François Charton, Mark Harman et al.ICLR 2022 · 161 citations
- Language Model Prior for Low-Resource Neural Machine TranslationChristos Baziotis, Barry Haddow, Alexandra BirchEMNLP 2020 · 11 citations
- Bridging the Data Gap between Training and Inference for Unsupervised Neural Machine TranslationZhiwei He, Xing Wang, Rui Wang, Shuming Shi et al.ACL 2022
Related papers
- On Learning Language-Invariant Representations for Universal Machine TranslationHan Zhao, Junjie Hu, Andrej RisteskiICML 2020 · 8 citations
- Fine-Tuning Large Language Models to Translate: Will a Touch of Noisy Data in Misaligned Languages Suffice?Dawei Zhu, Pinzhen Chen, Miaoran Zhang, Barry Haddow et al.EMNLP 2024 · 3 citations
- What Languages are Easy to Language-Model? A Perspective from Learning Probabilistic Regular LanguagesNadav Borenstein, Anej Svete, Robin Chan, Josef Valvoda et al.ACL 2024
- Can Large Language Models Translate Spoken-Only Languages through International Phonetic Transcription?Jiale Chen, Xuelian Dong, Qihao Yang, Wenxiu Xie et al.EMNLP 2025
- Explicit Learning and the LLM in Machine TranslationMalik Marmonier, Rachel Bawden, Benoît SagotEMNLP 2025
