By Tying Embeddings You Are Assuming the Distributional Hypothesis
Francesco Bertolotti, Walter Cazzola
Abstract
In this work, we analyze both theoretically and empirically the effect of tied input-output embeddings-a popular technique that reduces the model size while often improving training. Interestingly, we found that this technique is connected to Harris (1954)'s distributional hypothesis-often portrayed by the famous Firth (1957)'s quote "a word is characterized by the company it keeps". Specifically, our findings indicate that words (or, more broadly, symbols) with similar semantics tend to be encoded in similar input embeddings, while words that appear in similar contexts are encoded in similar output embeddings (thus explaining the semantic space arising in input and output embedding of foundational language models). As a consequence of these findings, the tying of the input and output embeddings is encouraged only when the distributional hypothesis holds for the underlying data. These results also provide insight into the embeddings of foundation language models (which are known to be semantically organized). Further, we complement the theoretical findings with several experiments supporting the claims.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 78b1bc61-ec39-4d97-9876-32af68c1f0f3Builds on15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer et al.ACL 2020 · 210 citations
Related papers
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 10 citations
- The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model PretrainingTing-Rui Chiang, Dani YogatamaEMNLP 2023 · 1 citation
- Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional SemanticsChun Hei Lo, Wai Lam, Hong Cheng, Guy EmersonACL 2024
- Grounded Compositional Outputs for Adaptive Language ModelingNikolaos Pappas, Phoebe Mulcaire, Noah A. SmithEMNLP 2020 · 1 citation
- Token Embeddings Violate the Manifold HypothesisMichael Robinson, Sourya Dey, Tony ChiangNeurIPS 2025 · 19 citations
