By Tying Embeddings You Are Assuming the Distributional Hypothesis
Francesco Bertolotti, Walter Cazzola
摘要
In this work, we analyze both theoretically and empirically the effect of tied input-output embeddings-a popular technique that reduces the model size while often improving training. Interestingly, we found that this technique is connected to Harris (1954)'s distributional hypothesis-often portrayed by the famous Firth (1957)'s quote "a word is characterized by the company it keeps". Specifically, our findings indicate that words (or, more broadly, symbols) with similar semantics tend to be encoded in similar input embeddings, while words that appear in similar contexts are encoded in similar output embeddings (thus explaining the semantic space arising in input and output embedding of foundational language models). As a consequence of these findings, the tying of the input and output embeddings is encouraged only when the distributional hypothesis holds for the underlying data. These results also provide insight into the embeddings of foundation language models (which are known to be semantically organized). Further, we complement the theoretical findings with several experiments supporting the claims.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 被引用 624 次
- Emerging Cross-lingual Structure in Pretrained Language ModelsAlexis Conneau, Shijie Wu, Haoran Li, Luke Zettlemoyer 等ACL 2020 · 被引用 210 次
相关 Paper
- On the Emergence of Linear Analogies in Word EmbeddingsDaniel J. Korchinski, Dhruva Karkada, Yasaman Bahri, Matthieu WyartNeurIPS 2025 · 被引用 10 次
- The Distributional Hypothesis Does Not Fully Explain the Benefits of Masked Language Model PretrainingTing-Rui Chiang, Dani YogatamaEMNLP 2023 · 被引用 1 次
- Distributional Inclusion Hypothesis and Quantifications: Probing for Hypernymy in Functional Distributional SemanticsChun Hei Lo, Wai Lam, Hong Cheng, Guy EmersonACL 2024
- Grounded Compositional Outputs for Adaptive Language ModelingNikolaos Pappas, Phoebe Mulcaire, Noah A. SmithEMNLP 2020 · 被引用 1 次
- Token Embeddings Violate the Manifold HypothesisMichael Robinson, Sourya Dey, Tony ChiangNeurIPS 2025 · 被引用 19 次
