Lune

ICML2024Top-tier venue

By Tying Embeddings You Are Assuming the Distributional Hypothesis

Francesco Bertolotti, Walter Cazzola

2024Year
4Citations

Abstract

In this work, we analyze both theoretically and empirically the effect of tied input-output embeddings-a popular technique that reduces the model size while often improving training. Interestingly, we found that this technique is connected to Harris (1954)'s distributional hypothesis-often portrayed by the famous Firth (1957)'s quote "a word is characterized by the company it keeps". Specifically, our findings indicate that words (or, more broadly, symbols) with similar semantics tend to be encoded in similar input embeddings, while words that appear in similar contexts are encoded in similar output embeddings (thus explaining the semantic space arising in input and output embedding of foundational language models). As a consequence of these findings, the tying of the input and output embeddings is encouraged only when the distributional hypothesis holds for the underlying data. These results also provide insight into the embeddings of foundation language models (which are known to be semantically organized). Further, we complement the theoretical findings with several experiments supporting the claims.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 78b1bc61-ec39-4d97-9876-32af68c1f0f3

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines