Prediction Hubs are Context-Informed Frequent Tokens in LLMs
Beatrix Miranda Ginn Nielsen, Iuri Macocco, Marco Baroni
摘要
Hubness, the tendency for a few points to be among the nearest neighbours of a disproportionate number of other points, commonly arises when applying standard distance measures to high-dimensional data, often negatively impacting distance-based analysis. As autoregressive large language models (LLMs) operate on high-dimensional representations, we ask whether they are also affected by hubness. We first prove that the only large-scale representation comparison operation performed by LLMs, namely that between context and unembedding vectors to determine continuation probabilities, is not characterized by the concentration of distances phenomenon that typically causes the appearance of nuisance hubness. We then empirically show that this comparison still leads to a high degree of hubness, but the hubs in this case do not constitute a disturbance. They are rather the result of context-modulated frequent tokens often appearing in the pool of likely candidates for next token prediction. However, when other distances are used to compare LLM representations, we do not have the same theoretical guarantees, and, indeed, we see nuisance hubs appear. There are two main takeaways. First, hubness, while omnipresent in high-dimensional spaces, is not a negative property that needs to be mitigated when LLMs are being used for next token prediction. Second, when comparing representations from LLMs using Euclidean or cosine distance, there is a high risk of nuisance hubs and practitioners should use mitigation techniques if relevant.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Cross Modal Retrieval with Querybank NormalisationSimion-Vlad Bogolin, Ioana Croitoru, Hailin Jin, Yang Liu 等CVPR 2022 · 被引用 84 次
- Balance Act: Mitigating Hubness in Cross-Modal Retrieval with Query and Gallery BanksYimu Wang, Xiangru Jian, Bo XueEMNLP 2023 · 被引用 7 次
- Bridging Information-Theoretic and Geometric Compression in Language ModelsEmily Cheng, Corentin Kervadec, Marco BaroniEMNLP 2023 · 被引用 5 次
相关 Paper
- All Bark and No Bite: Rogue Dimensions in Transformer Language Models Obscure Representational QualityWilliam Timkey, Marten van SchijndelEMNLP 2021 · 被引用 59 次
- Adversarial Hubness in Multi-Modal RetrievalTingwei Zhang, Fnu Suya, Rishi D. Jha, Collin Zhang 等S&P 2026 · 被引用 1 次
- One Single Hub Text Breaks CLIP: Identifying Vulnerabilities in Cross-Modal Encoders via HubnessHiroyuki Deguchi, Katsuki Chousa, Yusuke SakaiACL 2026
- A Text is Worth Several Tokens: Text Embedding from LLMs Secretly Aligns Well with The Key TokensZhijie Nie, Richong Zhang, Zhanyu WuACL 2025 · 被引用 5 次
- A Multi-Perspective Analysis of Memorization in Large Language ModelsBowen Chen, Namgi Han, Yusuke MiyaoEMNLP 2024 · 被引用 2 次
