Wasserstein Distances, Neuronal Entanglement, and Sparsity
Shashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh, Nir Shavit
摘要
Disentangling polysemantic neurons is at the core of many current approaches to interpretability of large language models. Here we attempt to study how disentanglement can be used to understand performance, particularly under weight sparsity, a leading post-training optimization technique. We suggest a novel measure for estimating neuronal entanglement: the Wasserstein distance of a neuron's output distribution to a Gaussian. Moreover, we show the existence of a small number of highly entangled "Wasserstein Neurons" in each linear layer of an LLM, characterized by their highly non-Gaussian output distributions, their role in mapping similar inputs to dissimilar outputs, and their significant impact on model accuracy. To study these phenomena, we propose a new experimental framework for disentangling polysemantic neurons. Our framework separates each layer's inputs to create a mixture of experts where each neuron's output is computed by a mixture of neurons of lower Wasserstein distance, each better at maintaining accuracy when sparsified without retraining. We provide strong evidence that this is because the mixture of sparse experts is effectively disentangling the input-output relationship of individual neurons, in particular the difficult Wasserstein neurons.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Sparse CLIP: Co-Optimizing Interpretability and Performance in Contrastive LearningChuan Qin, Constantin Venhoff, Sonia Joseph, Fanyi Xiao 等ICLR 2026 · 被引用 4 次
- Negative Pre-activations Differentiate SyntaxLinghao Kong, Angelina Ning, Micah Adler, Nir N ShavitICLR 2026 · 被引用 2 次
- Resting Neurons, Active Insights: Robustifying Activation Sparsity in LLMs via SpontaneityHaotian Xu, Jiannan Yang, Tian Gao, Lily Weng 等ICML 2026
它引用的顶会 Paper14
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 被引用 1,240 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
相关 Paper
- The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert LevelJeremy Herbst, Stefan Wermter, Jae Hee LeeICML 2026 · 被引用 9 次
- Monet: Mixture of Monosemantic Experts for TransformersJungwoo Park, Ahn Young Jin, Kee-Eung Kim, Jaewoo KangICLR 2025
- Rethinking Evaluation of Sparse Autoencoders through the Representation of Polysemous WordsGouki Minegishi, Hiroki Furuta, Yusuke Iwasawa, Yutaka MatsuoICLR 2025
- Mixture of Experts Made Intrinsically InterpretableXingyi Yang, Constantin Venhoff, Ashkan Khakzar, Christian Schröder de Witt 等ICML 2025
- Understanding Cross-layer Contributions to Mixture-of-Experts Routing in LLMsWengang Li, Lingqi Zhang, Toshio Endo, Mohamed WahibICLR 2026
