Capturing Polysemanticity with PRISM: A Multi-Concept Feature Description Framework
Laura Kopf, Nils Feldhus, Kirill Bykov, Philine Lou Bommer, Anna Hedström, Marina M.-C. Höhne, Oliver Eberle
Abstract
Automated interpretability research aims to identify concepts encoded in neural network features to enhance human understanding of model behavior. Within the context of large language models (LLMs) for natural language processing (NLP), current automated neuron-level feature description methods face two key challenges: limited robustness and the assumption that each neuron encodes a single concept (monosemanticity), despite increasing evidence of polysemanticity. This assumption restricts the expressiveness of feature descriptions and limits their ability to capture the full range of behaviors encoded in model internals. To address this, we introduce Polysemantic FeatuRe Identification and Scoring Method (PRISM), a novel framework specifically designed to capture the complexity of features in LLMs. Unlike approaches that assign a single description per neuron, common in many automated interpretability methods in NLP, PRISM produces more nuanced descriptions that account for both monosemantic and polysemantic behavior. We apply PRISM to LLMs and, through extensive benchmarking against existing methods, demonstrate that our approach produces more accurate and faithful feature descriptions, improving both overall description quality (via a description score) and the ability to capture distinct concepts when polysemanticity is present (via a polysemanticity score).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f19838fc-017d-40fd-baf8-dfc5c166b079Cited by top-tier papers4
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh et al.NeurIPS 2025 · 23 citations
- Circuit Insights: Towards Interpretability Beyond ActivationsElena Golimblevskaia, Aakriti Jain, Bruno Puri, Ammar Ibrahim et al.ICLR 2026 · 4 citations
- Semantic Regexes: Auto-Interpreting LLM Features with a Structured LanguageAngie W. Boggust, Donghao Ren, Yannick Assogba, Dominik Moritz et al.ICLR 2026 · 3 citations
- AlgoTrace: Algorithmic Primitives and Compositional Geometry of Reasoning in Language ModelsSamuel Lippl, Thomas McGee, Kimberly Lopez, Ziwen Pan et al.ICML 2026
Builds on21
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 1,792 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- XAI for Transformers: Better Explanations through Conservative PropagationAmeen Ali, Thomas Schnake, Oliver Eberle, Grégoire Montavon et al.ICML 2022 · 144 citations
- Inducing Causal Structure for Interpretable Neural NetworksAtticus Geiger, Zhengxuan Wu, Hanson Lu, Josh Rozner et al.ICML 2022 · 104 citations
Related papers
- Revising and Falsifying Sparse Autoencoder Feature ExplanationsGeorge Ma, Samuel Pfrommer, Somayeh SojoudiNeurIPS 2025 · 7 citations
- Enhancing Automated Interpretability with Output-Centric Feature DescriptionsYoav Gur-Arieh, Roy Mayan, Chen Agassy, Atticus Geiger et al.ACL 2025
- Wasserstein Distances, Neuronal Entanglement, and SparsityShashata Sawmya, Linghao Kong, Ilia Markov, Dan Alistarh et al.ICLR 2025
- Beyond Interpretability: The Gains of Feature Monosemanticity on Model RobustnessQi Zhang, Yifei Wang, Jingyi Cui, Xiang Pan et al.ICLR 2025
- AND: Audio Network Dissection for Interpreting Deep Acoustic ModelsTung-Yu Wu, Yu-Xiang Lin, Tsui-Wei WengICML 2024 · 6 citations
