Signal in the Noise: Polysemantic Interference Transfers and Predicts Cross-Model Influence
Bofan Gong, Shiyang Lai, James Evans, Dawn Song
Abstract
Polysemanticity is pervasive in language models and remains a major challenge for interpretation and model behavioral control. Leveraging sparse autoencoders (SAEs), we map the polysemantic topology of two small models (Pythia-70M and GPT-2-Small) to identify SAE feature pairs that are semantically unrelated yet exhibit interference within models. We intervene at four foci (prompt, token, feature, neuron) and measure induced shifts in the next-token prediction distribution, uncovering polysemantic structures that expose a systematic vulnerability in these models. Critically, interventions distilled from counterintuitive interference patterns shared by two small models transfer reliably to larger instruction-tuned models (Llama-3.1-8B/70B-Instruct and Gemma-2-9B-Instruct), yielding predictable behavioral shifts without access to model internals. These findings challenge the view that polysemanticity is purely stochastic, demonstrating instead that interference structures generalize across scale and family. Such generalization suggests a convergent, higher-order organization of internal representations, which is only weakly aligned with intuition and structured by latent regularities, offering new possibilities for both black-box control and theoretical insight into human and artificial cognition. Code and data are available here. * Equal contribution, alphabetical ordered. 1 Nevertheless, several studies have also documented limitations of SAEs (see Appendix P).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2a9eb578-fd52-4b3a-8ac1-93e40dbf6565Cited by top-tier papers1
Ask how each one uses itBuilds on11
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- Backdoor Defense via Decoupling the Training ProcessKunzhe Huang, Yiming Li, Baoyuan Wu, Zhan Qin et al.ICLR 2022 · 253 citations
- Transcoders find interpretable LLM feature circuitsJacob Dunefsky, Philippe Chlenski, Neel NandaNeurIPS 2024 · 222 citations
- A is for Absorption: Studying Feature Splitting and Absorption in Sparse AutoencodersDavid Chanin, James Wilken-Smith, Tomás Dulka, Hardik Bhatnagar et al.NeurIPS 2025 · 168 citations
- Don't trust your eyes: on the (un)reliability of feature visualizationsRobert Geirhos, Roland S. Zimmermann, Blair L. Bilodeau, Wieland Brendel et al.ICML 2024 · 38 citations
Related papers
- When Truthful Representations Flip Under Deceptive Instructions?Xianxuan Long, Yao Fu, Runchao Li, Mu Sheng et al.EMNLP 2025
- Endogenous Resistance to Activation Steering in Language ModelsAlex McKenzie, Keenan Pepper, Stijn Servaes, Martin Leitgab et al.ICML 2026 · 3 citations
- Breaking Bad Tokens: Detoxification of LLMs Using Sparse AutoencodersAgam Goyal, Vedant Rathi, William Yeh, Yian Wang et al.EMNLP 2025 · 1 citation
- Sparse Autoencoder Features for Classifications and TransferabilityJack Gallifant, Shan Chen, Kuleen Sasse, Hugo J. W. L. Aerts et al.EMNLP 2025
- Sparse Autoencoders Trained on the Same Data Learn Different FeaturesGonçalo Paulo, Nora BelroseICLR 2026 · 96 citations
