Scaling Sparse Feature Circuits For Studying In-Context Learning
Dmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel Nanda
Abstract
Sparse autoencoders (SAEs) are a popular tool for interpreting large language model activations, but their utility in addressing open questions in interpretability remains unclear. In this work, we demonstrate their effectiveness by using SAEs to deepen our understanding of the mechanism behind in-context learning (ICL). We identify abstract SAE features that (i) encode the model's knowledge of which task to execute and (ii) whose latent vectors causally induce the task zero-shot. This aligns with prior work showing that ICL is mediated by task vectors. We further demonstrate that these task vectors are well approximated by a sparse sum of SAE latents, including these task-execution features. To explore the ICL mechanism, we adapt the sparse feature circuits methodology of Marks et al. (2024) to work for the much larger Gemma-1 2B model, with 30 times as many parameters, and to the more complex task of ICL. Through circuit finding, we discover task-detecting features with corresponding SAE latents that activate earlier in the prompt, that detect when tasks have been performed. They are causally linked with task-execution features through the attention and MLP sublayers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cb99afff-1ef1-413c-a93a-c2ece7234558Cited by top-tier papers3
- Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language ModelsZihao Li, Xu Wang, Yuzhe Yang, Ziyu Yao et al.EMNLP 2025 · 14 citations
- How Few-Shot Examples Add Up: A Causal Decomposition of Function Vectors in In-Context LearningEntang Wang, Yiwei Wang, Aleksandra Bakalova, Michael HahnICML 2026 · 1 citation
- Shared Lexical Task Representations Explain Behavioral Variability In LLMsZhuonan Yang, Jacob Xiaochen Li, Francisco Velez, Eric Todd et al.ICML 2026
Builds on24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart et al.ICLR 2024 · 1,072 citations
- An Explanation of In-context Learning as Implicit Bayesian InferenceSang Michael Xie, Aditi Raghunathan, Percy Liang, Tengyu MaICLR 2022 · 1,030 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Transformers Learn In-Context by Gradient DescentJohannes von Oswald, Eyvind Niklasson, Ettore Randazzo, João Sacramento et al.ICML 2023 · 729 citations
Related papers
- Sparse Autoencoder Features for Classifications and TransferabilityJack Gallifant, Shan Chen, Kuleen Sasse, Hugo J. W. L. Aerts et al.EMNLP 2025
- Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language ModelsIkhyun Cho, Julia HockenmaierEMNLP 2025 · 3 citations
- Uncovering Sentiment Analysis Circuit in Large Language ModelShichen Li, Zhouyang Wang, Zhongqing Wang, Peifeng LiACL 2026
- Sparse Autoencoders Reveal Temporal Difference Learning in Large Language ModelsCan Demircan, Tankred Saanum, Akshay Kumar Jagadish, Marcel Binz et al.ICLR 2025 · 1 citation
- Are Sparse Autoencoders Useful? A Case Study in Sparse ProbingSubhash Kantamneni, Joshua Engels, Senthooran Rajamanoharan, Max Tegmark et al.ICML 2025
