Toward Efficient Sparse Autoencoder-Guided Steering for Improved In-Context Learning in Large Language Models
Ikhyun Cho, Julia Hockenmaier
摘要
Sparse autoencoders (SAEs) have emerged as a powerful analytical tool in mechanistic interpretability for large language models (LLMs), with growing success in applications beyond interpretability. Building on this momentum, we present a novel approach that leverages SAEs to enhance the general in-context learning (ICL) performance of LLMs. Specifically, we introduce Feature Detection through Prompt Variation (FDPV), which leverages the SAE's remarkable ability to capture subtle differences between prompts, enabling efficient feature selection for downstream steering. In addition, we propose a novel steering method tailored to ICL-Selective In-Context Steering (SISTER)-grounded in recent insights from ICL research that LLMs utilize label words as key anchors. Our method yields a 3.5% average performance improvement across diverse text classification tasks and exhibits greater robustness to hyperparameter variations compared to standard steering approaches. Our code is available at https://github.com/ihcho2/SAE-ICL .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister 等NeurIPS 2023 · 被引用 1,549 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- A Survey on In-context LearningQingxiu Dong, Lei Li, Damai Dai, Ce Zheng 等EMNLP 2024 · 被引用 479 次
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 被引用 244 次
相关 Paper
- Scaling Sparse Feature Circuits For Studying In-Context LearningDmitrii Kharlapenko, Stepan Shabalin, Arthur Conmy, Neel NandaICML 2025
- Does Higher Interpretability Imply Better Utility? A Pairwise Analysis on Sparse AutoencodersXu Wang, Yan Hu, Benyou Wang, Difan ZouICLR 2026 · 被引用 9 次
- Route Sparse Autoencoder to Interpret Large Language ModelsWei Shi, Sihang Li, Tao Liang, Mingyang Wan 等EMNLP 2025 · 被引用 1 次
- DLM-Scope: Mechanistic Interpretability of Diffusion Language Models via Sparse AutoencodersXu Wang, Bingqing Jiang, Yu Wan, Baosong Yang 等ICML 2026
- Language Models Can Explain Visual Features via SteeringJavier Ferrando, Enrique Lopez-Cuena, Pablo Agustin Martin-Torres, Daniel Hinjos 等CVPR 2026 · 被引用 2 次
