To Steer or Not to Steer? Mechanistic Error Reduction with Abstention for Language Models
Anna Hedström, Salim I. Amoukou, Tom Bewley, Saumitra Mishra, Manuela Veloso
Abstract
We introduce Mechanistic Error Reduction with Abstention (MERA), a principled framework for steering language models (LMs) to mitigate errors through selective, adaptive interventions. Unlike existing methods that rely on fixed, manually tuned steering strengths, often resulting in under or oversteering, MERA addresses these limitations by (i) optimising the intervention direction, and (ii) calibrating when, and how much to steer, thereby provably improving performance or abstaining when no confident correction is possible. Experiments across diverse datasets, and LM families demonstrate safe, effective, non-degrading error correction, and that MERA outperforms existing baselines. Moreover, MERA can be applied on top of existing steering techniques to further enhance their performance, establishing it as a general-purpose, and efficient approach to mechanistic activation steering.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2f0ae6a5-2358-401a-aded-cb19469dc8d8Cited by top-tier papers10
- Into the Rabbit Hull: From Task-Relevant Concepts in DINO to Minkowski GeometryThomas Fel, Binxu Wang, Michael A. Lepori, Matthew Kowal et al.ICLR 2026 · 28 citations
- Enhancing Instruction Following of LLMs via Activation Steering with Dynamic RejectionMinjae Kang, Jaehyung KimICLR 2026 · 6 citations
- Exploring Diverse Generation Paths via Inference-time Stiefel Activation SteeringDongxuan Zhu, Ly Tran Ho Khanh, Andy Yat-Ming Cheung, Man-Chung Yue et al.ICLR 2026 · 4 citations
- Local Linearity of LLMs Enables Activation Steering via Model-Based Linear Optimal ControlJulian Skifstad, Xinyue Annie Yang, Glen ChouICML 2026 · 2 citations
- Steering at the Source: Style Modulation Heads for Robust Persona ControlYoshihiro Izawa, Gouki Minegishi, Koshi Eguchi, Sosuke Hosokawa et al.ICML 2026 · 2 citations
Builds on18
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
Related papers
- A Simple Yet Effective Method for Non-Refusing Context Relevant Fine-grained Safety Steering in LLMsShaona Ghosh, Amrita Bhattacharjee, Yftah Ziser, Christopher ParisienEMNLP 2025
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang et al.NeurIPS 2025 · 9 citations
- Steer Like the LLM: Activation Steering that Mimics PromptingGeert Heyman, Frederik VandeputteICML 2026
- Angular Steering: Behavior Control via Rotation in Activation SpaceMinh Hieu Vu, Tan M. NguyenNeurIPS 2025 · 53 citations
- Improving Instruction-Following in Language Models through Activation SteeringAlessandro Stolfo, Vidhisha Balachandran, Safoora Yousefi, Eric Horvitz et al.ICLR 2025
