Learning Distribution-wise Control in Representation Space for Language Models
Chunyuan Deng, Ruidi Chang, Hanjie Chen
Abstract
Interventions in language models (LMs) are applied strategically to steer model behavior during the forward pass. Learnable interventions, also known as representation fine-tuning, aim to apply pointwise control within the concept subspace and have proven effective in altering highlevel behaviors. In this work, we extend this approach to the distribution level, enabling the model to learn not only pointwise transformations but also the surrounding regions of the concept subspace. We demonstrate that these methods perform effectively in early layers, with larger standard deviations correlating strongly with improved performance. Across eight commonsense reasoning and seven arithmetic reasoning benchmarks, our distribution-wise interventions consistently outperform pointwise interventions in controllability and robustness. These results illustrate that distribution-wise interventions provide a more comprehensive method for steering model behavior and enabling finer-grained control over language models. The code is at: https://github.com/chili-lab/D-Intervention .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a8c0f3f-34f6-4f1c-baa0-5de63e9bd5e8Cited by top-tier papers2
- Spherical Steering: Geometry-Aware Activation Rotation for Language ModelsZejia You, Chunyuan Deng, Hanjie ChenICML 2026 · 12 citations
- Steering Information Utility in Key-Value Memory for Language Model Post-TrainingChunyuan Deng, Ruidi Chang, Hanjie ChenNeurIPS 2025 · 2 citations
Builds on29
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Inference-Time Intervention: Eliciting Truthful Answers from a Language ModelKenneth Li, Oam Patel, Fernanda B. Viégas, Hanspeter Pfister et al.NeurIPS 2023 · 1,549 citations
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim et al.NeurIPS 2023 · 861 citations
- Causal Abstractions of Neural NetworksAtticus Geiger, Hanson Lu, Thomas Icard, Christopher PottsNeurIPS 2021 · 516 citations
Related papers
- Improving Reasoning Performance in Large Language Models via Representation EngineeringBertram Højer, Oliver Simon Jarvis, Stefan HeinrichICLR 2025
- Concept Concentration for Faithful Representation InterventionHongzheng Yang, Yongqiang Chen, Zeyu Qin, Tongliang Liu et al.ICML 2026
- Steering into New Embedding Spaces: Analyzing Cross-Lingual Alignment Induced by Model Interventions in Multilingual Language ModelsAnirudh Sundar, Sinead Williamson, Katherine Metcalf, Barry-John Theobald et al.ACL 2025 · 8 citations
- InT: Self-Proposed Interventions Enable Credit Assignment in LLM ReasoningMatthew Y. R. Yang, Hao Bai, Ian Wu, Gene Yang et al.ICLR 2026 · 12 citations
- Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMsHaoming Meng, Kexin Huang, Shaohang Wei, Chiyu Ma et al.ICLR 2026 · 24 citations
