The Logical Implication Steering Method for Conditional Interventions on Transformer Generation
Damjan Kalajdzievski
Abstract
The field of mechanistic interpretability in pretrained transformer models has demonstrated substantial evidence supporting the "linear representation hypothesis", which is the idea that high level concepts are encoded as vectors in the space of activations of a model. Studies also show that model generation behavior can be steered toward a given concept by adding the concept's vector to the corresponding activations. We show how to leverage these properties to build a form of logical implication into models, enabling transparent and interpretable adjustments that induce a chosen generation behavior in response to the presence of any given concept. Our method, Logical Implication Model Steering (LIMS), unlocks new hand-engineered reasoning capabilities by integrating neuro-symbolic logic into pre-trained transformer models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e371ef90-9243-481f-82e3-b4237e9b5a96Builds on13
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Chain-of-Thought Reasoning Without PromptingXuezhi Wang, Denny ZhouNeurIPS 2024 · 305 citations
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 258 citations
Related papers
- Learning to Disentangle Latent Reasoning Rules with Language VAEs: A Systematic StudyYingji Zhang, Marco Valentino, Danilo S. Carvalho, André FreitasAAAI 2026 · 1 citation
- Latent Concept Disentanglement in Transformer-based Language ModelsGuanzhe Hong, Bhavya Vasudeva, Vatsal Sharan, Cyrus Rashtchian et al.ICLR 2026 · 4 citations
- Towards a Unified Paradigm of Concept Editing in Large Language ModelsZhuowen Han, Xinwei Wu, Dan Shi, Renren Jin et al.EMNLP 2025
- On the Origins of Linear Representations in Large Language ModelsYibo Jiang, Goutham Rajendran, Pradeep Kumar Ravikumar, Bryon Aragam et al.ICML 2024 · 68 citations
- ActivationReasoning: Logical Reasoning in Latent Activation SpacesLukas Helff, Ruben Härle, Wolfgang Stammer, Felix Friedrich et al.ICLR 2026 · 6 citations
