Lune

ICML2025Top-tier venue

The Logical Implication Steering Method for Conditional Interventions on Transformer Generation

Damjan Kalajdzievski

2025Year

Abstract

The field of mechanistic interpretability in pretrained transformer models has demonstrated substantial evidence supporting the "linear representation hypothesis", which is the idea that high level concepts are encoded as vectors in the space of activations of a model. Studies also show that model generation behavior can be steered toward a given concept by adding the concept's vector to the corresponding activations. We show how to leverage these properties to build a form of logical implication into models, enabling transparent and interpretable adjustments that induce a chosen generation behavior in response to the presence of any given concept. Our method, Logical Implication Model Steering (LIMS), unlocks new hand-engineered reasoning capabilities by integrating neuro-symbolic logic into pre-trained transformer models.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e371ef90-9243-481f-82e3-b4237e9b5a96

Builds on13

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines