Lune

EMNLP2024Top-tier venue

Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Michael Lan, Philip Torr, Fazl Barez

2024Year
1Citations
5Top-tier citations

Abstract

While transformer models exhibit strong capabilities on linguistic tasks, their complex architectures make them difficult to interpret.Recent work has aimed to reverse engineer transformer models into human-readable representations called circuits that implement algorithmic functions.We extend this research by analyzing and comparing circuits for similar sequence continuation tasks, which include increasing sequences of Arabic numerals, number words, and months.By applying circuit interpretability analysis, we identify a key sub-circuit in both GPT-2 Small and Llama-2-7B responsible for detecting sequence members and for predicting the next member in a sequence.Our analysis reveals that semantically related sequences rely on shared circuit subgraphs with analogous roles.Additionally, we show that this sub-circuit has effects on various math-related prompts, such as on intervaled circuits, Spanish number word and months continuation, and natural language word problems.This mechanistic understanding of transformers is a critical step towards building more robust, aligned, and interpretable language models. 1 To encourage reuse and further development our code and datasets can be found here:

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext caf4653a-4e27-4914-8bfd-43b85b833ff0

Cited by top-tier papers5

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines