Lune

EMNLP2024顶会

Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language Models

Michael Lan, Philip Torr, Fazl Barez

2024年份
1被引次数
5顶会引用

摘要

While transformer models exhibit strong capabilities on linguistic tasks, their complex architectures make them difficult to interpret.Recent work has aimed to reverse engineer transformer models into human-readable representations called circuits that implement algorithmic functions.We extend this research by analyzing and comparing circuits for similar sequence continuation tasks, which include increasing sequences of Arabic numerals, number words, and months.By applying circuit interpretability analysis, we identify a key sub-circuit in both GPT-2 Small and Llama-2-7B responsible for detecting sequence members and for predicting the next member in a sequence.Our analysis reveals that semantically related sequences rely on shared circuit subgraphs with analogous roles.Additionally, we show that this sub-circuit has effects on various math-related prompts, such as on intervaled circuits, Spanish number word and months continuation, and natural language word problems.This mechanistic understanding of transformers is a critical step towards building more robust, aligned, and interpretable language models. 1 To encourage reuse and further development our code and datasets can be found here:

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖