Discursive Circuits: How Do Language Models Understand Discourse Relations?
Yisong Miao, Min-Yen Kan
摘要
Which components in transformer language models are responsible for discourse understanding? We hypothesize that sparse computational graphs, termed as discursive circuits, control how models process discourse relations. Unlike simpler tasks, discourse relations involve longer spans and complex reasoning. To make circuit discovery feasible, we introduce a task called Completion under Discourse Relation (CUDR), where a model completes a discourse given a specified relation. To support this task, we construct a corpus of minimal contrastive pairs tailored for activation patching in circuit discovery. Experiments show that sparse circuits (≈ 0.2% of a full GPT-2 model) recover discourse understanding in the English PDTB-based CUDR task. These circuits generalize well to unseen discourse frameworks such as RST and SDRT. Further analysis shows lower layers capture linguistic features such as lexical semantics and coreference, while upper layers encode discourse-level abstractions. Feature utility is consistent across frameworks (e.g., coreference supports Expansion-like relations).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Illusions of Confidence? Diagnosing LLM Truthfulness via Neighborhood ConsistencyHaoming Xu, Ningyuan Zhao, Yunzhi Yao, Weihong Xu 等ACL 2026 · 被引用 2 次
- Latent Agents: A Post-Training Procedure for Internalized Multi-Agent DebateJohn Seon Keun Yi, Aaron Mueller, Dokyun LeeACL 2026 · 被引用 1 次
- Interpreting and Enhancing Emotional Circuits in Large Vision-Language Models via Cross-Modal Information FlowChengsheng Zhang, Chenghao Sun, Zhining Xie, Xinmei TianICML 2026
- Improving Implicit Discourse Relation Recognition with Natural Language Explanations from LLMsHeng Wang, Changxing WuAAAI 2026
它引用的顶会 Paper21
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
- Towards Automated Circuit Discovery for Mechanistic InterpretabilityArthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim 等NeurIPS 2023 · 被引用 861 次
- How does GPT-2 compute greater-than?: Interpreting mathematical abilities in a pre-trained language modelMichael Hanna, Ollie Liu, Alexandre VariengienNeurIPS 2023 · 被引用 251 次
- Towards Best Practices of Activation Patching in Language Models: Metrics and MethodsFred Zhang, Neel NandaICLR 2024 · 被引用 233 次
- Circuit Component Reuse Across Tasks in Transformer Language ModelsJack Merullo, Carsten Eickhoff, Ellie PavlickICLR 2024 · 被引用 108 次
相关 Paper
- Connective Prediction for Implicit Discourse Relation Recognition via Knowledge DistillationHongyi Wu, Hao Zhou, Man Lan, Yuanbin Wu 等ACL 2023 · 被引用 7 次
- Towards Interpretable Sequence Continuation: Analyzing Shared Circuits in Large Language ModelsMichael Lan, Philip Torr, Fazl BarezEMNLP 2024 · 被引用 1 次
- Discourse-Aware Neural Extractive Text SummarizationJiacheng Xu, Zhe Gan, Yu Cheng, Jingjing LiuACL 2020 · 被引用 264 次
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov 等ICLR 2025
- Not Just Classification: Recognizing Implicit Discourse Relation on Joint Modeling of Classification and GenerationFeng Jiang, Yaxin Fan, Xiaomin Chu, Peifeng Li 等EMNLP 2021 · 被引用 14 次
