Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and Architectures
Yutong Gao, Qinglin Meng, Yuan Zhou, Liangming Pan
摘要
While Large Language Models (LLMs) have achieved strong performance across many NLP tasks, their opaque internal mechanisms hinder trustworthiness and safe deployment. Existing surveys in explainable AI largely focus on post-hoc explanation methods that interpret trained models through external approximations. In contrast, intrinsic interpretability, which builds transparency directly into model architectures and computations, has recently emerged as a promising alternative. This paper presents a systematic review of the recent advances in intrinsic interpretability for LLMs, categorizing existing approaches into five design paradigms: functional transparency, concept alignment, representational decomposability, explicit modularization, and latent sparsity induction. We further discuss open challenges and outline future research directions in this emerging field. The paper list is available at: https://github.com/PKU-PILLAR-Group/Survey-Intrinsic-Interpretability-of-LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper28
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 被引用 3,415 次
- Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought PromptingMiles Turpin, Julian Michael, Ethan Perez, Samuel R. BowmanNeurIPS 2023 · 被引用 1,792 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- Sparse Autoencoders Find Highly Interpretable Features in Language ModelsRobert Huben, Hoagy Cunningham, Logan Riggs Smith, Aidan Ewart 等ICLR 2024 · 被引用 1,072 次
相关 Paper
- Explainability and Interpretability of Multilingual Large Language Models: A SurveyLucas Resck, Isabelle Augenstein, Anna KorhonenEMNLP 2025
- Concept Bottleneck Large Language ModelsChung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei WengICLR 2025
- Mechanisms vs. Outcomes: Probing for Syntax Fails to Explain Performance on Targeted Syntactic EvaluationsAnanth Agarwal, Jasper Jian, Christopher D. Manning, Shikhar MurtyEMNLP 2025 · 被引用 5 次
- Making Sense of LLM Decisions: A Prototype-based Framework for Explainable ClassificationBowen Wei, Mehrdad Fazli, Ziwei ZhuAAAI 2026
- Bridging Internal Consistency and External Alignment: A Causal and Dynamic Interpretability Framework for LLM GenerationShuyao Xiao, Shengling Wang, Ke ChaoACL 2026
