Concept Bottleneck Large Language Models
Chung-En Sun, Tuomas P. Oikarinen, Berk Ustun, Tsui-Wei Weng
摘要
We introduce Concept Bottleneck Large Language Models (CB-LLMs), a novel framework for building inherently interpretable Large Language Models (LLMs). In contrast to traditional black-box LLMs that rely on limited post-hoc interpretations, CB-LLMs integrate intrinsic interpretability directly into the LLMs -allowing accurate explanations with scalability and transparency. We build CB-LLMs for two essential NLP tasks: text classification and text generation. In text classification, CB-LLMs is competitive with, and at times outperforms, traditional black-box models while providing explicit and interpretable reasoning. For the more challenging task of text generation, interpretable neurons in CB-LLMs enable precise concept detection, controlled generation, and safer outputs. The embedded interpretability empowers users to transparently identify harmful content, steer model behavior, and unlearn undesired concepts -significantly enhancing the safety, reliability, and trustworthiness of LLMs, which are critical capabilities notably absent in existing language models. Our code is available at https://github.com/Trustworthy- ML-Lab/CB-LLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Object-Centric Concept-BottlenecksDavid Steinmann, Wolfgang Stammer, Antonia Wüst, Kristian KerstingNeurIPS 2025 · 被引用 12 次
- An Analysis of Concept Bottleneck Models: Measuring, Understanding, and Mitigating the Impact of Noisy AnnotationsSeonghwan Park, Jueun Mun, Donghyun Oh, Namhoon LeeNeurIPS 2025 · 被引用 10 次
- Reasoning Scaffolding: Distilling the Flow of Thought from LLMsXiangyu Wen, Junhua Huang, Zeju Li, Min Li 等ICLR 2026 · 被引用 7 次
- Interpretable and Steerable Concept Bottleneck Sparse AutoencodersAkshay Kulkarni, Tsui-Wei Weng, Vivek Narayanaswamy, Shusen Liu 等CVPR 2026 · 被引用 6 次
- Interpretable Next-token Prediction via the Generalized Induction HeadEunji Kim, Sriya Mantena, Weiwei Yang, Chandan Singh 等NeurIPS 2025 · 被引用 3 次
它引用的顶会 Paper10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann 等ICML 2020 · 被引用 1,233 次
- VLG-CBM: Training Concept Bottleneck Models with Vision-Language GuidanceDivyansh Srivastava, Ge Yan, Lily WengNeurIPS 2024 · 被引用 87 次
- Concept Bottleneck Generative ModelsAya Abdelsalam Ismail, Julius Adebayo, Héctor Corrada Bravo, Stephen Ra 等ICLR 2024 · 被引用 43 次
相关 Paper
- Bayesian Concept Bottleneck Models with LLM PriorsJean Feng, Avni Kothari, Lucas Zier, Chandan Singh 等NeurIPS 2025 · 被引用 23 次
- Concept Bottleneck Language Models For Protein DesignAya Abdelsalam Ismail, Tuomas P. Oikarinen, Amy Wang, Julius Adebayo 等ICLR 2025
- Hybrid Concept Bottleneck ModelsYang Liu, Tianwei Zhang, Shi GuCVPR 2025
- Explanation Bottleneck ModelsShin'ya Yamaguchi, Kosuke NishidaAAAI 2025 · 被引用 4 次
- Towards Intrinsic Interpretability of Large Language Models: A Survey of Design Principles and ArchitecturesYutong Gao, Qinglin Meng, Yuan Zhou, Liangming PanACL 2026 · 被引用 3 次
