Can LLMs Facilitate Interpretation of Pre-trained Language Models?
Basel Mousi, Nadir Durrani, Fahim Dalvi
Abstract
Work done to uncover the knowledge encoded within pre-trained language models rely on annotated corpora or human-in-the-loop methods. However, these approaches are limited in terms of scalability and the scope of interpretation. We propose using a large language model, Chat-GPT, as an annotator to enable fine-grained interpretation analysis of pre-trained language models. We discover latent concepts within pre-trained language models by applying agglomerative hierarchical clustering over contextualized representations and then annotate these concepts using ChatGPT. Our findings demonstrate that ChatGPT produces accurate and semantically richer annotations compared to human-annotated concepts. Additionally, we showcase how GPT-based annotations empower interpretation analysis methodologies of which we demonstrate two: probing frameworks and neuron interpretation. To facilitate further exploration and experimentation in the field, we make available a substantial Concept-Net dataset (TCN) comprising 39,000 annotated concepts. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a5d178fa-392e-48db-8d34-17d65d11616aCited by top-tier papers5
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon et al.ICML 2024 · 197 citations
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 11 citations
- Do Activation Verbalization Methods Convey Privileged Information?Millicent Li, Alberto Mario Ceballos Arroyo, Giordano Rogers, Naomi Saphra et al.ICML 2026 · 5 citations
- Latent Concept-based Explanation of NLP ModelsXuemin Yu, Fahim Dalvi, Nadir Durrani, Marzia Nouri et al.EMNLP 2024 · 3 citations
- Exploring Alignment in Shared Cross-lingual SpacesBasel Mousi, Nadir Durrani, Fahim Dalvi, Majd Hawasly et al.ACL 2024 · 1 citation
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Locating and Editing Factual Associations in GPTKevin Meng, David Bau, Alex Andonian, Yonatan BelinkovNeurIPS 2022 · 3,415 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani et al.ICLR 2022 · 74 citations
- Overlooked Factors in Concept-Based Explanations: Dataset Choice, Concept Learnability, and Human CapabilityVikram V. Ramaswamy, Sunnie S. Y. Kim, Ruth Fong, Olga RussakovskyCVPR 2023
- Bootstrapping Variational Information Pursuit with Large Language and Vision Models for Interpretable Image ClassificationAditya Chattopadhyay, Kwan Ho Ryan Chan, René VidalICLR 2024 · 12 citations
- HiBug: On Human-Interpretable Model DebugMuxi Chen, Yu Li, Qiang XuNeurIPS 2023 · 22 citations
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
