Discovering Latent Concepts Learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, Hassan Sajjad
Abstract
A large number of studies that analyze deep neural network models and their ability to encode various linguistic and non-linguistic concepts provide an interpretation of the inner mechanics of these models. The scope of the analyses is limited to pre-defined concepts that reinforce the traditional linguistic knowledge and do not reflect on how novel concepts are learned by the model. We address this limitation by discovering and analyzing latent concepts learned in neural network models in an unsupervised fashion and provide interpretations from the model's perspective. In this work, we study: i) what latent concepts exist in the pre-trained BERT model, ii) how the discovered latent concepts align or diverge from classical linguistic hierarchy and iii) how the latent concepts evolve across layers. Our findings show: i) a model learns novel concepts (e.g. animal categories and demographic groups), which do not strictly adhere to any pre-defined categorization (e.g. POS, semantic tags), ii) several latent concepts are based on multiple properties which may include semantics, syntax, and morphology, iii) the lower layers in the model dominate in learning shallow lexical concepts while the higher layers learn semantic relations and iv) the discovered latent concepts highlight potential biases learned in the model. We also release 1 a novel BERT ConceptNet dataset (BCN) consisting of 174 concept labels and 1M annotated instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Zero-Shot Robustification of Zero-Shot ModelsDyah Adila, Changho Shin, Linrong Cai, Frederic SalaICLR 2024 · 31 citations
- COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsHao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin et al.EMNLP 2022 · 16 citations
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 11 citations
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez et al.EMNLP 2025 · 10 citations
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 8 citations
Builds on5
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson et al.ICML 2020 · 52 citations
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 5 citations
- Similarity Analysis of Contextual Word Representation ModelsJohn M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani et al.ACL 2020 · 3 citations
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 3 citations
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 3 citations
Related papers
- On the Transformation of Latent Space in Fine-Tuned NLP ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj AlamEMNLP 2022 · 3 citations
- From Tokens to Lattices: Emergent Lattice Structures in Language ModelsBo Xiong, Steffen StaabICLR 2025
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 1 citation
- Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based LearningDeepika Vemuri, Sayanta Adhikari, Ankit Saha, Krishn Vishwas Kher et al.ICML 2026
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 7 citations
