Discovering Latent Concepts Learned in BERT
Fahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani, Jia Xu, Hassan Sajjad
摘要
A large number of studies that analyze deep neural network models and their ability to encode various linguistic and non-linguistic concepts provide an interpretation of the inner mechanics of these models. The scope of the analyses is limited to pre-defined concepts that reinforce the traditional linguistic knowledge and do not reflect on how novel concepts are learned by the model. We address this limitation by discovering and analyzing latent concepts learned in neural network models in an unsupervised fashion and provide interpretations from the model's perspective. In this work, we study: i) what latent concepts exist in the pre-trained BERT model, ii) how the discovered latent concepts align or diverge from classical linguistic hierarchy and iii) how the latent concepts evolve across layers. Our findings show: i) a model learns novel concepts (e.g. animal categories and demographic groups), which do not strictly adhere to any pre-defined categorization (e.g. POS, semantic tags), ii) several latent concepts are based on multiple properties which may include semantics, syntax, and morphology, iii) the lower layers in the model dominate in learning shallow lexical concepts while the higher layers learn semantic relations and iv) the discovered latent concepts highlight potential biases learned in the model. We also release 1 a novel BERT ConceptNet dataset (BCN) consisting of 174 concept labels and 1M annotated instances.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Zero-Shot Robustification of Zero-Shot ModelsDyah Adila, Changho Shin, Linrong Cai, Frederic SalaICLR 2024 · 被引用 31 次
- COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsHao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin 等EMNLP 2022 · 被引用 16 次
- Evaluating Neuron Interpretation Methods of NLP ModelsYimin Fan, Fahim Dalvi, Nadir Durrani, Hassan SajjadNeurIPS 2023 · 被引用 11 次
- Precise In-Parameter Concept Erasure in Large Language ModelsYoav Gur-Arieh, Clara Suslik, Yihuai Hong, Fazl Barez 等EMNLP 2025 · 被引用 10 次
- Causal Differentiating Concepts: Interpreting LM Behavior via Causal Representation LearningNavita Goyal, Hal Daumé III, Alexandre Drouin, Dhanya SridharNeurIPS 2025 · 被引用 8 次
它引用的顶会 Paper5
- Emergence of Separable Manifolds in Deep Language RepresentationsJonathan Mamou, Hang Le, Miguel Del Rio, Cory Stephenson 等ICML 2020 · 被引用 52 次
- Analyzing Individual Neurons in Pre-trained Language ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Yonatan BelinkovEMNLP 2020 · 被引用 5 次
- Similarity Analysis of Contextual Word Representation ModelsJohn M. Wu, Yonatan Belinkov, Hassan Sajjad, Nadir Durrani 等ACL 2020 · 被引用 3 次
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 被引用 3 次
- Intrinsic Probing through Dimension SelectionLucas Torroba Hennigen, Adina Williams, Ryan CotterellEMNLP 2020 · 被引用 3 次
相关 Paper
- On the Transformation of Latent Space in Fine-Tuned NLP ModelsNadir Durrani, Hassan Sajjad, Fahim Dalvi, Firoj AlamEMNLP 2022 · 被引用 3 次
- From Tokens to Lattices: Emergent Lattice Structures in Language ModelsBo Xiong, Steffen StaabICLR 2025
- Can LLMs Facilitate Interpretation of Pre-trained Language Models?Basel Mousi, Nadir Durrani, Fahim DalviEMNLP 2023 · 被引用 1 次
- Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based LearningDeepika Vemuri, Sayanta Adhikari, Ankit Saha, Krishn Vishwas Kher 等ICML 2026
- Finding Universal Grammatical Relations in Multilingual BERTEthan A. Chi, John Hewitt, Christopher D. ManningACL 2020 · 被引用 7 次
