From Tokens to Lattices: Emergent Lattice Structures in Language Models
Bo Xiong, Steffen Staab
Abstract
Pretrained masked language models (MLMs) have demonstrated an impressive capability to comprehend and encode conceptual knowledge, revealing a lattice structure among concepts. This raises a critical question: how does this conceptualization emerge from MLM pretraining? In this paper, we explore this problem from the perspective of Formal Concept Analysis (FCA), a mathematical framework that derives concept lattices from the observations of object-attribute relationships. We show that the MLM's objective implicitly learns a formal context that describes objects, attributes, and their dependencies, which enables the reconstruction of a concept lattice through FCA. We propose a novel framework for concept lattice construction from pretrained MLMs and investigate the origin of the inductive biases of MLMs in lattice structure learning. Our framework differs from previous work because it does not rely on human-defined concepts and allows for discovering "latent" concepts that extend beyond human definitions. We create three datasets for evaluation, and the empirical results verify our hypothesis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3a519a66-2a78-4bc7-816d-09928e22e2b1Cited by top-tier papers1
Ask how each one uses itBuilds on6
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Discovering Latent Concepts Learned in BERTFahim Dalvi, Abdul Rafae Khan, Firoj Alam, Nadir Durrani et al.ICLR 2022 · 74 citations
- Hyperbolic Embedding Inference for Structured Multi-Label PredictionBo Xiong, Michael Cochez, Mojtaba Nayyeri, Steffen StaabNeurIPS 2022 · 23 citations
- COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsHao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin et al.EMNLP 2022 · 16 citations
- Do PLMs Know and Understand Ontological Knowledge?Weiqi Wu, Chengyue Jiang, Yong Jiang, Pengjun Xie et al.ACL 2023 · 13 citations
Related papers
- Formal Concept Lattices are Good Semantic Scaffolds for Concept-Based LearningDeepika Vemuri, Sayanta Adhikari, Ankit Saha, Krishn Vishwas Kher et al.ICML 2026
- Asking without Telling: Exploring Latent Ontologies in Contextual RepresentationsJulian Michael, Jan A. Botha, Ian TenneyEMNLP 2020 · 3 citations
- Pre-training Language Models with Deterministic Factual KnowledgeShaobo Li, Xiaoguang Li, Lifeng Shang, Chengjie Sun et al.EMNLP 2022 · 13 citations
- Exploiting Structured Knowledge in Text via Graph-Guided Representation LearningTao Shen, Yi Mao, Pengcheng He, Guodong Long et al.EMNLP 2020 · 60 citations
- Knowledgeable or Educated Guess? Revisiting Language Models as Knowledge BasesBoxi Cao, Hongyu Lin, Xianpei Han, Le Sun et al.ACL 2021
