Are Large Language Models a Good Replacement of Taxonomies?
Yushi Sun, Xin Hao, Kai Sun, Yifan Xu, Xiao Yang, Xin Luna Dong, Nan Tang, Lei Chen
Abstract
Large language models (LLMs) demonstrate an impressive ability to internalize knowledge and answer natural language questions. Although previous studies validate that LLMs perform well on general knowledge while presenting poor performance on long-tail nuanced knowledge, the community is still doubtful about whether the traditional knowledge graphs should be replaced by LLMs. In this paper, we ask if the schema of knowledge graph (i.e., taxonomy) is made obsolete by LLMs. Intuitively, LLMs should perform well on common taxonomies and at taxonomy levels that are common to people. Unfortunately, there lacks a comprehensive benchmark that evaluates the LLMs over a wide range of taxonomies from common to specialized domains and at levels from root to leaf so that we can draw a confident conclusion. To narrow the research gap, we constructed a novel taxonomy hierarchical structure discovery benchmark named TaxoGlimpse to evaluate the performance of LLMs over taxonomies. TaxoGlimpse covers ten representative taxonomies from common to specialized domains with in-depth experiments of different levels of entities in this taxonomy from root to leaf. Our comprehensive experiments of eighteen LLMs under three prompting settings validate that LLMs perform miserably poorly in handling specialized taxonomies and leaf-level entities. Specifically, the QA accuracy of the best LLM drops by up to 30% as we go from common to specialized domains and from root to leaf levels of taxonomies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesKeke Huang, Yimin Shi, Dujian Ding, Yifei Li et al.VLDB 2025 · 18 citations
- Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language ModelsRuixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks et al.ACL 2026 · 6 citations
- You Are What You Bought: Generating Customer Personas for E-commerce ApplicationsYimin Shi, Yang Fei, Shiqi Zhang, Haixun Wang et al.SIGIR 2025 · 6 citations
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li et al.SIGMOD 2026 · 4 citations
- CLLMate: A Multimodal Benchmark for Weather and Climate Events ForecastingHaobo Li, Zhaowei Wang, Jiachen Wang, Yueya Wang et al.EMNLP 2025 · 2 citations
Builds on9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das et al.ACL 2023 · 233 citations
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 145 citations
Related papers
- MetaBench: A Multi-task Benchmark for Assessing LLMs in MetabolomicsYuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi et al.ACL 2026 · 1 citation
- KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsYuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan et al.WWW 2024 · 6 citations
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning et al.ICML 2025
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 33 citations
- Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge ExtractionYuheng Yang, Siqi Zhu, Tao Feng, Ge Liu et al.ICML 2026
