Are Large Language Models a Good Replacement of Taxonomies?
Yushi Sun, Xin Hao, Kai Sun, Yifan Xu, Xiao Yang, Xin Luna Dong, Nan Tang, Lei Chen
摘要
Large language models (LLMs) demonstrate an impressive ability to internalize knowledge and answer natural language questions. Although previous studies validate that LLMs perform well on general knowledge while presenting poor performance on long-tail nuanced knowledge, the community is still doubtful about whether the traditional knowledge graphs should be replaced by LLMs. In this paper, we ask if the schema of knowledge graph (i.e., taxonomy) is made obsolete by LLMs. Intuitively, LLMs should perform well on common taxonomies and at taxonomy levels that are common to people. Unfortunately, there lacks a comprehensive benchmark that evaluates the LLMs over a wide range of taxonomies from common to specialized domains and at levels from root to leaf so that we can draw a confident conclusion. To narrow the research gap, we constructed a novel taxonomy hierarchical structure discovery benchmark named TaxoGlimpse to evaluate the performance of LLMs over taxonomies. TaxoGlimpse covers ten representative taxonomies from common to specialized domains with in-depth experiments of different levels of entities in this taxonomy from root to leaf. Our comprehensive experiments of eighteen LLMs under three prompting settings validate that LLMs perform miserably poorly in handling specialized taxonomies and leaf-level entities. Specifically, the QA accuracy of the best LLM drops by up to 30% as we go from common to specialized domains and from root to leaf levels of taxonomies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- ThriftLLM: On Cost-Effective Selection of Large Language Models for Classification QueriesKeke Huang, Yimin Shi, Dujian Ding, Yifei Li 等VLDB 2025 · 被引用 18 次
- Sparse Feature Coactivation Reveals Causal Semantic Modules in Large Language ModelsRuixuan Deng, Xiaoyang Hu, Miles Gilberti, Shane Storks 等ACL 2026 · 被引用 6 次
- You Are What You Bought: Generating Customer Personas for E-commerce ApplicationsYimin Shi, Yang Fei, Shiqi Zhang, Haixun Wang 等SIGIR 2025 · 被引用 6 次
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li 等SIGMOD 2026 · 被引用 4 次
- CLLMate: A Multimodal Benchmark for Weather and Climate Events ForecastingHaobo Li, Zhaowei Wang, Jiachen Wang, Yueya Wang 等EMNLP 2025 · 被引用 2 次
它引用的顶会 Paper9
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric MemoriesAlex Mallen, Akari Asai, Victor Zhong, Rajarshi Das 等ACL 2023 · 被引用 233 次
- Revisiting Relation Extraction in the era of Large Language ModelsSomin Wadhwa, Silvio Amir, Byron C. WallaceACL 2023 · 被引用 145 次
相关 Paper
- MetaBench: A Multi-task Benchmark for Assessing LLMs in MetabolomicsYuxing Lu, Xukai Zhao, J. Ben Tamo, Micky C. Nnamdi 等ACL 2026 · 被引用 1 次
- KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language ModelsYuyang Bai, Shangbin Feng, Vidhisha Balachandran, Zhaoxuan Tan 等WWW 2024 · 被引用 6 次
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning 等ICML 2025
- End-to-End Ontology Learning with Large Language ModelsAndy Lo, Albert Q. Jiang, Wenda Li, Mateja JamnikNeurIPS 2024 · 被引用 33 次
- Probing the Knowledge Boundary: An Interactive Agentic Framework for Deep Knowledge ExtractionYuheng Yang, Siqi Zhu, Tao Feng, Ge Liu 等ICML 2026
