Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect Clustering
Kun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang, Xiaocheng Feng, Bing Qin
摘要
The rapid growth of scientific literature demands efficient methods to organize and synthesize research findings. Existing taxonomy construction methods, leveraging unsupervised clustering or direct prompting of large language models (LLMs), often lack coherence and granularity. We propose a novel context-aware hierarchical taxonomy generation framework that integrates LLM-guided multi-aspect encoding with dynamic clustering. Our method leverages LLMs to identify key aspects of each paper (e.g., methodology, dataset, evaluation) and generates aspect-specific paper summaries, which are then encoded and clustered along each aspect to form a coherent hierarchy. In addition, we introduce a new benchmark of 156 expert-crafted taxonomies encompassing 11.6 k papers, providing the first naturally annotated dataset for this task. Experimental results demonstrate that our method significantly outperforms prior approaches, achieving stateof-the-art performance in taxonomy coherence, granularity, and interpretability. 1 * Work was done during an internship at SMU. † Corresponding Author 1 Code and dataset are available in https://github.com/ zhukun1020/TaxoBench-CS .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AgenticScholar: Agentic Data Management with Pipeline Orchestration for Scholarly CorporaHai Lan, Tingting Wang, Zhifeng Bao, Guoliang Li 等SIGMOD 2026 · 被引用 4 次
- Nonparametric Deep Fine-grained Clustering with Low-Rank Guided Vision-Language Modelxulun ye, Benyu Wu, Jie Hong, Kun ZhouCVPR 2026
它引用的顶会 Paper13
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 被引用 794 次
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu 等ICLR 2024 · 被引用 395 次
- Specializing Smaller Language Models towards Multi-Step ReasoningYao Fu, Hao Peng, Litu Ou, Ashish Sabharwal 等ICML 2023 · 被引用 347 次
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang 等NeurIPS 2024 · 被引用 151 次
- The Truth is in There: Improving Reasoning in Language Models with Layer-Selective Rank ReductionPratyusha Sharma, Jordan T. Ash, Dipendra MisraICLR 2024 · 被引用 135 次
相关 Paper
- TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research CorporaPriyanka Kargupta, Nan Zhang, Yunyi Zhang, Rui Zhang 等ACL 2025
- TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsAvishek Lahiri, Yufang Hou, Debarshi Kumar SanyalEMNLP 2025
- Leveraging Large Language Models for NLG Evaluation: Advances and ChallengesZhen Li, Xiaohan Xu, Tao Shen, Can Xu 等EMNLP 2024 · 被引用 17 次
- Compress and Mix: Advancing Efficient Taxonomy Completion with Large Language ModelsHongyuan Xu, Yuhang Niu, Yanlong Wen, Xiaojie YuanWWW 2025 · 被引用 6 次
- TaxoCom: Topic Taxonomy Completion with Hierarchical Discovery of Novel Topic ClustersDongha Lee, Jiaming Shen, Seongku Kang, Susik Yoon 等WWW 2022 · 被引用 46 次
