TaxoAlign: Scholarly Taxonomy Generation Using Language Models
Avishek Lahiri, Yufang Hou, Debarshi Kumar Sanyal
摘要
Taxonomies play a crucial role in helping researchers structure and navigate knowledge in a hierarchical manner. They also form an important part in the creation of comprehensive literature surveys. The existing approaches to automatic survey generation do not compare the structure of the generated surveys with those written by human experts. To address this gap, we present our own method for automated taxonomy creation that can bridge the gap between human-generated and automaticallycreated taxonomies. For this purpose, we create the CS-TAXOBENCH benchmark which consists of 460 taxonomies that have been extracted from human-written survey papers. We also include an additional test set of 80 taxonomies curated from conference survey papers. We propose TAXOALIGN, a threephase topic-based instruction-guided method for scholarly taxonomy generation. Additionally, we propose a stringent automated evaluation framework that measures the structural alignment and semantic coherence of automatically generated taxonomies in comparison to those created by human experts. We evaluate our method and various baselines on CS-TAXOBENCH, using both automated evaluation metrics and human evaluation studies. The results show that TAXOALIGN consistently surpasses the baselines on nearly all metrics. The code and data can be found at https: //github.com/AvishekLahiri/TaxoAlign . Gold Standard Taxonomy Generated Taxonomy Human Image Generation : A Comprehensive Survey |--DATA-DRIVEN METHODS ON HUMAN IMAGE GENERATION | |--Method Taxonomy Based on Fundamental Models | |--Method Taxonomy Based on Task Settings | +--Main Components in Data-Driven Methods |--HYBRID METHODS |--KNOWLEDGE-GUIDED METHODS ON | | HUMAN IMAGE GENERATION | |--Fundamental Models of Knowledge-Guided Methods | |--Pixel Warping Pipeline | +--Virtual Rendering Pipeline |--APPLICATIONS
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- S2ORC: The Semantic Scholar Open Research CorpusKyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney 等ACL 2020 · 被引用 424 次
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang 等NeurIPS 2024 · 被引用 151 次
相关 Paper
- Context-Aware Hierarchical Taxonomy Generation for Scientific Papers via LLM-Guided Multi-Aspect ClusteringKun Zhu, Lizi Liao, Yuxuan Gu, Lei Huang 等EMNLP 2025 · 被引用 8 次
- Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom's TaxonomyFei Zhang, Zhe Zhao, Haibin Wen, Tianshuo Wei 等ACL 2026
- TaxoAdapt: Aligning LLM-Based Multidimensional Taxonomy Construction to Evolving Research CorporaPriyanka Kargupta, Nan Zhang, Yunyi Zhang, Rui Zhang 等ACL 2025
- RubricBench: Aligning Model-Generated Rubrics with Human StandardsJunyi Zhou, Qiyuan Zhang, Yufei Wang, Fuyuan Lyu 等ACL 2026 · 被引用 7 次
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsTong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi ZhangEMNLP 2025 · 被引用 2 次
