Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom's Taxonomy
Fei Zhang, Zhe Zhao, Haibin Wen, Tianshuo Wei, Zaixi Zhang, Chao Yang, Ye Wei
Abstract
The rapid advance of automatic survey generation (ASG) has created a critical evaluation challenge. Existing evaluation methods suffer from both cognitive dimensional simplification and methodological unreliability, primarily due to the over-reliance on the "LLM-as-a-Judge" approach. To bridge this gap, we establish Bloom-Eval, a six-tiered benchmark based on Bloom's Taxonomy that reliably evaluates ASG systems by prioritizing deterministic algorithms and introducing our GRADE approach for abstract abilities. Furthermore, we construct a large-scale, cross-disciplinary dataset of over 3,000 high-quality papers. Our empirical study on this benchmark reveals that while leading ASG systems are proficient format organizers, they remain unqualified knowledge integrators. This work aims to redefine ASG evaluation standards, shifting the research focus from the formal mimicry of surface structure to the cognitive deepening of intellectual content. Our method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c8d56222-cb2b-4d7d-8f6b-ac151ba71d4eBuilds on7
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang et al.ICLR 2026 · 250 citations
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 173 citations
- Multi-Granularity Interaction Network for Extractive and Abstractive Multi-Document SummarizationHanqi Jin, Tianming Wang, Xiaojun WanACL 2020 · 92 citations
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang et al.ACL 2022 · 62 citations
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 39 citations
Related papers
- AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMsQingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu et al.ACL 2026
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning et al.ICML 2025
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang et al.EMNLP 2024 · 37 citations
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsTong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi ZhangEMNLP 2025 · 2 citations
- TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsAvishek Lahiri, Yufang Hou, Debarshi Kumar SanyalEMNLP 2025
