Bloom-Eval: A Hierarchical Evaluation Benchmark for Automatic Survey Generation Based on Bloom's Taxonomy
Fei Zhang, Zhe Zhao, Haibin Wen, Tianshuo Wei, Zaixi Zhang, Chao Yang, Ye Wei
摘要
The rapid advance of automatic survey generation (ASG) has created a critical evaluation challenge. Existing evaluation methods suffer from both cognitive dimensional simplification and methodological unreliability, primarily due to the over-reliance on the "LLM-as-a-Judge" approach. To bridge this gap, we establish Bloom-Eval, a six-tiered benchmark based on Bloom's Taxonomy that reliably evaluates ASG systems by prioritizing deterministic algorithms and introducing our GRADE approach for abstract abilities. Furthermore, we construct a large-scale, cross-disciplinary dataset of over 3,000 high-quality papers. Our empirical study on this benchmark reveals that while leading ASG systems are proficient format organizers, they remain unqualified knowledge integrators. This work aims to redefine ASG evaluation standards, shifting the research focus from the formal mimicry of surface structure to the cognitive deepening of intellectual content. Our method provides the ASG field with a systematic, reproducible, and theoretically grounded benchmark to guide future research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research AgentsMingxuan Du, Benfeng Xu, Chiwei Zhu, Licheng Zhang 等ICLR 2026 · 被引用 250 次
- BooookScore: A systematic exploration of book-length summarization in the era of LLMsYapei Chang, Kyle Lo, Tanya Goyal, Mohit IyyerICLR 2024 · 被引用 173 次
- Multi-Granularity Interaction Network for Extractive and Abstractive Multi-Document SummarizationHanqi Jin, Tianming Wang, Xiaojun WanACL 2020 · 被引用 92 次
- DYLE: Dynamic Latent Extraction for Abstractive Long-Input SummarizationZiming Mao, Chen Henry Wu, Ansong Ni, Yusen Zhang 等ACL 2022 · 被引用 62 次
- Automatic Generation of Citation Texts in Scholarly Papers: A Pilot StudyXinyu Xing, Xiaosheng Fan, Xiaojun WanACL 2020 · 被引用 39 次
相关 Paper
- AutoTaskEval: Towards Domain-Specific and Fine-Grained Evaluation for LLMsQingqing Lyu, Linjuan Wu, Yongliang Shen, Hengwei Liu 等ACL 2026
- Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem SolvingYuxuan Zhou, Xien Liu, Chenwei Yan, Chen Ning 等ICML 2025
- Humans or LLMs as the Judge? A Study on Judgement BiasGuiming Hardy Chen, Shunian Chen, Ziche Liu, Feng Jiang 等EMNLP 2024 · 被引用 37 次
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsTong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi ZhangEMNLP 2025 · 被引用 2 次
- TaxoAlign: Scholarly Taxonomy Generation Using Language ModelsAvishek Lahiri, Yufang Hou, Debarshi Kumar SanyalEMNLP 2025
