Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review Composition
Xuemei Tang, Xufeng Duan, Zhenguang G. Cai
摘要
Large language models (LLMs) have emerged as a potential solution to automate the complex processes involved in writing literature reviews, such as literature collection, organization, and summarization. However, it is yet unclear how good LLMs are at automating comprehensive and reliable literature reviews. This study introduces a framework to automatically evaluate the performance of LLMs in three key tasks of literature review writing: reference generation, abstract writing, and literature review composition. We introduce multidimensional evaluation metrics that assess the hallucination rates in generated references and measure the semantic coverage and factual consistency of the literature summaries and compositions against human-written counterparts. The experimental results reveal that even the most advanced models still generate hallucinated references, despite recent progress. Moreover, we observe that the performance of different models varies across disciplines when it comes to writing literature reviews. These findings highlight the need for further research and development to improve the reliability of LLMs in automating academic literature reviews. The dataset and code used in this study are publicly available in our GitHub repository 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation DetectionYuechen Jiang, Zhiwei Liu, Yupeng Cao, Yueru He 等ACL 2026 · 被引用 9 次
- SurveyGen: Quality-Aware Scientific Survey Generation with Large Language ModelsTong Bao, Mir Tafseer Nayeem, Davood Rafiei, Chengzhi ZhangEMNLP 2025 · 被引用 2 次
- LitReview Arena: Evaluating Literature Review Agents with Battle-style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue 等ICML 2026
- MoRI: Learning Motivation-Grounded Reasoning for Scientific Ideation in Large Language ModelsChenyang Gu, Jiahao Cheng, Meicong Zhang, Pujun Zheng 等ACL 2026
它引用的顶会 Paper4
- Enabling Large Language Models to Generate Text with CitationsTianyu Gao, Howard Yen, Jiatong Yu, Danqi ChenEMNLP 2023 · 被引用 152 次
- AutoSurvey: Large Language Models Can Automatically Write SurveysYidong Wang, Qi Guo, Wenjin Yao, Hongbo Zhang 等NeurIPS 2024 · 被引用 151 次
- Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage ScholarsJiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan 等CHI 2024 · 被引用 21 次
- EFUF: Efficient Fine-Grained Unlearning Framework for Mitigating Hallucinations in Multimodal Large Language ModelsShangyu Xing, Fei Zhao, Zhen Wu, Tuo An 等EMNLP 2024 · 被引用 6 次
相关 Paper
- Appraising the Potential Uses and Harms of LLMs for Medical Systematic ReviewsHye Sun Yun, Iain James Marshall, Thomas A. Trikalinos, Byron C. WallaceEMNLP 2023 · 被引用 11 次
- LLMs can Perform Multi-Dimensional Analytic Writing Assessments: A Case Study of L2 Graduate-Level Academic English WritingZhengxiang Wang, Veronika Makarova, Zhi Li, Jordan Kodner 等ACL 2025 · 被引用 5 次
- Systematic Task Exploration with LLMs: A Study in Citation Text GenerationFurkan Sahinuç, Ilia Kuznetsov, Yufang Hou, Iryna GurevychACL 2024
- Mixture of Knowledge Minigraph Agents for Literature Review GenerationZhi Zhang, Yan Liu, Sheng-hua Zhong, Gong Chen 等AAAI 2025 · 被引用 1 次
- CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based VerificationYuchen Tian, Weixiang Yan, Qian Yang, Xuandong Zhao 等AAAI 2025 · 被引用 41 次
