Evaluating Large Language Models on Academic Literature Understanding and Review: An Empirical Study among Early-stage Scholars
Jiyao Wang, Haolong Hu, Zuyuan Wang, Song Yan, Youyu Sheng, Dengbo He
Abstract
The rapid advancement of large language models (LLMs) such as ChatGPT makes LLM-based academic tools possible. However, little research has empirically evaluated how scholars perform different types of academic tasks with LLMs. Through an empirical study followed by a semi-structured interview, we assessed 48 early-stage scholars’ performance in conducting core academic activities (i.e., paper reading and literature reviews) under different levels of time pressure. Before conducting the tasks, participants received different training programs regarding the limitations and capabilities of the LLMs. After completing the tasks, participants completed an interview. Quantitative data regarding the influence of time pressure, task type, and training program on participants’ performance in academic tasks was analyzed. Semi-structured interviews provided additional information on the influential factors of task performance, participants’ perceptions of LLMs, and concerns about integrating LLMs into academic workflows. The findings can guide more appropriate usage and design of LLM-based tools in assisting academic work.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 88429187-2cea-4b98-9965-390af470736aCited by top-tier papers8
- Understanding the LLM-ification of CHI: Unpacking the Impact of LLMs at CHI through a Systematic Literature ReviewRock Yuren Pang, Hope Schroeder, Kynnedy Simone Smith, Solon Barocas et al.CHI 2025 · 51 citations
- How CO2STLY Is CHI? The Carbon Footprint of Generative AI in HCI Research and What We Should Do About ItNanna Inie, Jeanette Falk, Raghavendra SelvanCHI 2025 · 33 citations
- Co-Writing with AI, on Human Terms: Aligning Research with User Demands Across the Writing ProcessMohi Reza, Jeb Thomas-Mitchell, Peter Dushniku, Nathan Laundry et al.CSCW 2025 · 29 citations
- Small, Medium, Large? A Meta-Study of Effect Sizes at CHI to Aid Interpretation of Effect Sizes and Power CalculationAnna-Marie Ortloff, Florin Martius, Mischa Meier, Theo Raimbault et al.CHI 2025 · 17 citations
- Large Language Models for Automated Literature Review: An Evaluation of Reference Generation, Abstract Writing, and Review CompositionXuemei Tang, Xufeng Duan, Zhenguang G. CaiEMNLP 2025 · 5 citations
Related papers
- Understanding the Effect of Risk Perception on the Acceptance and Use of Large Language Models Among University StudentsMichael T. Rücker, Carolin Büchting, Thomas KoschCSCW 2025 · 4 citations
- "If the Machine Is As Good As Me, Then What Use Am I?" - How the Use of ChatGPT Changes Young Professionals' Perception of Productivity and AccomplishmentCharlotte Kobiella, Yarhy Said Flores López, Franz Waltenberger, Fiona Draxler et al.CHI 2024 · 54 citations
- An Empirical Study to Understand How Students Use ChatGPT for Writing EssaysAndrew Jelson, Daniel Manesh, Alice Jang, Daniel Dunlap et al.CHI 2026 · 3 citations
- Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic ProcrastinationAnanya Bhattacharjee, Yuchen Zeng, Sarah Yi Xu, Dana Kulzhabayeva et al.CHI 2024 · 39 citations
- Beyond Code Generation: An Observational Study of ChatGPT Usage in Software Engineering PracticeRanim Khojah, Mazen Mohamad, Philipp Leitner, Francisco Gomes de Oliveira NetoFSE 2024 · 56 citations
