"Here the GPT made a choice, and every choice can be biased": How Students Critically Engage with LLMs through End-User Auditing Activity
Snehal Prabhudesai, Ananya Prashant Kasi, Anmol Mansingh, Anindya Das Antar, Hua Shen, Nikola Banovic
Abstract
Despite recognizing that Large Language Models (LLMs) can generate inaccurate or unacceptable responses, universities are increasingly making such models available to their students.Existing university policies defer the responsibility of checking for correctness and appropriateness of LLM responses to students and assume that they will have the required knowledge and skills to do so on their own.In this work, we conducted a series of user studies with students (N=47) from a large North American public research university to understand if and how they critically engage with LLMs.Our participants evaluated an LLM provided by the university in a quasi-experimental setup; first by themselves, and then with a scaffolded design probe that guided them through an end-user auditing exercise.Qualitative analysis of participant think-aloud and LLM interaction data showed that students without basic AI literacy skills struggle to conceptualize and evaluate LLM biases on their own.However, they transition to focused thinking and purposeful interactions when provided with structured guidance.We highlight areas where current university policies may fall short and offer policy and design recommendations to better support students.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 8e5a4784-2c9f-4927-92b3-37470555bca8Cited by top-tier papers8
- The Siren Song of LLMs: How Users Perceive and Respond to Dark Patterns in Large Language ModelsYike Shi, Qing Xiao, Qing Hu, Hong Shen et al.CHI 2026 · 7 citations
- Exploring Teacher-Chatbot Interaction and Affect in Block-Based ProgrammingBahare Riahi, Ally Limke, Xiaoyi Tian, Viktoriia Storozhevykh et al.CHI 2026 · 2 citations
- What Do People Want to Know about Artificial Intelligence (AI)? The Importance of Answering End-user Questions to Explain Autonomous Vehicle (AV) DecisionsSomayeh Molaei, Lionel Peter Robert, Nikola BanovicCSCW 2025 · 2 citations
- Power Echoes: Investigating Moderation Biases in Online Power-Asymmetric ConflictsYaqiong Li, Peng Zhang, Peixu Hou, Kainan Tu et al.CHI 2026 · 2 citations
- AI as We Describe It: How Large Language Models and Their Applications in Health are Represented Across Channels of Public DiscourseJiawei Zhou, Lei Zhang, Mei Li, Benjamin D. Horne et al.CHI 2026 · 1 citation
Related papers
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and InconsistenciesSunnie S. Y. Kim, Jennifer Wortman Vaughan, Q. Vera Liao, Tania Lombrozo et al.CHI 2025 · 118 citations
- How Beginning Programmers and Code LLMs (Mis)read Each OtherSydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha et al.CHI 2024 · 66 citations
- Understanding the Effect of Risk Perception on the Acceptance and Use of Large Language Models Among University StudentsMichael T. Rücker, Carolin Büchting, Thomas KoschCSCW 2025 · 4 citations
- Understanding the Role of Large Language Models in Personalizing and Scaffolding Strategies to Combat Academic ProcrastinationAnanya Bhattacharjee, Yuchen Zeng, Sarah Yi Xu, Dana Kulzhabayeva et al.CHI 2024 · 39 citations
- Effects of LLM-based Search on Decision Making: Speed, Accuracy, and OverrelianceSofia Eleni Spatharioti, David M. Rothschild, Daniel G. Goldstein, Jake M. HofmanCHI 2025 · 28 citations
