When AI Gives Advice: Evaluating AI and Human Responses to Online Advice-Seeking for Well-Being
Harsh Kumar, Jasmine Chahal, Yinuo Zhao, Zeling Zhang, Annika Z. Wei, Louis Tay, Ashton Anderson
Abstract
Seeking advice is a core human behavior that the internet has reinvented twice: first through forums and Q&A communities that crowdsource public guidance, and now through large language models (LLMs). Yet the quality of this LLM advice for everyday well-being scenarios remains unclear. How does it compare, not only against human comments, but against the wisdom of the online crowd? We ran two studies (N=210) in which experts compared top-voted Reddit advice with LLM-generated advice. LLMs ranked significantly higher overall and on effectiveness, warmth, and willingness to seek advice again. GPT-4o beat GPT-5 on all metrics except sycophancy, suggesting that benchmark gains need not improve advice-giving. In Study-2, we examined how human and algorithmic advice could be combined, and found that human advice can be unobtrusively polished to compete with AI-generated comments. We conclude with design implications for advice-giving agents and ecosystems blending AI, crowd input, and expert oversight.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Towards Understanding Sycophancy in Language ModelsMrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud et al.ICLR 2024 · 762 citations
- Measuring and Understanding Trust Calibrations for Automated Systems: A Survey of the State-Of-The-Art and Future DirectionsMagdalena Wischnewski, Nicole C. Krämer, Emmanuel MüllerCHI 2023 · 135 citations
- "Are You Really Sure?" Understanding the Effects of Human Self-Confidence Calibration in AI-Assisted Decision MakingShuai Ma, Xinru Wang, Ying Lei, Chuhan Shi et al.CHI 2024 · 54 citations
Related papers
- Objection Overruled! Lay People can Distinguish Large Language Models from Lawyers, but still Favour Advice from an LLMEike Schneiders, Tina Seabrooke, Joshua Krook, Richard Hyde et al.CHI 2025 · 15 citations
- Generating Automatic Feedback on UI Mockups with Large Language ModelsPeitong Duan, Jeremy Warner, Yang Li, Bjoern HartmannCHI 2024 · 81 citations
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 509 citations
- How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?Ryan Liu, Theodore R. Sumers, Ishita Dasgupta, Thomas L. GriffithsICML 2024 · 33 citations
- "It's the only thing I can trust": Envisioning Large Language Model Use by Autistic Workers for Communication AssistanceJiWoong Jang, Sanika Moharana, Patrick Carrington, Andrew BegelCHI 2024 · 58 citations
