Plan-Then-Execute: An Empirical Study of User Trust and Team Performance When Using LLM Agents As A Daily Assistant
Gaole He, Gianluca Demartini, Ujwal Gadiraju
Abstract
Since the explosion in popularity of ChatGPT, large language models (LLMs) have continued to impact our everyday lives. Equipped with external tools that are designed for a specific purpose (e.g., for flight booking or an alarm clock), LLM agents exercise an increasing capability to assist humans in their daily work. Although LLM agents have shown a promising blueprint as daily assistants, there is a limited understanding of how they can provide daily assistance based on planning and sequential decision making capabilities. We draw inspiration from recent work that has highlighted the value of 'LLM-modulo' setups in conjunction with humans-in-the-loop for planning tasks. We conducted an empirical study (N = 248) of LLM agents as daily assistants in six commonly occurring tasks with different levels of risk typically associated with them (e.g., flight ticket booking and credit card payments). To ensure user agency and control over the LLM agent, we adopted LLM agents in a plan-then-execute manner, wherein the agents conducted step-wise planning and step-by-step execution in a simulation environment. We analyzed how user involvement at each stage affects their trust and collaborative team performance. Our findings demonstrate that LLM agents can be a double-edged sword -- (1) they can work well when a high-quality plan and necessary user involvement in execution are available, and (2) users can easily mistrust the LLM agents with plans that seem plausible. We synthesized key insights for using LLM agents as daily assistants to calibrate user trust and achieve better overall task outcomes. Our work has important implications for the future design of daily assistants and human-AI collaboration with LLM agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42151629-7fbe-4ece-8562-239fbc3439bbCited by top-tier papers9
- Semantic-Aware Logical Reasoning via a Semiotic FrameworkYunyao Zhang, Xinglang Zhang, Junxi Sheng, Wenbing Li et al.ACL 2026 · 27 citations
- Mind the Gap! Choice Independence in Using Multilingual LLMs for Persuasive Co-Writing Tasks in Different LanguagesShreyan Biswas, Alexander Erlei, Ujwal GadirajuCHI 2025 · 11 citations
- Cocoa: Co-Planning and Co-Execution with AI AgentsK. J. Kevin Feng, Kevin Pu, Matt Latzke, Tal August et al.CHI 2026 · 6 citations
- When Workout Buddies Are Virtual: AI Agents and Human Peers in a Longitudinal Physical Activity StudyAlessandro Silacci, Mauro Cherubini, Arianna Boldi, Amon Rapp et al.CHI 2026 · 2 citations
- Who Does What? Archetypes of Roles Assigned to LLMs During Human-AI Decision-MakingShreya Chappidi, Jatinder Singh, Andra Valentina KrauzeCHI 2026 · 2 citations
Builds on32
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-makingZana Buçinca, Maja Barbara Malaya, Krzysztof Z. GajosCSCW 2021 · 962 citations
- On the Planning Abilities of Large Language Models - A Critical InvestigationKarthik Valmeekam, Matthew Marquez, Sarath Sreedharan, Subbarao KambhampatiNeurIPS 2023 · 509 citations
- AI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model PromptsTongshuang Wu, Michael Terry, Carrie Jun CaiCHI 2022 · 465 citations
- Explanations Can Reduce Overreliance on AI Systems During Decision-MakingHelena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg et al.CSCW 2023 · 362 citations
Related papers
- How Do Analysts Understand and Verify AI-Assisted Data Analyses?Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang et al.CHI 2024 · 36 citations
- Are Generative AI Agents Effective Personalized Financial Advisors?Takehiro Takayanagi, Kiyoshi Izumi, Javier Sanz-Cruzado, Richard McCreadie et al.SIGIR 2025 · 9 citations
- Can Large Language Model Agents Simulate Human Trust Behavior?Chengxing Xie, Canyu Chen, Feiran Jia, Ziyu Ye et al.NeurIPS 2024 · 183 citations
- Vibe Check: Understanding the Effects of LLM-Based Conversational Agents' Personality and Alignment on User Perceptions in Goal-Oriented TasksHasibur Rahman, Smit DesaiCHI 2026 · 1 citation
- From Text to Trust: Empowering AI-assisted Decision Making with Adaptive LLM-powered AnalysisZhuoyan Li, Hangxiao Zhu, Zhuoran Lu, Ziang Xiao et al.CHI 2025 · 30 citations
