Impatient Users Confuse AI Agents: High-fidelity Simulations of Human Traits for Testing Agents
Muyu He, Anand Kumar, Soumyadeep Bakshi, James Zou, Nazneen Rajani
Abstract
Despite rapid progress in building conversational AI agents, robustness is still largely untested. Small shifts in user behavior, such as being more impatient, incoherent, or skeptical, can cause sharp drops in agent performance, revealing how brittle current AI agents are. Today's benchmarks fail to capture this fragility: agents may perform well under standard evaluations but degrade spectacularly in more realistic and varied settings. We address this robustness testing gap by introducing TraitBasis, a lightweight, model-agnostic method for systematically stress testing AI agents. TraitBasis learns directions in activation space corresponding to steerable user traits (e.g., impatience or incoherence), which can be controlled, scaled, composed, and applied at inference time without any fine-tuning or extra data. Using TraitBasis, we extend τ -Bench to τ -Trait, where user behaviors are altered via controlled trait vectors. We observe an average 4%-20% performance degradation on τ -Trait across frontier models, highlighting the lack of robustness of current AI agents to variations in user behavior. Together, these results highlight both the critical role of robustness testing and the promise of TraitBasis as a simple, data-efficient, and compositional tool. By powering simulation-driven stress tests and training loops, TraitBasis opens the door to building AI agents that remain reliable in the unpredictable dynamics of real-world human interactions. We plan to open-source τ -Trait across four domains: airline, retail, telecom, and telehealth, so the community can systematically QA their agents under realistic, behaviorally diverse intents and trait scenarios. We have open-sourced τ -Trait across four domains: airline, retail, telecom, and telehealth, so the community can systematically QA their agents under realistic, behaviorally diverse intents and trait scenarios: https://github.com/collinear-ai/tau-trait .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic EvaluationsPreethi Seshadri, Samuel Cahyawijaya, Ayomide Odumakinde, Sameer Singh et al.ACL 2026 · 19 citations
- InfoPO: Information-Driven Policy Optimization for User-Centric AgentsFanqi Kong, Jiayi Zhang, Mingyi Deng, Chenglin Wu et al.ICML 2026
Builds on10
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 244 citations
- MCP-Bench: Benchmarking Tool-Using LLM Agents with Complex Real-World Tasks via MCP ServersZhenting Wang, Qi Chang, Hemani Patel, Shashank Biju et al.ICLR 2026 · 109 citations
- Quantifying the Persona Effect in LLM SimulationsTiancheng Hu, Nigel CollierACL 2024 · 22 citations
- Democratizing Large Language Models via Personalized Parameter-Efficient Fine-tuningZhaoxuan Tan, Qingkai Zeng, Yijun Tian, Zheyuan Liu et al.EMNLP 2024 · 17 citations
Related papers
- Non-Collaborative User Simulators for Tool AgentsJeonghoon Shim, Woojung Song, Cheyon Jin, Seungwon KooK et al.ICLR 2026 · 17 citations
- AgentInspect: Diagnosing Behavioral Failures in Artificial Intelligence AgentsRuchira Manke, Mohammad Wardat, Foutse Khomh, Hridesh RajanISSTA 2026
- τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World DomainsShunyu Yao, Noah Shinn, Pedram Razavi, Karthik R. NarasimhanICLR 2025
- -Knowledge: Evaluating Conversational Agents over Unstructured KnowledgeQuan Shi, Alexandra Zytek, Pedram Razavi, Karthik Narasimhan et al.ICML 2026 · 15 citations
- TAI3: Testing Agent Integrity in Interpreting User IntentShiwei Feng, Xiangzhe Xu, Xuan Chen, Kaiyuan Zhang et al.NeurIPS 2025 · 9 citations
