Characterizing Agents in Production
Melissa Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi
Abstract
LLM-based agents already operate in production across many industries, yet we lack a clear understanding of which technical methods make these deployments successful. We present the first systematic study of Characterizing Agents in Production (CAP) using first-hand data from agent developers. We conducted 20 in-depth case studies through interviews and surveyed 306 practitioners across 26 domains. We examine why organizations build agents, how they build them, how they evaluate them, and the key challenges they face in deployment. Our findings show that production agents rely on simple, controllable approaches: 68% execute at most 10 steps before human intervention, 70% rely on prompting off-the-shelf models rather than weight tuning, and 74% depend primarily on human evaluation. Reliability—defined as consistent correct behavior over time—emerges as the dominant challenge, which practitioners address through system-level design choices. CAP documents the current state of production agents, providing the research community with visibility into real-world deployment practices and underexplored research opportunities.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34c7c7b8-563b-41d6-bb4b-8938f0933fb3Builds on5
- SWE-agent: Agent-Computer Interfaces Enable Automated Software EngineeringJohn Yang, Carlos E. Jimenez, Alexander Wettig, Kilian Lieret et al.NeurIPS 2024 · 2,059 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
- Semantic Operators and Their Optimization: Towards AI-Based Data Analytics with Accuracy GuaranteesLiana Patel, Siddharth Jha, Melissa Z. Pan, Harshit Gupta et al.VLDB 2025 · 16 citations
- PlanGEN: A Multi-Agent Framework for Generating Planning and Reasoning Trajectories for Complex Problem SolvingMihir Parmar, Xin Liu, Palash Goyal, Yanfei Chen et al.EMNLP 2025 · 4 citations
- EnIGMA: Interactive Tools Substantially Assist LM Agents in Finding Security VulnerabilitiesTalor Abramovich, Meet Udeshi, Minghao Shao, Kilian Lieret et al.ICML 2025
Related papers
- DiLLS: Interactive Diagnosis of LLM-based Multi-agent Systems via Layered Summary of Agent BehaviorsRui Sheng, Yukun Yang, Chuhan Shi, Yanna Lin et al.CHI 2026 · 2 citations
- Holistic Agent Leaderboard: The Missing Infrastructure for AI Agent EvaluationSayash Kapoor, Benedikt Stroebl, Peter Kirgis, Nitya Nadgir et al.ICLR 2026 · 86 citations
- Speaker Verification in Agent-generated ConversationsYizhe Yang, Palakorn Achananuparp, Heyan Huang, Jing Jiang et al.ACL 2024
- Are Your Agents Upward Deceivers?Dadi Guo, Qingyu Liu, Dongrui Liu, Qihan Ren et al.ICML 2026 · 5 citations
- ReCreate: Reasoning and Creating Domain Agents Driven by ExperienceZhezheng Hao, Hong Wang, Jian Luo, Jianqing Zhang et al.ACL 2026 · 16 citations
