Efficient LLM Serving for Agentic Workflows: A Data Systems Perspective
Noppanat Wadlom, Junyi Shen, Yao Lu
Abstract
Agentic workflows are composed of sequences of interdependent Large Language Model (LLM) calls, and they have become a dominant workload in modern AI systems. These workflows exhibit extensive redundancy from overlapping prompts and intermediate results due to speculative and parallel exploration. Existing LLM serving systems, such as vLLM, focus on optimizing individual inference calls and overlook cross-call dependencies, leading to significant inefficiencies. This paper rethinks LLM and agent serving from a data systems perspective and introduces Helium, a workflow-aware serving framework that models agentic workloads as query plans and treats LLM invocations as first-class operators. Helium integrates proactive caching and cache-aware scheduling to maximize reuse across prompts, KV states, and workflows. Through these techniques, Helium bridges classic query optimization principles with LLM serving, achieving up to 1.56× speedup over state-of-the-art agent serving systems on various workloads. Our results demonstrate that end-to-end optimization across workflows is essential for scalable and efficient LLM-based agents.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4a15e1eb-7a27-4f8e-863b-01dc720fa6f5Builds on36
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Generative Agents: Interactive Simulacra of Human BehaviorJoon Sung Park, Joseph C. O'Brien, Carrie Jun Cai, Meredith Ringel Morris et al.UIST 2023 · 1,882 citations
Related papers
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu et al.NeurIPS 2025 · 77 citations
- HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End OptimizationSize Li, Zhiqing Tang, Hongrui Liang, Jianxiong Guo et al.ICML 2026
- ThunderAgent: A Fast, Simple, and Program-Aware Agentic Inference SystemHao Kang, Ziyang Li, Xinyu Yang, Weili Xu et al.ICML 2026 · 14 citations
- Agentix: An Efficient Serving Engine for LLM Agents as General ProgramsMichael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang et al.NSDI 2026 · 27 citations
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
