Parrot: Efficient Serving of LLM-based Applications with Semantic Variable
Chaofan Lin, Zhenhua Han, Chengruidong Zhang, Yuqing Yang, Fan Yang, Chen Chen, Lili Qiu
摘要
The rise of large language models (LLMs) has enabled LLM-based applications (a.k.a. AI agents or co-pilots), a new software paradigm that combines the strength of LLM and conventional software. Diverse LLM applications from different tenants could design complex workflows using multiple LLM requests to accomplish one task. However, they have to use the over-simplified request-level API provided by today's public LLM services, losing essential application-level information. Public LLM services have to blindly optimize individual LLM requests, leading to sub-optimal end-to-end performance of LLM applications. This paper introduces Parrot, an LLM service system that focuses on the end-to-end experience of LLM-based applications. Parrot proposes Semantic Variable, a unified abstraction to expose application-level knowledge to public LLM services. A Semantic Variable annotates an input/output variable in the prompt of a request, and creates the data pipeline when connecting multiple LLM requests, providing a natural way to program LLM applications. Exposing Semantic Variables to the public LLM service allows it to perform conventional data flow analysis to uncover the correlation across multiple LLM requests. This correlation opens a brand-new optimization space for the end-to-end performance of LLMbased applications. Extensive evaluations demonstrate that Parrot can achieve up to an order-of-magnitude improvement for popular and practical use cases of LLM applications. * This work is partially done while Chaofan Lin's internship and Dr. Chen Chen's visting scholar in Microsoft Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper38
- KVFlow: Efficient Prefix Caching for Accelerating LLM-Based Multi-Agent WorkflowsZaifeng Pan, Ajjkumar Patel, Yipeng Shen, Zhengding Hu 等NeurIPS 2025 · 被引用 77 次
- KVCache Cache in the Wild: Characterizing and Optimizing KVCache Cache at a Large Cloud ProviderJiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen 等USENIX ATC 2025 · 被引用 70 次
- CacheBlend: Fast Large Language Model Serving for RAG with Cached Knowledge FusionJiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray 等EuroSys 2025 · 被引用 68 次
- Twilight: Adaptive Attention Sparsity with Hierarchical Top- PruningChaofan Lin, Jiaming Tang, Shuo Yang, Hanshuo Wang 等NeurIPS 2025 · 被引用 53 次
- Efficiently Scaling LLM Reasoning Programs with CertaindexYichao Fu, Junda Chen, Siqi Zhu, Zheyu Fu 等NeurIPS 2025 · 被引用 42 次
它引用的顶会 Paper16
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Orca: A Distributed Serving System for Transformer-Based Generative ModelsGyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim 等OSDI 2022 · 被引用 690 次
相关 Paper
- Agentix: An Efficient Serving Engine for LLM Agents as General ProgramsMichael Luo, Xiaoxiang Shi, Colin Cai, Tianjun Zhang 等NSDI 2026 · 被引用 27 次
- Parrot: Enhancing Multi-Turn Instruction Following for Large Language ModelsYuchong Sun, Che Liu, Kun Zhou, Jinwen Huang 等ACL 2024
- AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT ApplicationsLeming Shen, Qiang Yang, Yuanqing Zheng, Mo LiMobiCom 2025 · 被引用 14 次
- From Commands to Prompts: LLM-based Semantic File System for AIOSZeru Shi, Kai Mei, Mingyu Jin, Yongye Su 等ICLR 2025
- AXIS: Efficient Human-Agent-Computer Interaction with API-First LLM-Based AgentsJunting Lu, Zhiyang Zhang, Fangkai Yang, Jue Zhang 等ACL 2025 · 被引用 9 次
