Towards End-to-End Optimization of LLM-based Applications with Ayo
Xin Tan, Yimin Jiang, Yitao Yang, Hong Xu
摘要
Large language model (LLM)-based applications consist of both LLM and non-LLM components, each contributing to the end-to-end latency. Despite great efforts to optimize LLM inference, end-to-end workflow optimization has been overlooked. Existing frameworks employ coarse-grained orchestration with task modules, which confines optimizations to within each module and yields suboptimal scheduling decisions.
We propose fine-grained end-to-end orchestration, which utilizes task primitives as the basic units and represents each query's workflow as a primitive-level dataflow graph. This explicitly exposes a much larger design space, enables optimizations in parallelization and pipelining across primitives of different modules, and enhances scheduling to improve application-level performance. We build Teola, a novel orchestration framework for LLM-based applications that implements this scheme. Comprehensive experiments show that Teola can achieve up to 2.09x speedup over existing systems across various popular LLM applications. The code is available at https://github.com/NetX-lab/Ayo.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- JITServe: SLO-aware LLM Serving with Imprecise Request InformationWei Zhang, Zhiyu Wu, Yi Mu, Rui Ning 等NSDI 2026 · 被引用 29 次
- Murakkab: Resource-Efficient Agentic Workflow Orchestration in Cloud PlatformsGohar Irfan Chaudhry, Esha Choukse, Haoran Qiu, Íñigo Goiri 等OSDI 2026 · 被引用 26 次
- Efficient LLM Serving for Agentic Workflows: A Data Systems PerspectiveNoppanat Wadlom, Junyi Shen, Yao LuSIGMOD 2026 · 被引用 14 次
- ChainBuddy: An AI-assisted Agent System for Generating LLM PipelinesJingyue Zhang, Ian ArawjoCHI 2025 · 被引用 13 次
- The Cost of Dynamic Reasoning: Demystifying AI Agents and Test-Time Scaling from an AI Infrastructure PerspectiveJiin Kim, Byeongjun Shin, Jinha Chung, Minsoo RhuHPCA 2026 · 被引用 7 次
它引用的顶会 Paper23
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging FaceYongliang Shen, Kaitao Song, Xu Tan, Dongsheng Li 等NeurIPS 2023 · 被引用 1,778 次
相关 Paper
- HeraSys: Collaborative Serving of Multiple LLM Workflows via Fine-Grained End-to-End OptimizationSize Li, Zhiqing Tang, Hongrui Liang, Jianxiong Guo 等ICML 2026
- Optimas: An Intelligent Analytics-Informed Generative AI Framework for Performance OptimizationMohammad Zaeed, Tanzima Z. Islam, Vladimir IndicKDD 2026
- Prompting Is Programming: A Query Language for Large Language ModelsLuca Beurer-Kellner, Marc Fischer, Martin T. VechevPLDI 2023 · 被引用 114 次
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang 等ICDE 2026
- BOND: A Co-Designed Framework for LLM-Powered Analytics Over Relational DataLixiang Chen, Qin Zheng, Zhicheng Pan, Chengcheng Yang 等ICDE 2026
