SAGE: A Dataflow-Native Framework for Modular, Controllable, and Transparent LLM-Augmented Reasoning
Jun Liu, Peilin Liu, Ruicheng Zhang, Senlei Zhang, Yanbo Chen, Ziao Wang, Jinyun Yang, mingqi wang, Shuhao Zhang, Xiaofei Liao, Hai Jin
Abstract
Large Language Model (LLM) applications increasingly execute as end-to-end inference pipelines that couple generation with retrieval, stateful memory, context refinement, and tool use under strict tail-latency and Service-Level Objective (SLO) constraints. Today, these stages are often stitched together as RPC-connected services, obscuring cross-stage queueing and interference and limiting pipeline-level compilation and resource sharing. We present SAGE (Streaming-Augmented Generative Execution), a full-stack system that treats inference pipelines as first-class compilation targets. SAGE exposes pipelines as declarative dataflows and compiles them into distributed execution plans with bounded-queue backpressure. It integrates vector search, streaming semantic state, structured memory, and refinement as operators with explicit resource/state contracts, enabling operator-level diagnosis of tail behavior. SAGE integrates pluggable generation and embedding backends and provides a unified control plane for engine management, batching, and admission under mixed workloads. On a 16-node cluster, SAGE sustains 16 requests/s at tokens/request with 1 ms median scheduling overhead, and achieves near-linear scale-out to 16 nodes (11.4 throughput at 16 nodes), and reduces P99 latency by 57% under multi-pipeline contention versus simultaneous admission.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-ServeAmey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan et al.OSDI 2024 · 537 citations
- MLPerf Inference BenchmarkVijay Janapa Reddi, Christine Cheng, David Kanter, Peter Mattson et al.ISCA 2020 · 517 citations
- LLMLingua: Compressing Prompts for Accelerated Inference of Large Language ModelsHuiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang et al.EMNLP 2023 · 94 citations
Related papers
- HEXGEN-FLOW: Optimizing LLM Inference Request Scheduling for Agentic Text-to-SQLYou Peng, Youhe Jiang, Wenqi Jiang, Chen Wang et al.ICDE 2026
- SGLang: Efficient Execution of Structured Language Model ProgramsLianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun et al.NeurIPS 2024 · 1,586 citations
- Aqua: Network-Accelerated Memory Offloading for LLMs in Scale-Up GPU DomainsAbhishek Vijaya Kumar, Gianni Antichi, Rachee SinghASPLOS 2025 · 4 citations
- Splitwise: Efficient Generative LLM Inference Using Phase SplittingPratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah et al.ISCA 2024 · 282 citations
- PipeInfer: Accelerating LLM Inference using Asynchronous Pipelined SpeculationBranden Butler, Sixing Yu, Arya Mazaheri, Ali JannesariSC 2024 · 12 citations
