Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
Shengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian, Xiaowen Chu, Jian Tang, Xu Chen
摘要
Generative large language models (LLMs) have garnered significant attention due to their exceptional capabilities in various AI tasks. Traditionally deployed in cloud datacenters, LLMs are now increasingly moving towards more accessible edge platforms to protect sensitive user data and ensure privacy preservation. The limited computational resources of individual edge devices, however, can result in excessively prolonged in-ference latency and overwhelmed memory usage. While existing research has explored collaborative edge computing to break the resource wall of individual devices, these solutions yet suffer from massive communication overhead and under-utilization of edge resources. Furthermore, they focus exclusively on optimizing the prefill phase, neglecting the crucial autoregressive decoding phase for generative LLMs. To address that, we propose Jupiter, a fast, scalable, and resource-efficient collaborative edge AI system for generative LLM inference. Jupiter introduces a flexible pipelined architecture as a principle and differentiates its system design according to the differentiated characteristics of the prefill and decoding phases. For prefill phase, Jupiter submits a novel intra-sequence pipeline parallelism and develops a meticulous parallelism planning strategy to maximize resource efficiency; For decoding, Jupiter devises an effective outline-based pipeline parallel decoding mechanism combined with speculative decoding, which further magnifies inference acceleration. Extensive evaluation based on realistic implementation demon-strates that Jupiter remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 26.1 × end-to-end latency reduction while rendering on-par generation quality.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home ClustersZonghang Li, Tao Li, Wenjiao Feng, Rongxing Xiao 等ICLR 2026 · 被引用 3 次
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video UnderstandingShengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng 等INFOCOM 2026 · 被引用 2 次
- ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block ActivationChuiyang Meng, Ming Tang, Vincent W. S. WongINFOCOM 2026
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui 等AAAI 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer 等NeurIPS 2023 · 被引用 1,486 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
相关 Paper
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi 等ICML 2026 · 被引用 2 次
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 被引用 20 次
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang 等INFOCOM 2026
- A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language ModelsZuan Xie, Yang Xu, Hongli Xu, Yunming Liao 等INFOCOM 2026 · 被引用 10 次
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou 等INFOCOM 2024 · 被引用 43 次
