Jupiter: Fast and Resource-Efficient Collaborative Inference of Generative LLMs on Edge Devices
Shengyuan Ye, Bei Ouyang, Liekang Zeng, Tianyi Qian, Xiaowen Chu, Jian Tang, Xu Chen
Abstract
Generative large language models (LLMs) have garnered significant attention due to their exceptional capabilities in various AI tasks. Traditionally deployed in cloud datacenters, LLMs are now increasingly moving towards more accessible edge platforms to protect sensitive user data and ensure privacy preservation. The limited computational resources of individual edge devices, however, can result in excessively prolonged in-ference latency and overwhelmed memory usage. While existing research has explored collaborative edge computing to break the resource wall of individual devices, these solutions yet suffer from massive communication overhead and under-utilization of edge resources. Furthermore, they focus exclusively on optimizing the prefill phase, neglecting the crucial autoregressive decoding phase for generative LLMs. To address that, we propose Jupiter, a fast, scalable, and resource-efficient collaborative edge AI system for generative LLM inference. Jupiter introduces a flexible pipelined architecture as a principle and differentiates its system design according to the differentiated characteristics of the prefill and decoding phases. For prefill phase, Jupiter submits a novel intra-sequence pipeline parallelism and develops a meticulous parallelism planning strategy to maximize resource efficiency; For decoding, Jupiter devises an effective outline-based pipeline parallel decoding mechanism combined with speculative decoding, which further magnifies inference acceleration. Extensive evaluation based on realistic implementation demon-strates that Jupiter remarkably outperforms state-of-the-art approaches under various edge environment setups, achieving up to 26.1 × end-to-end latency reduction while rendering on-par generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb9318b2-85ca-4543-a22d-cf113b0c80dbCited by top-tier papers4
- Prima.cpp: Fast 30-70B LLM Inference on Heterogeneous and Low-Resource Home ClustersZonghang Li, Tao Li, Wenjiao Feng, Rongxing Xiao et al.ICLR 2026 · 3 citations
- Venus: An Efficient Edge Memory-and-Retrieval System for VLM-based Online Video UnderstandingShengyuan Ye, Bei Ouyang, Tianyi Qian, Liekang Zeng et al.INFOCOM 2026 · 2 citations
- ZorBA: Zeroth-order Federated Fine-tuning of LLMs with Heterogeneous Block ActivationChuiyang Meng, Ming Tang, Vincent W. S. WongINFOCOM 2026
- DIAA: A Decoding-Efficient Inference Acceleration Approach for On-Device Large Language ModelsHao Tian, Sheng Lu, Fuwen Tian, Guangming Cui et al.AAAI 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
Related papers
- PipeSD: An Efficient Cloud-Edge Collaborative Pipeline Inference Framework with Speculative DecodingHan Yunhe, Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi et al.ICML 2026 · 2 citations
- SpecEdge: Scalable Edge-Assisted Serving Framework for Interactive LLMsJinwoo Park, Seunggeun Cho, Dongsu HanNeurIPS 2025 · 20 citations
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang et al.INFOCOM 2026
- A Novel Hat-Shaped Device-Cloud Collaborative Inference Framework for Large Language ModelsZuan Xie, Yang Xu, Hongli Xu, Yunming Liao et al.INFOCOM 2026 · 10 citations
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou et al.INFOCOM 2024 · 43 citations
