HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the Edge
Kaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li Li
Abstract
Deploying large language models (LLMs) on edge devices enables real-time responses while preserving user privacy. However, constrained memory and compute resources pose significant challenges for high-quality, single-device inference. To address this, we propose HCInfer, a hierarchical coordination framework for collaborative LLM inference across edge devices. By leveraging idle neighboring devices, HCInfer alleviates performance bottlenecks typical in isolated deployments. HCInfer employs a two-level coordination strategy. At the inter-device level, it leverages idle neighboring devices to collaboratively process attention computations, significantly reducing synchronization overhead. At the intra-device level, it applies finegrained memory and compute optimizations to fully exploit local hardware capabilities. Building on this architecture, HCInfer integrates three key components: (1) Asymmetric Transformer decomposition decouples attention and FFN computation, enabling selective and parallel execution across devices. (2) Layer-wise subdeadline scheduling dynamically profiles execution latency and adapts precision or structure to meet real-time constraints (3) An Overhead Mitigation Module efficiently manages on-device resource usage to support scalability without overwhelming hardware. We evaluate HCInfer on PC, smart home, and mobile platforms using OPT-13B, Qwen2.5-14B, and Llama2-13B models. Experiments show HCInfer achieves 1.67× to 4.3× speedup in TTFT and 1.16× to 17.15× speedup in TPOT compared to existing baselines, maintaining a sub-deadline miss rate (SubDMR) of 15.3% under worst-case conditions while keeping model accuracy degradation within 8% for typical cases and up to 11% in extreme scenarios. These results demonstrate HCInfer's potential to enable efficient and responsive LLM inference in real-world edge environments.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c810289d-2269-44ad-b65b-4f23dfbca4c5Related papers
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen et al.INFOCOM 2026 · 2 citations
- EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge NetworksYiming Yao, Jianwei Niu, Bin Dai, Tao RenACL 2026
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang et al.INFOCOM 2026
- Multi-Tier Multi-Node Scheduling of LLM for Collaborative AI ComputingMulei Ma, Chenyu Gong, Liekang Zeng, Yang YangINFOCOM 2025 · 12 citations
- DynamicInfer: Runtime-Aware Sparse Offloading for LLMs Inference on a Consumer-Grade GPUZhui Zhu, Weichen Zhang, Zhenghan Zhou, Yunhao Liu et al.ICLR 2026
