TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge Devices
Bowen Zhang, Junyang Zhang, Jiahui Hou, Yixin Wang
Abstract
Recently, large language models (LLMs) have demonstrated powerful capabilities in various domains, including intelligent voice. To preserve privacy while offering better services, it is meaningful to run LLMs on heterogeneous smart home devices for home assistants. Current works predominantly center on the inference capabilities of LLMs using GPU clusters. However, few works study efficient inference on edge devices with constrained computational resources, which presents unique challenges. In this paper, we introduce TensAllo, an innovative system that optimizes the inference of LLMs and ensures reasonable workloads deployed on heterogeneous devices for home assistants. TensAllo employs tensor parallelism for distributed inference and utilizes quantization techniques to reduce memory usage. To reduce inference latency and power consumption, TensAllo intelligently selects devices and determines the optimal allocation of tensors under the constraints of the set power consumption. Extensive experimentation on various heterogeneous edge devices confirms TensAllo's superiority. The results demonstrate that TensAllo increases inference speed by 50 % and 37 %, respectively.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c7aa65fc-ba34-486f-ac33-baf7ec21f4baCited by top-tier papers1
Ask how each one uses itRelated papers
- VCU-LLM: Prompt-efficient On-device Large Language Model for Vague Command Understanding in Smart HomesZhengyuan Zhang, Dong Zhao, Tiancheng He, Zilong Wang et al.UbiComp 2026
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou et al.INFOCOM 2024 · 43 citations
- MELTing Point: Mobile Evaluation of Language TransformersStefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, Hamed HaddadiMobiCom 2024 · 32 citations
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen et al.INFOCOM 2026 · 2 citations
- DynamoLLM: Designing LLM Inference Clusters for Performance and Energy EfficiencyJovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas et al.HPCA 2025 · 106 citations
