Lune

INFOCOM2025Top-tier venue

TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge Devices

Bowen Zhang, Junyang Zhang, Jiahui Hou, Yixin Wang

2025Year
10Citations
1Top-tier citations

Abstract

Recently, large language models (LLMs) have demonstrated powerful capabilities in various domains, including intelligent voice. To preserve privacy while offering better services, it is meaningful to run LLMs on heterogeneous smart home devices for home assistants. Current works predominantly center on the inference capabilities of LLMs using GPU clusters. However, few works study efficient inference on edge devices with constrained computational resources, which presents unique challenges. In this paper, we introduce TensAllo, an innovative system that optimizes the inference of LLMs and ensures reasonable workloads deployed on heterogeneous devices for home assistants. TensAllo employs tensor parallelism for distributed inference and utilizes quantization techniques to reduce memory usage. To reduce inference latency and power consumption, TensAllo intelligently selects devices and determines the optimal allocation of tensors under the constraints of the set power consumption. Extensive experimentation on various heterogeneous edge devices confirms TensAllo's superiority. The results demonstrate that TensAllo increases inference speed by 50 % and 37 %, respectively.

Ask about this paper

Ask your agent about it.

Lune has read the top-tier papers around this one, so every answer names the papers it rests on.

Questions to start from

Your agent calls

Lunesearch_papers

Ask in Lune

Free to start. No credit card required.

lune papers get c7aa65fc-ba34-486f-ac33-baf7ec21f4ba

Cited by top-tier papers1

Ask how each one uses it

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines