Lune

INFOCOM2025顶会

TensAllo: Adaptive Deployment of LLMs on Resource-Constrained Heterogeneous Edge Devices

Bowen Zhang, Junyang Zhang, Jiahui Hou, Yixin Wang

2025年份
10被引次数
1顶会引用

摘要

Recently, large language models (LLMs) have demonstrated powerful capabilities in various domains, including intelligent voice. To preserve privacy while offering better services, it is meaningful to run LLMs on heterogeneous smart home devices for home assistants. Current works predominantly center on the inference capabilities of LLMs using GPU clusters. However, few works study efficient inference on edge devices with constrained computational resources, which presents unique challenges. In this paper, we introduce TensAllo, an innovative system that optimizes the inference of LLMs and ensures reasonable workloads deployed on heterogeneous devices for home assistants. TensAllo employs tensor parallelism for distributed inference and utilizes quantization techniques to reduce memory usage. To reduce inference latency and power consumption, TensAllo intelligently selects devices and determines the optimal allocation of tensors under the constraints of the set power consumption. Extensive experimentation on various heterogeneous edge devices confirms TensAllo's superiority. The results demonstrate that TensAllo increases inference speed by 50 % and 37 %, respectively.

问问这篇 Paper

问问你的智能体。

Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。

可以从这些问题问起

智能体调用

Lunesearch_papers

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖