EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge Networks
Yiming Yao, Jianwei Niu, Bin Dai, Tao Ren
摘要
Recent breakthroughs in Transformer-based large models, have driven widespread tasks, yet their reliance on centralized cloud deployment raises significant privacy risks due to sensitive data exposure. While edgebased collaborative inference offers a privacypreserving alternative, existing methods face critical limitations: static model partitioning cannot adapt to dynamic edge resource fluctuations, and rigid multi-head attention handling overlooks semantic-critical prioritization and parallelism. We propose EdgeFormer, a latency-aware framework for distributed Transformer inference in resource-constrained edge networks. EdgeFormer dynamically allocates model blocks across devices via efficiencystorage trade-off optimization and introduces collaborative Multi-Head Attention (cMHA), which distributes semantic-critical attention heads across devices while pruning redundant ones under real-time constraints. We further develop LiScore, a composite metric integrating attention diversity and latency costs, alongside a similarity-based retrieval method to reduce recomputation overhead. Extensive experiments demonstrate that EdgeFormer achieves up to 2.01× inference acceleration over state-of-theart baselines with ≤1.06% accuracy loss, maintaining robustness under varying edge conditions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 被引用 291 次
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 被引用 236 次
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 被引用 17 次
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 被引用 10 次
相关 Paper
- HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeKaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li LiRTSS 2025 · 被引用 1 次
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen 等INFOCOM 2026 · 被引用 2 次
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou 等INFOCOM 2024 · 被引用 43 次
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang 等INFOCOM 2026
- PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and CombinationBo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao 等AAAI 2026
