EdgeFormer: Latency-Aware Collaborative Multi-Head Attention of Transformer Inference in Edge Networks
Yiming Yao, Jianwei Niu, Bin Dai, Tao Ren
Abstract
Recent breakthroughs in Transformer-based large models, have driven widespread tasks, yet their reliance on centralized cloud deployment raises significant privacy risks due to sensitive data exposure. While edgebased collaborative inference offers a privacypreserving alternative, existing methods face critical limitations: static model partitioning cannot adapt to dynamic edge resource fluctuations, and rigid multi-head attention handling overlooks semantic-critical prioritization and parallelism. We propose EdgeFormer, a latency-aware framework for distributed Transformer inference in resource-constrained edge networks. EdgeFormer dynamically allocates model blocks across devices via efficiencystorage trade-off optimization and introduces collaborative Multi-Head Attention (cMHA), which distributes semantic-critical attention heads across devices while pruning redundant ones under real-time constraints. We further develop LiScore, a composite metric integrating attention diversity and latency costs, alongside a similarity-based retrieval method to reduce recomputation overhead. Extensive experiments demonstrate that EdgeFormer achieves up to 2.01× inference acceleration over state-of-theart baselines with ≤1.06% accuracy loss, maintaining robustness under varying edge conditions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cd9ad99b-1da1-4a56-b2c1-939551dcbd0eBuilds on8
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Privacy Risks of General-Purpose Language ModelsXudong Pan, Mi Zhang, Shouling Ji, Min YangS&P 2020 · 291 citations
- Structured Pruning Learns Compact and Accurate ModelsMengzhou Xia, Zexuan Zhong, Danqi ChenACL 2022 · 236 citations
- LLM-Pilot: Characterize and Optimize Performance of your LLM Inference ServicesMalgorzata Lazuka, Andreea Anghel, Thomas P. ParnellSC 2024 · 17 citations
- Online Resource Allocation for Edge Intelligence with Colocated Model Retraining and InferenceHuaiguang Cai, Zhi Zhou, Qianyi HuangINFOCOM 2024 · 10 citations
Related papers
- HCInfer: Hierarchical Coordination for Real-Time Collaborative Inference of LLM on the EdgeKaiyuan Liu, Lizi Zhang, Chengzhong Xu, Li LiRTSS 2025 · 1 citation
- HALO: Semantic-Aware Distributed LLM Inference in Lossy Edge NetworkPeirong Zheng, Wenchao Xu, Haozhao Wang, Jinyu Chen et al.INFOCOM 2026 · 2 citations
- Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer InferenceShengyuan Ye, Jiangsu Du, Liekang Zeng, Wenzhong Ou et al.INFOCOM 2024 · 43 citations
- EdgeSpec: Distributed Speculative Decoding for Large Language Models at EdgeYulin Chen, Meng Tian, Chao Qiu, Xiaofei Wang et al.INFOCOM 2026
- PCFormer: Accelerating Privacy-preserving Transformer Inference by Partition and CombinationBo Zeng, Zhi Pang, Yuyang Zhang, Kai Zhao et al.AAAI 2026
