FLASH: Towards a High-performance Hardware Acceleration Architecture for Cross-silo Federated Learning
Junxue Zhang, Xiaodian Cheng, Wei Wang, Liu Yang, Jinbin Hu, Kai Chen
摘要
Cross-silo federated learning (FL) adopts various cryptographic operations to preserve data privacy, which introduces significant performance overhead. In this paper, we identify nine widely-used cryptographic operations and design an efficient hardware architecture to accelerate them. However, directly offloading them on hardware statically leads to (1) inadequate hardware acceleration due to the limited resources allocated to each operation; (2) insufficient resource utilization, since different operations are used at different times. To address these challenges, we propose FLASH, a high-performance hardware acceleration architecture for cross-silo FL systems. At its heart, FLASH extracts two basic operators-modular exponentiation and multiplicationbehind the nine cryptographic operations and implements them as highly-performant engines to achieve adequate acceleration. Furthermore, it leverages a dataflow scheduling scheme to dynamically compose different cryptographic operations based on these basic engines to obtain sufficient resource utilization. We have implemented a fully-functional FLASH prototype with Xilinx VU13P FPGA and integrated it with FATE, the most widely-adopted cross-silo FL framework. Experimental results show that, for the nine cryptographic operations, FLASH achieves up to 14.0× and 3.4× acceleration over CPU and GPU, translating to up to 6.8× and 2.0× speedup for realistic FL applications, respectively. We finally evaluate the FLASH design as an ASIC, and it achieves 23.6× performance improvement upon the FPGA prototype.
Inspired by the above observation, we present FLASH, a highperformance hardware acceleration architecture for cross-silo FL. This section describes how we design FLASH in detail. Please note that our design has been fully implemented in our FLASH prototype with FPGAs as well as rigorously evaluated 6 Appendix B and C provide more details of these operations. Clock Key 1 Key 2 Calculating parameters based on 𝑁 Converting input data into Montgomery form Computation in Montgomery form Converting back from Montgomery form
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Design and Operation of Shared Machine Learning Clusters on CampusKaiqiang Xu, Decang Sun, Hao Wang, Zhenghang Ren 等ASPLOS 2025 · 被引用 24 次
- Efficient Decentralized Federated Singular Vector DecompositionDi Chai, Junxue Zhang, Liu Yang, Yilun Jin 等USENIX ATC 2024 · 被引用 9 次
- Accelerating Privacy-Preserving Machine Learning With GeniBatchXinyang Huang, Junxue Zhang, Xiaodian Cheng, Hong Zhang 等EuroSys 2024 · 被引用 8 次
- Accelerating Secure Collaborative Machine Learning with Protocol-Aware RDMAZhenghang Ren, Mingxuan Fan, Zilong Wang, Junxue Zhang 等USENIX Security 2024 · 被引用 6 次
- Sequoia: An Accessible and Extensible Framework for Privacy-Preserving Machine Learning over Distributed DataKaiqiang Xu, Di Chai, Junxue Zhang, Fan Lai 等SIGMOD 2025 · 被引用 1 次
它引用的顶会 Paper15
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan 等CCS 2016 · 被引用 7,620 次
- SecureML: A System for Scalable Privacy-Preserving Machine LearningPayman Mohassel, Yupeng ZhangS&P 2017 · 被引用 2,107 次
- BatchCrypt: Efficient Homomorphic Encryption for Cross-Silo Federated LearningChengliang Zhang, Suyi Li, Junzhe Xia, Wei Wang 等USENIX ATC 2020 · 被引用 967 次
- F1: A Fast and Programmable Accelerator for Fully Homomorphic EncryptionNikola Samardzic, Axel Feldmann, Aleksandar Krastev, Srinivas Devadas 等MICRO 2021 · 被引用 294 次
- HEAX: An Architecture for Computing on Encrypted DataM. Sadegh Riazi, Kim Laine, Blake Pelton, Wei DaiASPLOS 2020 · 被引用 244 次
相关 Paper
- Flagger: Cooperative Acceleration for Large-Scale Cross-Silo Federated Learning AggregationXiurui Pan, Yuda An, Shengwen Liang, Bo Mao 等ISCA 2024 · 被引用 6 次
- CROPHE: Cross-Operator Dataflow Optimization for Fully Homomorphic Encryption AcceleratorsXinhua Chen, Jiangbin Dong, Hongren Zheng, Tian Tang 等HPCA 2026 · 被引用 1 次
- EFFACT: A Highly Efficient Full-Stack FHE Acceleration PlatformYi Huang, Xinsheng Gong, Xiangyu Kong, Dibei Chen 等HPCA 2025 · 被引用 10 次
- Poseidon: Practical Homomorphic Encryption AcceleratorYinghao Yang, Huaizhi Zhang, Shengyu Fan, Hang Lu 等HPCA 2023 · 被引用 109 次
- Morphling: A Throughput-Maximized TFHE-based Accelerator using Transform-domain ReusePrasetiyo, Adiwena Putra, Joo-Young KimHPCA 2024 · 被引用 22 次
