Hierarchical Mixture of Experts: Generalizable Learning for High-Level Synthesis
Weikai Li, Ding Wang, Zijian Ding, Atefeh Sohrabizadeh, Zongyue Qin, Jason Cong, Yizhou Sun
Abstract
High-level synthesis (HLS) is a widely used tool in designing Field Programmable Gate Array (FPGA). HLS enables FPGA design with software programming languages by compiling the source code into an FPGA circuit. The source code includes a program (called "kernel") and several pragmas that instruct hardware synthesis, such as parallelization, pipeline, etc. While it is relatively easy for software developers to design the program, it heavily relies on hardware knowledge to design the pragmas, posing a big challenge for software developers. Recently, different machine learning algorithms, such as GNNs, have been proposed to automate the pragma design via performance prediction. However, when applying the trained model on new kernels, the significant domain shift often leads to unsatisfactory performance. We propose a more domain-generalizable model structure: a two-level hierarchical Mixture of Experts (MoE), that can be flexibly adapted to any GNN model. Different expert networks can learn to deal with different regions in the representation space, and they can utilize similar patterns between the old kernels and new kernels. In the low-level MoE, we apply MoE on three natural granularities of a program: node, basic block, and graph. The high-level MoE learns to aggregate the three granularities for the final decision. To train the hierarchical MoE stably, we further propose a two-stage training method to avoid expert polarization. Extensive experiments verify the effectiveness of the proposed hierarchical MoE. We publicized our codes at https://github.com/weikai-li/HierarchicalMoE .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34112dc7-1a9e-4ea3-8907-fbef53181e92Cited by top-tier papers2
- MalMoE: Mixture-of-Experts Enhanced Encrypted Malicious Traffic Detection Under Graph DriftYunpeng Tan, Qingyang Li, Mingxin Yang, Yannan Hu et al.INFOCOM 2026 · 1 citation
- CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesCheng Tang, Guochong Sui, Wenqi Lou, Zihan Wang et al.AAAI 2026
Builds on9
- Unsupervised Domain Adaptive Graph Convolutional NetworksMan Wu, Shirui Pan, Chuan Zhou, Xiaojun Chang et al.WWW 2020 · 221 citations
- ProGraML: A Graph-based Program Representation for Data Flow Analysis and Compiler OptimizationsChris Cummins, Zacharias V. Fisches, Tal Ben-Nun, Torsten Hoefler et al.ICML 2021 · 140 citations
- Graph Mixture of Experts: Learning on Large-Scale Graphs with Explicit Diversity ModelingHaotao Wang, Ziyu Jiang, Yuning You, Yan Han et al.NeurIPS 2023 · 104 citations
- On the Adequacy of Untuned Warmup for Adaptive OptimizationJerry Ma, Denis YaratsAAAI 2021 · 81 citations
- Meta-DMoE: Adapting to Domain Shift by Meta-Distillation from Mixture-of-ExpertsTao Zhong, Zhixiang Chi, Li Gu, Yang Wang et al.NeurIPS 2022 · 70 citations
Related papers
- High-level synthesis performance prediction using GNNs: benchmarking, modeling, and advancingNan Wu, Hang Yang, Yuan Xie, Pan Li et al.DAC 2022 · 57 citations
- Automated accelerator optimization aided by graph neural networksAtefeh Sohrabizadeh, Yunsheng Bai, Yizhou Sun, Jason CongDAC 2022 · 48 citations
- C-GNN-PRUNE: A Unified Graph-Based Framework for Structure-Aware Pruning of Mixture-of-Experts ModelsLin Li, Yan Wang, Zhuopeng WangAAAI 2026 · 1 citation
- NeuroSchedule: A Novel Effective GNN-based Scheduling Method for High-level SynthesisJun Zeng, Mingyang Kou, Hailong YaoNeurIPS 2022 · 5 citations
- Self-Adaptive Graph Mixture of ModelsMohit Meena, Yash Punjabi, Abhishek A, Vishal Sharma et al.AAAI 2026
