Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPs
Yikang Zhang, Zhuo Chen, Zhao Zhong
Abstract
In this paper, we propose a Collaboration of Experts (CoE) framework to assemble the expertise of multiple networks towards a common goal. Each expert is an individual network with expertise on a unique portion of the dataset, contributing to the collective capacity. Given a sample, delegator selects an expert and simultaneously outputs a rough prediction to trigger potential early termination. For each model in CoE, we propose a novel training algorithm with two major components: weight generation module (WGM) and label generation module (LGM). It fulfills the co-adaptation of experts and delegator. WGM partitions the training data into portions based on delegator via solving a balanced transportation problem, then impels each expert to focus on one portion by reweighting the losses. LGM generates the label to constitute the loss of delegator for expert selection. CoE achieves the state-of-the-art performance on ImageNet, 80.7% top-1 accuracy with 194M FLOPs. Combined with PWLU and CondConv, CoE further boosts the accuracy to 80.0% with only 100M FLOPs for the first time. Furthermore, experiment results on the translation task also demonstrate the strong generalizability of CoE. CoE is hardware-friendly, yielding a 3∼6x acceleration compared with existing conditional computation approaches.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 1,298 citations
- Boosted Dynamic Neural NetworksHaichao Yu, Haoxiang Li, Gang Hua, Gao Huang et al.AAAI 2023 · 16 citations
- Layer Compression of Deep Networks with Straight FlowsChengyue Gong, Xiaocong Du, Bhargav Bhushanam, Lemeng Wu et al.AAAI 2024 · 1 citation
Builds on16
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 854 citations
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 263 citations
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 253 citations
Related papers
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural NetworksHakbin Kim, Dong-Wan ChoiSIGMOD 2021 · 1 citation
- Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet ArchitecturesShuqing Luo, Ye Han, Pingzhi Li, Jiayin Qin et al.NeurIPS 2025 · 3 citations
- FLAME: Fully Leveraging MoE Sparsity for Transformer on FPGAXuanda Lin, Huinan Tian, Wenxiao Xue, Lanqi Ma et al.DAC 2024 · 9 citations
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye et al.NeurIPS 2022 · 162 citations
- Cooperation of Experts: Fusing Heterogeneous Information with Large MarginShuo Wang, Shunyang Huang, Jinghui Yuan, Zhixiang Shen et al.ICML 2025
