Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPs
Yikang Zhang, Zhuo Chen, Zhao Zhong
摘要
In this paper, we propose a Collaboration of Experts (CoE) framework to assemble the expertise of multiple networks towards a common goal. Each expert is an individual network with expertise on a unique portion of the dataset, contributing to the collective capacity. Given a sample, delegator selects an expert and simultaneously outputs a rough prediction to trigger potential early termination. For each model in CoE, we propose a novel training algorithm with two major components: weight generation module (WGM) and label generation module (LGM). It fulfills the co-adaptation of experts and delegator. WGM partitions the training data into portions based on delegator via solving a balanced transportation problem, then impels each expert to focus on one portion by reweighting the losses. LGM generates the label to constitute the loss of delegator for expert selection. CoE achieves the state-of-the-art performance on ImageNet, 80.7% top-1 accuracy with 194M FLOPs. Combined with PWLU and CondConv, CoE further boosts the accuracy to 80.0% with only 100M FLOPs for the first time. Furthermore, experiment results on the translation task also demonstrate the strong generalizability of CoE. CoE is hardware-friendly, yielding a 3∼6x acceleration compared with existing conditional computation approaches.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Scaling Up Your Kernels to 31×31: Revisiting Large Kernel Design in CNNsXiaohan Ding, Xiangyu Zhang, Jungong Han, Guiguang DingCVPR 2022 · 被引用 1,298 次
- Boosted Dynamic Neural NetworksHaichao Yu, Haoxiang Li, Gang Hua, Gao Huang 等AAAI 2023 · 被引用 16 次
- Layer Compression of Deep Networks with Straight FlowsChengyue Gong, Xiaocong Du, Bhargav Bhushanam, Lemeng Wu 等AAAI 2024 · 被引用 1 次
它引用的顶会 Paper16
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- S4L: Self-Supervised Semi-Supervised LearningLucas Beyer, Xiaohua Zhai, Avital Oliver, Alexander KolesnikovICCV 2019 · 被引用 854 次
- Hyperparameter Ensembles for Robustness and Uncertainty QuantificationFlorian Wenzel, Jasper Snoek, Dustin Tran, Rodolphe JenattonNeurIPS 2020 · 被引用 263 次
- LogME: Practical Assessment of Pre-trained Models for Transfer LearningKaichao You, Yong Liu, Jianmin Wang, Mingsheng LongICML 2021 · 被引用 253 次
相关 Paper
- Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural NetworksHakbin Kim, Dong-Wan ChoiSIGMOD 2021 · 被引用 1 次
- Mozart: Modularized and Efficient MoE Training on 3.5D Wafer-Scale Chiplet ArchitecturesShuqing Luo, Ye Han, Pingzhi Li, Jiayin Qin 等NeurIPS 2025 · 被引用 3 次
- FLAME: Fully Leveraging MoE Sparsity for Transformer on FPGAXuanda Lin, Huinan Tian, Wenxiao Xue, Lanqi Ma 等DAC 2024 · 被引用 9 次
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye 等NeurIPS 2022 · 被引用 162 次
- Cooperation of Experts: Fusing Heterogeneous Information with Large MarginShuo Wang, Shunyang Huang, Jinghui Yuan, Zhixiang Shen 等ICML 2025
