Pool of Experts: Realtime Querying Specialized Knowledge in Massive Neural Networks
Hakbin Kim, Dong-Wan Choi
Abstract
In spite of the great success of deep learning technologies, training and delivery of a practically serviceable model is still a highly time-consuming process. Furthermore, a resulting model is usually too generic and heavyweight, and hence essentially goes through another expensive model compression phase to fit in a resource-limited device like embedded systems. Inspired by the fact that a machine learning task specifically requested by mobile users is often much simpler than it is supported by a massive generic model, this paper proposes a framework, called Pool of Experts (PoE), that instantly builds a lightweight and task-specific model without any training process. For a realtime model querying service, PoE first extracts a pool of primitive components, called experts, from a well-trained and sufficiently generic network by exploiting a novel conditional knowledge distillation method, and then performs our train-free knowledge consolidation to quickly combine necessary experts into a lightweight network for a target task. Thanks to this train-free property, in our thorough empirical study, PoE can build a fairly accurate yet compact model in a realtime manner, whereas it takes a few minutes per query for the other training methods to achieve a similar level of the accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on3
- Online Model Distillation for Efficient Video InferenceRavi Teja Mullapudi, Steven Chen, Keyi Zhang, Deva Ramanan et al.ICCV 2019 · 131 citations
- BlazeIt: Optimizing Declarative Aggregation and Limit Queries for Neural Network-Based Video AnalyticsDaniel Kang, Peter Bailis, Matei ZahariaVLDB 2020 · 103 citations
- Video Monitoring QueriesNick Koudas, Raymond Li, Ioannis XarchakosICDE 2020 · 33 citations
Related papers
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
- Collaboration of Experts: Achieving 80% Top-1 Accuracy on ImageNet with 100M FLOPsYikang Zhang, Zhuo Chen, Zhao ZhongICML 2022 · 11 citations
- D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM ServingHaodong Wang, Qihua Zhou, Zicong Hong, Song GuoMobiCom 2025 · 8 citations
- LightTS: Lightweight Time Series Classification with Adaptive Ensemble DistillationDavid Campos, Miao Zhang, Bin Yang, Tung Kieu et al.SIGMOD 2023 · 105 citations
- Mixture of Lookup ExpertsShibo Jie, Yehui Tang, Kai Han, Yitong Li et al.ICML 2025
