CompOFA - Compound Once-For-All Networks for Faster Multi-Platform Deployment
Manas Sahni, Shreya Varshini, Alind Khare, Alexey Tumanov
Abstract
The emergence of CNNs in mainstream deployment has necessitated methods to design and train efficient architectures tailored to maximize the accuracy under diverse hardware & latency constraints. To scale these resource-intensive tasks with an increasing number of deployment targets, Once-For-All (OFA) proposed an approach to jointly train several models at once with a constant training cost. However, this cost remains as high as 40-50 GPU days and also suffers from a combinatorial explosion of sub-optimal model configurations. We seek to reduce this search space -and hence the training budget -by constraining search to models close to the accuracy-latency Pareto frontier. We incorporate insights of compound relationships between model dimensions to build CompOFA, a design space smaller by several orders of magnitude. Through experiments on ImageNet, we demonstrate that even with simple heuristics we can achieve a 2x reduction in training time 1 and 216x speedup in model search/extraction time compared to the state of the art, without loss of Pareto optimality! We also show that this smaller design space is dense enough to support equally accurate models for a similar diversity of hardware and latency targets, while also reducing the complexity of the training and subsequent extraction algorithms. 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 50acec48-bb7c-4620-bcea-98914d97f587Cited by top-tier papers9
- Stimulative Training of Residual Networks: A Social Psychology Perspective of LoafingPeng Ye, Shengji Tang, Baopu Li, Tao Chen et al.NeurIPS 2022 · 15 citations
- Boosting Residual Networks with Group KnowledgeShengji Tang, Peng Ye, Baopu Li, Weihao Lin et al.AAAI 2024 · 7 citations
- KLAS: Using Similarity to Stitch Neural Networks for Improved Accuracy-Efficiency TradeoffsDebopam Sanyal, Anantharaman S. Iyer, Alind Khare, Trisha Jain et al.ICLR 2026 · 2 citations
- VillainNet: Targeted Poisoning Attacks Against SuperNets Along the Accuracy-Latency Pareto FrontierDavid Oygenblik, Abhinav Vemulapalli, Animesh Agrawal, Debopam Sanyal et al.CCS 2025
- Bespoke: A Block-Level Neural Network Optimization Framework for Low-Cost DeploymentJong-Ryul Lee, Yong-Hyuk MoonAAAI 2023
Builds on7
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang et al.ICLR 2020 · 1,522 citations
- Universally Slimmable Networks and Improved Training TechniquesJiahui Yu, Thomas S. HuangICCV 2019 · 444 citations
- On Network Design Spaces for Visual RecognitionIlija Radosavovic, Justin Johnson, Saining Xie, Wan-Yen Lo et al.ICCV 2019 · 148 citations
- AOWS: Adaptive and Optimal Network Width Search With Latency ConstraintsMaxim Berman, Leonid Pishchulin, Ning Xu, Matthew B. Blaschko et al.CVPR 2020
Related papers
- SuperFast: Fast Supernet Training Using Initial KnowledgeMoritz Thoma, Emad Aghajanzadeh, Shambhavi Balamuthu Sampath, Pierpaolo Morì et al.DAC 2025
- APQ: Joint Search for Network Architecture, Pruning and Quantization PolicyTianzhe Wang, Kuan Wang, Han Cai, Ji Lin et al.CVPR 2020
- Best of Both Worlds: AutoML Codesign of a CNN and its Hardware AcceleratorMohamed S. Abdelfattah, Lukasz Dudziak, Thomas Chau, Royson Lee et al.DAC 2020 · 78 citations
- Searching for Fast Model Families on Datacenter AcceleratorsSheng Li, Mingxing Tan, Ruoming Pang, Andrew Li et al.CVPR 2021
- OPA: One-Predict-All For Efficient DeploymentJunpeng Guo, Shengqing Xia, Chunyi PengINFOCOM 2023 · 3 citations
