DivBO: Diversity-aware CASH for Ensemble Learning
Yu Shen, Yupeng Lu, Yang Li, Yaofeng Tu, Wentao Zhang, Bin Cui
Abstract
The Combined Algorithm Selection and Hyperparameters optimization (CASH) problem is one of the fundamental problems in Automated Machine Learning (AutoML). Motivated by the success of ensemble learning, recent AutoML systems build post-hoc ensembles to output the final predictions instead of using the best single learner. However, while most CASH methods focus on searching for a single learner with the best performance, they neglect the diversity among base learners (i.e., they may suggest similar configurations to previously evaluated ones), which is also a crucial consideration when building an ensemble. To tackle this issue and further enhance the ensemble performance, we propose DivBO, a diversity-aware framework to inject explicit search of diversity into the CASH problems. In the framework, we propose to use a diversity surrogate to predict the pair-wise diversity of two unseen configurations. Furthermore, we introduce a temporary pool and a weighted acquisition function to guide the search of both performance and diversity based on Bayesian optimization. Empirical results on 15 public datasets show that DivBO achieves the best average ranks (1.82 and 1.73) on both validation and test errors among 10 compared methods, including post-hoc designs in recent AutoML systems and state-of-the-art baselines for ensemble learning on CASH problems.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter TuningBeicheng Xu, Wei Liu, Keyao Ding, Yupeng Lu et al.AAAI 2026 · 2 citations
- CAPS: Cost-Aware ML Pipeline SelectionAntonios Kontaxakis, Dimitris Sacharidis, Alberto Abelló, Sergi Nadal et al.VLDB 2026
- CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter OptimizationBeicheng Xu, Keyao Ding, Wei Liu, Yupeng Lu et al.KDD 2026
Builds on7
- Neural Ensemble Search for Uncertainty Estimation and Dataset ShiftSheheryar Zaidi, Arber Zela, Thomas Elsken, Chris C. Holmes et al.NeurIPS 2021 · 97 citations
- An ADMM Based Framework for AutoML Pipeline ConfigurationSijia Liu, Parikshit Ram, Deepak Vijaykeerthy, Djallel Bouneffouf et al.AAAI 2020 · 82 citations
- PaSca: A Graph Neural Architecture Search System under the Scalable ParadigmWentao Zhang, Yu Shen, Zheyu Lin, Yang Li et al.WWW 2022 · 69 citations
- VolcanoML: Speeding up End-to-End AutoML via Scalable Search Space DecompositionYang Li, Yu Shen, Wentao Zhang, Jiawei Jiang et al.VLDB 2021 · 55 citations
- Efficient Automatic CASH via Rising BanditsYang Li, Jiawei Jiang, Jinyang Gao, Yingxia Shao et al.AAAI 2020 · 45 citations
Related papers
- CASH via Optimal Diversity for Ensemble LearningPranav Poduval, Sanjay Kumar Patnala, Gaurav Oberoi, Nitish Srivasatava et al.KDD 2024 · 2 citations
- Put CASH on Bandits: A Max K-Armed Problem for Automated Machine LearningAmir Rezaei Balef, Claire Vernade, Katharina EggenspergerNeurIPS 2025 · 4 citations
- Bayesian Optimization for Simultaneous Selection of Machine Learning Algorithms and Hyperparameters on Shared Latent SpaceKazuki Ishikawa, Ryota Ozaki, Yohei Kanzaki, Ichiro Takeuchi et al.KDD 2025 · 1 citation
- Weighted Sampling for Combined Model Selection and Hyperparameter TuningDimitrios Sarigiannis, Thomas P. Parnell, Haralampos PozidisAAAI 2020 · 3 citations
- Batch Multi-Fidelity Bayesian Optimization with Deep Auto-Regressive NetworksShibo Li, Robert M. Kirby, Shandian ZheNeurIPS 2021 · 14 citations
