StackGenVis: Alignment of Data, Algorithms, and Models for Stacking Ensemble Learning Using Performance Metrics
Angelos Chatzimparmpas, Rafael Messias Martins, Kostiantyn Kucher, Andreas Kerren
Abstract
In machine learning (ML), ensemble methods-such as bagging, boosting, and stacking-are widely-established approaches that regularly achieve top-notch predictive performance. Stacking (also called "stacked generalization") is an ensemble method that combines heterogeneous base models, arranged in at least one layer, and then employs another metamodel to summarize the predictions of those models. Although it may be a highly-effective approach for increasing the predictive performance of ML, generating a stack of models from scratch can be a cumbersome trial-and-error process. This challenge stems from the enormous space of available solutions, with different sets of data instances and features that could be used for training, several algorithms to choose from, and instantiations of these algorithms using diverse parameters (i.e., models) that perform differently according to various metrics. In this work, we present a knowledge generation model, which supports ensemble learning with the use of visualization, and a visual analytics system for stacked generalization. Our system, StackGenVis, assists users in dynamically adapting performance metrics, managing data instances, selecting the most important features for a given data set, choosing a set of top-performant and diverse algorithms, and measuring the predictive performance. In consequence, our proposed tool helps users to decide between distinct models and to reduce the complexity of the resulting stack by removing overpromising and underperforming models. The applicability and effectiveness of StackGenVis are demonstrated with two use cases: a real-world healthcare data set and a collection of data related to sentiment/stance detection in texts. Finally, the tool has been evaluated through interviews with three ML experts.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4144f987-cb81-435e-9a1c-e3c5c8ffddd5Cited by top-tier papers1
Ask how each one uses itRelated papers
- Theoretical Guarantees of Learning Ensembling Strategies with Applications to Time Series ForecastingHilaf Hasson, Danielle C. Maddix, Bernie Wang, Gaurav Gupta et al.ICML 2023 · 4 citations
- Towards Understanding Fairness and its Composition in Ensemble Machine LearningUsman Gohar, Sumon Biswas, Hridesh RajanICSE 2023 · 30 citations
- MAAT: a novel ensemble approach to addressing fairness and performance bugs for machine learning softwareZhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanFSE 2022 · 65 citations
- PSEO: Optimizing Post-hoc Stacking Ensemble Through Hyperparameter TuningBeicheng Xu, Wei Liu, Keyao Ding, Yupeng Lu et al.AAAI 2026 · 2 citations
- United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat OverfitUri Stern, Daniel Shwartz, Daphna WeinshallAAAI 2024
