Synthetic Model Combination: An Instance-wise Approach to Unsupervised Ensemble Learning
Alex J. Chan, Mihaela van der Schaar
Abstract
Consider making a prediction over new test data without any opportunity to learn from a training set of labelled data - instead given access to a set of expert models and their predictions alongside some limited information about the dataset used to train them. In scenarios from finance to the medical sciences, and even consumer practice, stakeholders have developed models on private data they either cannot, or do not want to, share. Given the value and legislation surrounding personal information, it is not surprising that only the models, and not the data, will be released - the pertinent question becoming: how best to use these models? Previous work has focused on global model selection or ensembling, with the result of a single final model across the feature space. Machine learning models perform notoriously poorly on data outside their training domain however, and so we argue that when ensembling models the weightings for individual instances must reflect their respective domains - in other words models that are more likely to have seen information on that instance should have more attention paid to them. We introduce a method for such an instance-wise ensembling of models, including a novel representation learning step for handling sparse high-dimensional domains. Finally, we demonstrate the need and generalisability of our method on classical machine learning tasks as well as highlighting a real world use case in the pharmacological setting of vancomycin precision dosing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 401f084c-3a76-429b-a16d-857812e36665Cited by top-tier papers1
Ask how each one uses itBuilds on8
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 287 citations
- Uncertainty Quantification and Deep EnsemblesRahul Rahaman, Alexandre H. ThiéryNeurIPS 2021 · 250 citations
- Dangers of Bayesian Model Averaging under Covariate ShiftPavel Izmailov, Patrick Nicholson, Sanae Lotfi, Andrew Gordon WilsonNeurIPS 2021 · 51 citations
- Unlabelled Data Improves Bayesian Uncertainty Calibration under Covariate ShiftAlex J. Chan, Ahmed M. Alaa, Zhaozhi Qian, Mihaela van der SchaarICML 2020 · 42 citations
- POETREE: Interpretable Policy Learning with Adaptive Decision TreesAlizée Pace, Alex J. Chan, Mihaela van der SchaarICLR 2022 · 18 citations
Related papers
- Domain constraints improve risk prediction when outcome data is missingSidhika Balachandar, Nikhil Garg, Emma PiersonICLR 2024 · 11 citations
- Aggregate Models, Not Explanations: Improving Feature Importance EstimationJoseph Paillard, Angel REYERO LOBO, Denis-Alexander Engemann, Thirion BertrandICML 2026 · 1 citation
- Test-time Collective PredictionCelestine Mendler-Dünner, Wenshuo Guo, Stephen Bates, Michael I. JordanNeurIPS 2021 · 4 citations
- WISER: Weak Supervision and Supervised Representation Learning to Improve Drug Response Prediction in CancerKumar Shubham, Aishwarya Jayagopal, Syed Mohammed Danish, Prathosh A. P. et al.ICML 2024 · 8 citations
- Soup-of-Experts: Pretraining Specialist Models via Parameters AveragingPierre Ablin, Angelos Katharopoulos, Skyler Seto, David GrangierICML 2025
