Fast, Accurate, and Simple Models for Tabular Data via Augmented Distillation
Rasool Fakoor, Jonas Mueller, Nick Erickson, Pratik Chaudhari, Alexander J. Smola
摘要
Automated machine learning (AutoML) can produce complex model ensembles by stacking, bagging, and boosting many individual models like trees, deep networks, and nearest neighbor estimators. While highly accurate, the resulting predictors are large, slow, and opaque as compared to their constituents. To improve the deployment of AutoML on tabular data, we propose FAST-DAD to distill arbitrarily complex ensemble predictors into individual models like boosted trees, random forests, and deep networks. At the heart of our approach is a data augmentation strategy based on Gibbs sampling from a self-attention pseudolikelihood estimator. Across 30 datasets spanning regression and binary/multiclass classification tasks, FAST-DAD distillation produces significantly better individual models than one obtains through standard training on the original data. Our individual distilled models are over 10x faster and more accurate than ensemble predictors produced by AutoML tools like H2O/AutoSklearn.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Does Knowledge Distillation Really Work?Samuel Stanton, Pavel Izmailov, Polina Kirichenko, Alexander A. Alemi 等NeurIPS 2021 · 被引用 318 次
- Gradient-Free Structured Pruning with Unlabeled DataAzade Nova, Hanjun Dai, Dale SchuurmansICML 2023 · 被引用 38 次
- Reinforce Data, Multiply Impact: Improved Model Accuracy and Robustness with Dataset ReinforcementFartash Faghri, Hadi Pouransari, Sachin Mehta, Mehrdad Farajtabar 等ICCV 2023 · 被引用 16 次
- Does your graph need a confidence boost? Convergent boosted smoothing on graphs with tabular node featuresJiuhai Chen, Jonas Mueller, Vassilis N. Ioannidis, Soji Adeshina 等ICLR 2022 · 被引用 13 次
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 被引用 4 次
它引用的顶会 Paper4
- On the Efficacy of Knowledge DistillationJang Hyun Cho, Bharath HariharanICCV 2019 · 被引用 741 次
- Neural Oblivious Decision Ensembles for Deep Learning on Tabular DataSergei Popov, Stanislav Morozov, Artem BabenkoICLR 2020 · 被引用 407 次
- Maximizing Overall Diversity for Improved Uncertainty Estimates in Deep EnsemblesSiddhartha Jain, Ge Liu, Jonas Mueller, David GiffordAAAI 2020 · 被引用 69 次
- Born-Again Tree EnsemblesThibaut Vidal, Maximilian SchifferICML 2020 · 被引用 62 次
相关 Paper
- HyperFast: Instant Classification for Tabular DataDavid Bonet, Daniel Mas Montserrat, Xavier Giró-i-Nieto, Alexander G. IoannidisAAAI 2024 · 被引用 29 次
- TabPack: Efficient Hyperparameter Ensembles for Tabular Deep LearningYury Gorishniy, Akim Kotelnikov, Ivan Rubachev, Artem BabenkoICML 2026
- TabM: Advancing tabular deep learning with parameter-efficient ensemblingYury Gorishniy, Akim Kotelnikov, Artem BabenkoICLR 2025
- iLTM: Integrated Large Tabular ModelDavid Bonet, Marçal Comajoan Cara, Alvaro Calafell, Daniel Mas Montserrat 等KDD 2026 · 被引用 4 次
- Doing More with Less: Characterizing Dataset Downsampling for AutoMLFatjon Zogaj, José Pablo Cambronero, Martin C. Rinard, Jürgen CitoVLDB 2021 · 被引用 20 次
