FeatureLTE: Learning to Estimate Feature Importance
Tianping Zhang, Zhaoyang Wang, Chen Qian, Jian Li, Yin Lou
摘要
Feature importance scores (FIS) estimation is an important problem in many data-intensive applications. Traditional approaches can be divided into two types; model-specific methods and model-agnostic methods. In this work, we present FeatureLTE, a novel learning-based approach to FIS estimation. For the first time, as we demonstrate through extensive experiments, it is possible to build general-purpose pre-trained models for FIS estimation. Therefore, FIS estimation reduces to prediction outputs from a pre-trained FeatureLTE model. Pre-trained FeatureLTE models enjoy several desired advantages, including accuracy, robustness, efficiency, and evolvability, and FeatureLTE models really begin to shine on large datasets where traditional methods often find themselves unable to scale. We build our pre-trained models for binary classification and regression problems using observations from nearly 1,000 public datasets. We systematically evaluate various design choices of FeatureLTE model construction and carefully design meta features to make sure that they are computationally lightweight. Based on our evaluation, FeatureLTE is on par with the best existing FIS estimators in terms of FIS quality, and achieves up to 339.48x speedup without sacrificing the quality of FIS estimates on large-scale datasets. Finally, we release two pre-trained FeatureLTE models for binary classification and regression problems that are ready to use on almost all tabular datasets, along with the repository of 701 binary classification datasets and 256 regression datasets with pre-computed feature importance scores to promote future research along this direction.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- PRICE: A Pretrained Model for Cross-Database Cardinality EstimationTianjing Zeng, Junwei Lan, Jiahong Ma, Wenqing Wei 等VLDB 2025 · 被引用 14 次
- OpenFE: Automated Feature Generation with Expert-level PerformanceTianping Zhang, Zheyu Aqa Zhang, Zhiyuan Fan, Haoyan Luo 等ICML 2023 · 被引用 60 次
- FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metrics for Automatic Text GenerationMoussa Kamal Eddine, Guokan Shang, Antoine J.-P. Tixier, Michalis VazirgiannisACL 2022
- Toward Efficient Automated Feature EngineeringKafeng Wang, Pengyang Wang, Chengzhong XuICDE 2023 · 被引用 6 次
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
