FeatureLTE: Learning to Estimate Feature Importance
Tianping Zhang, Zhaoyang Wang, Chen Qian, Jian Li, Yin Lou
Abstract
Feature importance scores (FIS) estimation is an important problem in many data-intensive applications. Traditional approaches can be divided into two types; model-specific methods and model-agnostic methods. In this work, we present FeatureLTE, a novel learning-based approach to FIS estimation. For the first time, as we demonstrate through extensive experiments, it is possible to build general-purpose pre-trained models for FIS estimation. Therefore, FIS estimation reduces to prediction outputs from a pre-trained FeatureLTE model. Pre-trained FeatureLTE models enjoy several desired advantages, including accuracy, robustness, efficiency, and evolvability, and FeatureLTE models really begin to shine on large datasets where traditional methods often find themselves unable to scale. We build our pre-trained models for binary classification and regression problems using observations from nearly 1,000 public datasets. We systematically evaluate various design choices of FeatureLTE model construction and carefully design meta features to make sure that they are computationally lightweight. Based on our evaluation, FeatureLTE is on par with the best existing FIS estimators in terms of FIS quality, and achieves up to 339.48x speedup without sacrificing the quality of FIS estimates on large-scale datasets. Finally, we release two pre-trained FeatureLTE models for binary classification and regression problems that are ready to use on almost all tabular datasets, along with the repository of 701 binary classification datasets and 256 regression datasets with pre-computed feature importance scores to promote future research along this direction.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get c2c46480-6064-4918-b669-4469a5387f79Related papers
- PRICE: A Pretrained Model for Cross-Database Cardinality EstimationTianjing Zeng, Junwei Lan, Jiahong Ma, Wenqing Wei et al.VLDB 2025 · 14 citations
- OpenFE: Automated Feature Generation with Expert-level PerformanceTianping Zhang, Zheyu Aqa Zhang, Zhiyuan Fan, Haoyan Luo et al.ICML 2023 · 60 citations
- FrugalScore: Learning Cheaper, Lighter and Faster Evaluation Metrics for Automatic Text GenerationMoussa Kamal Eddine, Guokan Shang, Antoine J.-P. Tixier, Michalis VazirgiannisACL 2022
- Toward Efficient Automated Feature EngineeringKafeng Wang, Pengyang Wang, Chengzhong XuICDE 2023 · 6 citations
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
