Learning a Data-Driven Policy Network for Pre-Training Automated Feature Engineering
Liyao Li, Haobo Wang, Liangyu Zha, Qingyi Huang, Sai Wu, Gang Chen, Junbo Zhao
Abstract
Feature engineering is widely acknowledged to be pivotal in tabular data analysis and prediction. Automated feature engineering (AutoFE) emerged to automate this process managed by experienced data scientists and engineers conventionally. In this area, most — if not all — prior work adopted an identical framework from the neural architecture search (NAS) method. While feasible, we posit that the NAS framework very much contradicts the way how human experts cope with the data since the inherent Markov decision process (MDP) setup differs. We point out that its data-unobserved setup consequentially results in an incapability to generalize across different datasets as well as also high computational cost. This paper proposes a novel AutoFE framework Feature Set Data-Driven Search (FETCH), a pipeline mainly for feature generation and selection. Notably, FETCH is built on a brand-new data-driven MDP setup using the tabular dataset as the state fed into the policy network. Further, we posit that the crucial merit of FETCH is its transferability where the yielded policy network trained on a variety of datasets is indeed capable to enact feature engineering on unseen data, without requiring additional exploration. To the best of our knowledge, this is a pioneer attempt to build a tabular data pre-training paradigm via AutoFE. Extensive experiments show that FETCH systematically surpasses the current state-of-the-art AutoFE methods and validates the transferability of AutoFE pre-training.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Cited by top-tier papers9
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- OpenFE: Automated Feature Generation with Expert-level PerformanceTianping Zhang, Zheyu Aqa Zhang, Zhiyuan Fan, Haoyan Luo et al.ICML 2023 · 60 citations
- MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementJaehyun Nam, Jinsung Yoon, Jiefeng Chen, Jinwoo Shin et al.NeurIPS 2025 · 58 citations
- Towards Cross-Table Masked Pretraining for Web Data MiningChao Ye, Guoshan Lu, Haobo Wang, Liyao Li et al.WWW 2024 · 23 citations
Related papers
- Catch: Collaborative Feature Set Search for Automated Feature EngineeringGuoshan Lu, Haobo Wang, Saisai Yang, Jing Yuan et al.WWW 2023 · 6 citations
- Toward Efficient Automated Feature EngineeringKafeng Wang, Pengyang Wang, Chengzhong XuICDE 2023 · 6 citations
- DiffPrep: Differentiable Data Preprocessing Pipeline Search for Learning over Tabular DataPeng Li, Zhiyi Chen, Xu Chu, Kexin RongSIGMOD 2023 · 24 citations
- CoFE: Collaborative Feature Engineering via Semantically-Guided Exploration and Diagnostic-Driven RefinementWeihao Jiang, Ziang Nan, Zhihui Shi, Ya Cong et al.KDD 2026
- DAGPipe: Differentiable DAG Learning for Automated Data PreparationJing Chang, Chang LiuKDD 2026
