Fastft: Accelerating Reinforced Feature Transformation via Advanced Exploration Strategies
Tianqi He, Xiaohan Huang, Yi Du, Qingqing Long, Ziyue Qiao, Min Wu, Yanjie Fu, Yuanchun Zhou, Meng Xiao
Abstract
Feature Transformation is crucial for classic machine learning that aims to generate feature combinations to enhance the performance of downstream tasks from a data-centric perspective. Current methodologies, such as manual expert-driven processes, iterative-feedback techniques, and exploration-generative tactics, have shown promise in automating such data engineering workflow by minimizing human involvement. However, three challenges remain in those frameworks: (1) It predominantly depends on downstream task performance metrics, as assessment is time-consuming, especially for large datasets. (2) The diversity of feature combinations will hardly be guaranteed after random exploration ends. (3) Rare significant transformations lead to sparse valuable feedback that hinders the learning processes or leads to less effective results. In response to these challenges, we introduce FASTFT, an innovative framework that leverages a trio of advanced strategies. We first decouple the feature transformation evaluation from the outcomes of the generated datasets via the performance predictor. To address the issue of reward sparsity, we developed a method to evaluate the novelty of generated transformation sequences. Incorporating this novelty into the reward function accelerates the model's exploration of effective transformations, thereby improving the search productivity. Additionally, we combine novelty and performance to create a prioritized memory buffer, ensuring that essential experiences are effectively revisited during exploration. Our extensive experimental evaluations validate the performance, efficiency, and traceability of our proposed framework, showcasing its superiority in handling complex feature transformation tasks11The code and data are publicly accessible via Github..
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 91272c4c-d21a-474c-9e12-716bcd79c4adBuilds on11
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 210 citations
- AutoDS: Towards Human-Centered Automation of Data ScienceDakuo Wang, Josh Andres, Justin D. Weisz, Erick Oduor et al.CHI 2021 · 77 citations
- OpenFE: Automated Feature Generation with Expert-level PerformanceTianping Zhang, Zheyu Aqa Zhang, Zhiyuan Fan, Haoyan Luo et al.ICML 2023 · 60 citations
- Anti-Exploration by Random Network DistillationAlexander Nikulin, Vladislav Kurenkov, Denis Tarasov, Sergey KolesnikovICML 2023 · 55 citations
- Exploration and Anti-Exploration with Distributional Random Network DistillationKai Yang, Jian Tao, Jiafei Lyu, Xiu LiICML 2024 · 37 citations
Related papers
- Sculpting Features from Noise: Reward-Guided Hierarchical Diffusion for Task-Optimal Feature TransformationNanxu Gong, Zijun Li, Sixun Dong, Haoyue Bai et al.NeurIPS 2025 · 15 citations
- Heterogeneous Multi-Agent Reinforcement Learning with Attention for Cooperative and Scalable Feature TransformationTao Zhe, Huazhen Fang, Kunpeng Liu, Qian Lou et al.KDD 2026
- Evolutionary Large Language Model for Automated Feature TransformationNanxu Gong, Chandan K. Reddy, Wangyang Ying, Haifeng Chen et al.AAAI 2025 · 39 citations
- Reinforcement-Enhanced Autoregressive Feature Transformation: Gradient-steered Search in Continuous Space for Postfix ExpressionsDongjie Wang, Meng Xiao, Min Wu, Pengfei Wang et al.NeurIPS 2023 · 34 citations
- The Semantic Architect: How FEAML Bridges Structured Data and LLMs for Multi-Label TasksWanfu Gao, Zebin He, Jun GaoAAAI 2026
