CoFE: Collaborative Feature Engineering via Semantically-Guided Exploration and Diagnostic-Driven Refinement
Weihao Jiang, Ziang Nan, Zhihui Shi, Ya Cong, Jun Xiao, Xiaoye Miao
摘要
In domains such as finance, healthcare, and industry, feature engineering remains the key bottleneck limiting the performance of machine learning models on tabular data. While Automated Feature Engineering (AutoFE) aims to reduce this manual effort, existing approaches still suffer from distinct limitations: data-driven exploration can waste substantial computation on semantically meaningless feature combinations, whereas knowledge-driven approaches using Large Language Models (LLMs) struggle to construct high-order interactions without rich structural context and are typically guided only by coarse global metrics. We propose CoFE (Collaborative Feature Engineering), a two-phase framework that tightly couples search-based exploration with LLM-driven reasoning. In the exploration phase, CoFE leverages an LLM-constructed semantic feature schema to guide a Monte Carlo Tree Search (MCTS), enforcing semantic constraints while encouraging a diverse pool of complex candidate features that provides the missing structural context for LLMs. In the refinement phase, CoFE introduces feature health reports, a diagnostic artifact that supplies the LLM with actionable sample-level and structural feedback for targeted corrections. Experiments on 16 public tabular benchmarks show that CoFE consistently outperforms state-of-the-art data-driven and LLM-based AutoFE methods on the majority of datasets, while offering favorable computational efficiency.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- CoFEH: LLM-driven Feature Engineering Empowered by Collaborative Bayesian Hyperparameter OptimizationBeicheng Xu, Keyao Ding, Wei Liu, Yupeng Lu 等KDD 2026
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 被引用 210 次
- MORE-FE: Multi-Operator and Reinforcement Learning-Enhanced Evolution for LLM Feature EngineeringChang-Yu Chao, Bryan Andersen, Xiao Xi Tan, Yi-Tse Lu 等KDD 2026
- Human-LLM Collaborative Feature Engineering for Tabular DataZhuoyan Li, Aditya Bansal, Jinzhao Li, Shishuang He 等ICLR 2026 · 被引用 2 次
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack 等NeurIPS 2024 · 被引用 78 次
