The Semantic Architect: How FEAML Bridges Structured Data and LLMs for Multi-Label Tasks
Wanfu Gao, Zebin He, Jun Gao
Abstract
Existing feature engineering methods based on large language models (LLMs) have not yet been applied to multi-label learning tasks. They lack the ability to model complex label dependencies and are not specifically adapted to the characteristics of multi-label tasks. To address the above issues, we propose Feature Engineering Automation for Multi-Label Learning (FEAML), an automated feature engineering method for multi-label classification which leverages the code generation capabilities of LLMs. By utilizing metadata and label co-occurrence matrices, LLMs are guided to understand the relationships between data features and task objectives, based on which high-quality features are generated.The newly generated features are evaluated in terms of model accuracy to assess their effectiveness, while Pearson correlation coefficients are used to detect redundancy. FEAML further incorporates the evaluation results as feedback to drive LLMs to continuously optimize code generation in subsequent iterations. By integrating LLMs with a feedback mechanism, FEAML realizes an efficient, interpretable and self-improving feature engineering paradigm. Empirical results on various multi-label datasets demonstrate that our FEAML outperforms other feature engineering methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LIFT: Language-Interfaced Fine-Tuning for Non-language Machine Learning TasksTuan Dinh, Yuchen Zeng, Ruisu Zhang, Ziqian Lin et al.NeurIPS 2022 · 222 citations
- Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningSungwon Han, Jinsung Yoon, Sercan Ö. Arik, Tomas PfisterICML 2024 · 81 citations
- Optimized Feature Generation for Tabular Data via LLMs with Decision Tree ReasoningJaehyun Nam, Kyuyoung Kim, Seunghyuk Oh, Jihoon Tack et al.NeurIPS 2024 · 78 citations
- Evolutionary Large Language Model for Automated Feature TransformationNanxu Gong, Chandan K. Reddy, Wangyang Ying, Haifeng Chen et al.AAAI 2025 · 39 citations
Related papers
- MORE-FE: Multi-Operator and Reinforcement Learning-Enhanced Evolution for LLM Feature EngineeringChang-Yu Chao, Bryan Andersen, Xiao Xi Tan, Yi-Tse Lu et al.KDD 2026
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 210 citations
- Memory-Augmented LLM-based Multi-Agent System for Automated Feature Generation on Tabular DataFengxian Dong, Zhi Zheng, Xiao Han, Wei Chen et al.ACL 2026
- FEA-Bench: A Benchmark for Evaluating Repository-Level Code Generation for Feature ImplementationWei Li, Xin Zhang, Zhongxin Guo, Shaoguang Mao et al.ACL 2025 · 40 citations
- Human-LLM Collaborative Feature Engineering for Tabular DataZhuoyan Li, Aditya Bansal, Jinzhao Li, Shishuang He et al.ICLR 2026 · 2 citations
