Towards Benchmarking Feature Type Inference for AutoML Platforms
Vraj Shah, Jonathan Lacanlale, Premanand Kumar, Kevin Yang, Arun Kumar
Abstract
The paradigm of AutoML has created an opportunity to enable ML for the masses. Emerging industrial-scale cloud AutoML platforms aim to automate the end-to-end ML workflow. While many works have looked into automated feature engineering, model selection, or hyper-parameter search in AutoML, little work has studied a crucial step that serves as an entry point to this workflow: ML feature type inference. The semantic gap between attribute types (e.g., strings, numbers) in databases/files and ML feature types (e.g., Numeric, Categorical) necessitates type inference. In this work, we formalize and standardize this task by creating the first ever benchmark labeled dataset, which we use to objectively evaluate existing AutoML tools. Our dataset has 9921 examples and a 9-class label vocabulary. Our labeled data also offers an alternative approach to automate this task than existing rule-based or syntax-based approaches: use ML itself to predict feature types. We collate a benchmark suite of 30 classification and regression tasks to assess the importance of type inference for downstream models. Empirical comparison on our labeled data shows that an ML-based approach delivers a lift of an average 14% and up to 38% in accuracy for identifying feature types compared to prominent industrial tools. Our downstream benchmark suite reveals that the ML-based approach outperforms existing industrial-strength tools for 47 out of 60 downstream models. We release our labeled dataset, models, and downstream benchmarks in a public repository with a leaderboard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Explaining Dataset Changes for Semantic Data Versioning with Explain-Da-VRoee Shraga, Renée J. MillerVLDB 2023 · 18 citations
- UPLIFT: Parallelization Strategies for Feature Transformations in Machine Learning WorkloadsArnab Phani, Lukas Erlbacher, Matthias BoehmVLDB 2022 · 13 citations
- Pollock: A Data Loading BenchmarkGerardo Vitagliano, Mazhar Hameed, Lan Jiang, Lucas Reisener et al.VLDB 2023 · 9 citations
- SchemaPile: A Large Collection of Relational Database SchemasTill Döhmen, Radu Geacu, Madelon Hulsebos, Sebastian SchelterSIGMOD 2024 · 9 citations
- How do Categorical Duplicates Affect ML? A New Benchmark and Empirical AnalysesVraj Shah, Thomas J. Parashos, Arun KumarVLDB 2024 · 8 citations
Builds on1
Related papers
- Large Language Models for Automated Data Science: Introducing CAAFE for Context-Aware Automated Feature EngineeringNoah Hollmann, Samuel Müller, Frank HutterNeurIPS 2023 · 210 citations
- SAPIENTML: Synthesizing Machine Learning Pipelines by Learning from Human-Written SolutionsRipon K. Saha, Akira Ura, Sonal Mahajan, Chenguang Zhu et al.ICSE 2022 · 11 citations
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
- Adda: Towards Efficient in-Database Feature Generation via LLM-based AgentsKuan Lu, Zhihui Yang, Sai Wu, Ruichen Xia et al.SIGMOD 2025 · 6 citations
- Choose Wisely: An Extensive Evaluation of Model Selection for Anomaly Detection in Time SeriesEmmanouil Sylligardos, Paul Boniol, John Paparrizos, Panos E. Trahanias et al.VLDB 2023 · 40 citations
