NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Lingze Zeng, Shaofeng Cai, Naili Xing, Jiaqi Zhu, Gang Chen, Peng Lu, Jian Pei, Beng Chin Ooi
Abstract
Relational Database Management Systems (RDBMS) manage complex, interrelated data and support a broad spectrum of analytical tasks. With the growing demand for predictive analytics, the deep integration of machine learning (ML) into RDBMS has become critical. However, a fundamental challenge hinders this evolution: conventional ML models are static and task-specific, whereas RDBMS environments are dynamic and must support diverse analytical queries. Each analytical task entails constructing a bespoke pipeline from scratch, which incurs significant development overhead and hence limits the wide adoption of ML in analytics.
We present NeurIDA, an autonomous end-to-end system for in-database analytics that dynamically "tweaks" the best available base model to better serve a given analytical task. In particular, we propose a novel paradigm of dynamic in-database modeling to pre-train a composable base model architecture over the relational data. Upon receiving a task, NeurIDA formulates the task and data profile to dynamically select and configure relevant components from the pool of base models and shared model components for prediction. For a friendly user experience, NeurIDA supports natural language queries; it interprets user intent to construct structured task profiles and generates analytical reports with dedicated LLM agents. By design, NeurIDA enables ease-of-use and yet effective and efficient in-database AI analytics. Extensive experimental studies show that NeurIDA consistently delivers up to 12% improvement in AUC-ROC and 25% relative reduction in MAE across ten tasks on five real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8ae9c9d2-622c-4d8f-8fb0-96caca677a13Builds on19
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 1,847 citations
- Understanding Architectures Learnt by Cell-based Neural Architecture SearchYao Shu, Wei Wang, Shaofeng CaiICLR 2020 · 92 citations
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu et al.ICLR 2024 · 72 citations
- Zero-Cost Proxies for Lightweight NASMohamed S. Abdelfattah, Abhinav Mehrotra, Lukasz Dudziak, Nicholas Donald LaneICLR 2021 · 65 citations
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 61 citations
Related papers
- Adda: Towards Efficient in-Database Feature Generation via LLM-based AgentsKuan Lu, Zhihui Yang, Sai Wu, Ruichen Xia et al.SIGMOD 2025 · 6 citations
- Towards Automatic and Efficient Prediction Query Processing in Analytical DatabaseYuchen Peng, Zhongle Xie, Ke Chen, Gang Chen et al.ICDE 2025 · 3 citations
- Powering In-Database Dynamic Model Slicing for Structured Data AnalyticsLingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen et al.VLDB 2024 · 7 citations
- MB2: Decomposed Behavior Modeling for Self-Driving Database Management SystemsLin Ma, William Zhang, Jie Jiao, Wuwen Wang et al.SIGMOD 2021 · 35 citations
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang et al.VLDB 2026 · 5 citations
