NeurIDA: Dynamic Modeling for Effective In-Database Analytics
Lingze Zeng, Shaofeng Cai, Naili Xing, Jiaqi Zhu, Gang Chen, Peng Lu, Jian Pei, Beng Chin Ooi
摘要
Relational Database Management Systems (RDBMS) manage complex, interrelated data and support a broad spectrum of analytical tasks. With the growing demand for predictive analytics, the deep integration of machine learning (ML) into RDBMS has become critical. However, a fundamental challenge hinders this evolution: conventional ML models are static and task-specific, whereas RDBMS environments are dynamic and must support diverse analytical queries. Each analytical task entails constructing a bespoke pipeline from scratch, which incurs significant development overhead and hence limits the wide adoption of ML in analytics.
We present NeurIDA, an autonomous end-to-end system for in-database analytics that dynamically "tweaks" the best available base model to better serve a given analytical task. In particular, we propose a novel paradigm of dynamic in-database modeling to pre-train a composable base model architecture over the relational data. Upon receiving a task, NeurIDA formulates the task and data profile to dynamically select and configure relevant components from the pool of base models and shared model components for prediction. For a friendly user experience, NeurIDA supports natural language queries; it interprets user intent to construct structured task profiles and generates analytical reports with dedicated LLM agents. By design, NeurIDA enables ease-of-use and yet effective and efficient in-database AI analytics. Extensive experimental studies show that NeurIDA consistently delivers up to 12% improvement in AUC-ROC and 25% relative reduction in MAE across ten tasks on five real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Revisiting Deep Learning Models for Tabular DataYury Gorishniy, Ivan Rubachev, Valentin Khrulkov, Artem BabenkoNeurIPS 2021 · 被引用 1,847 次
- Understanding Architectures Learnt by Cell-based Neural Architecture SearchYao Shu, Wei Wang, Shaofeng CaiICLR 2020 · 被引用 92 次
- Making Pre-trained Language Models Great on Tabular PredictionJiahuan Yan, Bo Zheng, Hongxia Xu, Yiheng Zhu 等ICLR 2024 · 被引用 72 次
- Zero-Cost Proxies for Lightweight NASMohamed S. Abdelfattah, Abhinav Mehrotra, Lukasz Dudziak, Nicholas Donald LaneICLR 2021 · 被引用 65 次
- Cerebro: A Data System for Optimized Deep Learning Model SelectionSupun Nakandala, Yuhao Zhang, Arun KumarVLDB 2020 · 被引用 61 次
相关 Paper
- Adda: Towards Efficient in-Database Feature Generation via LLM-based AgentsKuan Lu, Zhihui Yang, Sai Wu, Ruichen Xia 等SIGMOD 2025 · 被引用 6 次
- Towards Automatic and Efficient Prediction Query Processing in Analytical DatabaseYuchen Peng, Zhongle Xie, Ke Chen, Gang Chen 等ICDE 2025 · 被引用 3 次
- Powering In-Database Dynamic Model Slicing for Structured Data AnalyticsLingze Zeng, Naili Xing, Shaofeng Cai, Gang Chen 等VLDB 2024 · 被引用 7 次
- MB2: Decomposed Behavior Modeling for Self-Driving Database Management SystemsLin Ma, William Zhang, Jie Jiao, Wuwen Wang 等SIGMOD 2021 · 被引用 35 次
- SEMA: A High-performance System for LLM-based Semantic Query ProcessingKangkang Qi, Dongyang Xie, Wenbo Li, Hao Zhang 等VLDB 2026 · 被引用 5 次
