Reliable and Cost-Effective Exploratory Data Analysis via Graph-Guided RAG
Mossad Helali, Yutai Luo, Tae Jun Ham, Jim Plotts, Ashwin Chaugule, Jichuan Chang, Parthasarathy Ranganathan, Essam Mansour
Abstract
Automating Exploratory Data Analysis (EDA) is critical for accelerating the workflow of data scientists. While Large Language Models (LLMs) offer a promising solution, current LLM-only approaches often exhibit limited accuracy and code reliability on less-studied or private datasets. Moreover, their effectiveness significantly diminishes with open-source LLMs compared to proprietary ones, limiting their usability in enterprises that prefer local models for privacy and cost. To address these limitations, we introduce RAGvis: a novel two-stage graph-guided Retrieval-Augmented Generation (RAG) framework. RAGvis first builds a base knowledge graph (KG) of EDA notebooks and enriches it with structured EDA operation semantics. These semantics are extracted by an LLM guided by our empiricallydeveloped EDA operations taxonomy. Second, in the online generation stage for new datasets, RAGvis retrieves relevant operations from the KG, aligns them to the dataset's structure, refines them with LLM reasoning, and then employs a self-correcting agent to generate executable Python code. Experiments on two benchmarks demonstrate that RAGvis significantly improves code executability (pass rate), semantic accuracy, and visual quality in generated operations. This enhanced performance is achieved with substantially lower token usage compared to LLM-only baselines. Notably, our approach enables smaller, opensource LLMs to match the performance of proprietary models, presenting a reliable and costeffective pathway for automated EDA code generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9dcca5f-e915-468c-b62c-8964e2e891c4Builds on5
- InfiAgent-DABench: Evaluating Agents on Data Analysis TasksXueyu Hu, Ziyu Zhao, Shuang Wei, Ziwei Chai et al.ICML 2024 · 110 citations
- Large Language Models Meet NL2Code: A SurveyDaoguang Zan, Bei Chen, Fengji Zhang, Dianjie Lu et al.ACL 2023 · 104 citations
- VisEval: A Benchmark for Data Visualization in the Era of Large Language ModelsNan Chen, Yuge Zhang, Jiahang Xu, Kan Ren et al.IEEE VIS 2024 · 44 citations
- MetaInsight: Automatic Discovery of Structured Knowledge for Exploratory Data AnalysisPingchuan Ma, Rui Ding, Shi Han, Dongmei ZhangSIGMOD 2021 · 35 citations
- KGLiDS: A Platform for Semantic Abstraction, Linking, and Automation of Data ScienceMossad Helali, Niki Monjazeb, Shubham Vashisth, Philippe Carrier et al.ICDE 2024 · 7 citations
Related papers
- RTLFixer: Automatically Fixing RTL Syntax Errors with Large Language ModelYunda Tsai, Mingjie Liu, Haoxing RenDAC 2024 · 95 citations
- Can an LLM Find Its Way Around a Spreadsheet?Cho-Ting Lee, Andrew Neeser, Shengzhe Xu, Jay Katyan et al.ICSE 2025 · 1 citation
- Knowledge-Enhanced Program Repair for Data Science CodeShuyin Ouyang, Jie M. Zhang, Zeyu Sun, Albert Meroño-PeñuelaICSE 2025 · 2 citations
- ReMindRAG: Low-Cost LLM-Guided Knowledge Graph Traversal for Efficient RAGYikuan Hu, Jifeng Zhu, Lanrui Tang, Chen HuangNeurIPS 2025 · 10 citations
- VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality DocumentsShi Yu, Chaoyue Tang, Bokai Xu, Junbo Cui et al.ICLR 2025
