Improving Data Leakage Detection in Machine Learning Notebooks through Static Slicing and Structured LLM Prompts
Taha Draoui, Mohamed Wiem Mkaouer, Christian D. Newman
Abstract
Data leakage remains a critical yet under-diagnosed issue in machine learning pipelines, leading to inflated results and unreliable deployments. Existing detection approaches rely on static rules that often miss open-coded manipulations and fail to capture the diversity of real-world notebooks. This paper introduces a novel methodology that integrates static slicing with large language models (LLMs) to improve leakage detection. We use a Datalog-based static analysis that isolates compact, provenance-aware slices corresponding to model training and evaluation pairs, and we pair these with structured LLM prompts that guide step-by-step reasoning about potential leakage for each isolated slice. Evaluated on a curated benchmark of Python notebooks from Kaggle and GitHub, our approach achieves state-of-the-art performance in both preprocessing and overlap leakage detection, improving F1 scores over the previous state-of-the-art by 22% for preprocessing leakage and 15% for overlap leakage. Beyond these improvements, our slicing-based methodology substantially outperforms end-to-end prompting, demonstrating that precise program slicing is key to enabling LLMs to reliably detect leakage. Our findings highlight the effectiveness of combining program slicing and prompt engineering for data leakage detection and establish the first systematic LLM-based solution for detecting data leakage in machine learning code.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 17ac7104-fb2d-4ddb-9a9d-de8fcd62ae78Related papers
- Data Leakage in Notebooks: Static Detection and Better ProcessesChenyang Yang, Rachel A. Brower-Sinning, Grace A. Lewis, Christian KästnerASE 2022 · 25 citations
- Pig: Leveraging Large Language Models for Python Library MigrationsMiryeong Kang, Wonseok Oh, Gabin An, Hakjoo OhFSE 2026
- Blended Analysis for Predictive ExecutionYi Li, Hridya Dhulipala, Aashish Yadavally, Xiaokai Rong et al.FSE 2025 · 1 citation
- Training on the Benchmark Is Not All You NeedShiwen Ni, Xiangtao Kong, Chengming Li, Xiping Hu et al.AAAI 2025 · 24 citations
- NESA: Relational Neuro-Symbolic Static Program AnalysisChengpeng Wang, Yifei Gao, Wuqi Zhang, Xuwei Liu et al.FSE 2026 · 1 citation
