Data Imputation with Limited Data Redundancy Using Data Lakes
Chenyu Yang, Yuyu Luo, Chuanxuan Cui, Ju Fan, Chengliang Chai, Nan Tang
Abstract
Data imputation is essential for many data science applications. Existing methods rely heavily on sufficient data redundancy from within-table values. However, many real-world datasets often lack such data redundancy, necessitating external data sources. In this paper, we introduce a retrieval-augmented imputation framework, LakeFill , which combines large language models (LLMs) and data lakes to address this challenge. Unlike existing "table-level" retrieval methods designed for question answering, which retrieve data in the granularity of tables, LakeFill performs fine-grained "tuple-level" retrieval, optimized specifically for data imputation at the tuple level. It encodes (possibly incomplete) tuples to capture nuanced similarities and differences, enabling effective identification of candidate tuples. A novel reranking method that integrates checklist-based training data annotation with stratified training group construction further refines the retrieved tuples. Finally, a reasoner with a novel two-stage confidence-aware imputation ensures reliable imputation results. Extensive experiments show that LakeFill significantly outperforms state-of-the-art methods for data imputation when there is limited data redundancy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc53632a-0128-4273-8b50-25ad46bb404eCited by top-tier papers4
- Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging SystemShuyu Shen, Sirong Lu, Leixian Shen, Yuyu LuoCHI 2026 · 2 citations
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing et al.VLDB 2026
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan et al.ICML 2026
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
Builds on21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu et al.VLDB 2021 · 2,406 citations
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan et al.VLDB 2021 · 484 citations
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 325 citations
- Query Rewriting in Retrieval-Augmented Large Language ModelsXinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao et al.EMNLP 2023 · 191 citations
Related papers
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya et al.ICML 2026
- Improving Context Fidelity via Native Retrieval-Augmented ReasoningSuyuchen Wang, Jinlin Wang, Xinyu Wang, Shiqi Li et al.EMNLP 2025 · 1 citation
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie et al.ICDE 2026 · 1 citation
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 1 citation
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented ReasoningYaorui Shi, Sihang Li, Chang Wu, Zhiyuan Liu et al.NeurIPS 2025 · 30 citations
