Data Imputation with Limited Data Redundancy Using Data Lakes
Chenyu Yang, Yuyu Luo, Chuanxuan Cui, Ju Fan, Chengliang Chai, Nan Tang
摘要
Data imputation is essential for many data science applications. Existing methods rely heavily on sufficient data redundancy from within-table values. However, many real-world datasets often lack such data redundancy, necessitating external data sources. In this paper, we introduce a retrieval-augmented imputation framework, LakeFill , which combines large language models (LLMs) and data lakes to address this challenge. Unlike existing "table-level" retrieval methods designed for question answering, which retrieve data in the granularity of tables, LakeFill performs fine-grained "tuple-level" retrieval, optimized specifically for data imputation at the tuple level. It encodes (possibly incomplete) tuples to capture nuanced similarities and differences, enabling effective identification of candidate tuples. A novel reranking method that integrates checklist-based training data annotation with stratified training group construction further refines the retrieved tuples. Finally, a reasoner with a novel two-stage confidence-aware imputation ensures reliable imputation results. Extensive experiments show that LakeFill significantly outperforms state-of-the-art methods for data imputation when there is limited data redundancy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Debugging Defective Visualizations: Empirical Insights Informing a Human-AI Co‑Debugging SystemShuyu Shen, Sirong Lu, Leixian Shen, Yuyu LuoCHI 2026 · 被引用 2 次
- Auto-Fill: Learning to Predict Missing Values Accurately with Specialist Language ModelsYurong Liu, Yeye He, Haoyu Dong, Junjie Xing 等VLDB 2026
- TuneAhead: Predicting Fine-tuning Performance Before Training BeginsYuxiang Luo, Haonan Long, Chen Wang, Qiqi Duan 等ICML 2026
- A Risk Decomposition Framework for Pre-hoc Fine-tuning PredictionYuxiang Luo, Chen Wang, Nan TangICML 2026
它引用的顶会 Paper21
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Query Rewriting in Retrieval-Augmented Large Language ModelsXinbei Ma, Yeyun Gong, Pengcheng He, Hai Zhao 等EMNLP 2023 · 被引用 191 次
相关 Paper
- LakeQA: An Exploratory QA Benchmark over a Million-Scale Data LakeHaonan Wang, Jiaxiang Liu, Yurong Liu, Austin Wijaya 等ICML 2026
- Improving Context Fidelity via Native Retrieval-Augmented ReasoningSuyuchen Wang, Jinlin Wang, Xinyu Wang, Shiqi Li 等EMNLP 2025 · 被引用 1 次
- Text2sql-Flow: a Robust Sql-Aware Data Augmentation Framework for Text-To-SqlQifeng Cai, Hao Liang, Chang Xu, Tao Xie 等ICDE 2026 · 被引用 1 次
- LIMSSR: LLM-Driven Sequence-to-Score Reasoning under Training-Time Incomplete Multimodal ObservationsHuangbiao Xu, huanqi wu, Xiao Ke, Yuxin PengICML 2026 · 被引用 1 次
- Search and Refine During Think: Facilitating Knowledge Refinement for Improved Retrieval-Augmented ReasoningYaorui Shi, Sihang Li, Chang Wu, Zhiyuan Liu 等NeurIPS 2025 · 被引用 30 次
