Baran: Effective Error Correction via a Unified Context Representation and Transfer Learning
Mohammad Mahdavi, Ziawasch Abedjan
摘要
Traditional error correction solutions leverage handmaid rules or master data to find the correct values. Both are often amiss in real-world scenarios. Therefore, it is desirable to additionally learn corrections from a limited number of example repairs. To effectively generalize example repairs, it is necessary to capture the entire context of each erroneous value. A context comprises the value itself, the co-occurring values inside the same tuple, and all values that define the attribute type. Typically, an error corrector based on any of these context information undergoes an individual process of operations that is not always easy to integrate with other types of error correctors. In this paper, we present a new error correction system, Baran, which provides a unifying abstraction for integrating multiple error corrector models that can be pretrained and updated in the same way. Because of the holistic nature of our approach, we generate more correction candidates than state of the art and, because of the underlying context-aware data representation, we achieve high precision. We show that, by pretraining our models based on Wikipedia revisions, our system can further improve its overall precision and recall. In our experiments, Baran significantly outperforms state-of-the-art error correction systems in terms of effectiveness and human involvement requiring only 20 labeled tuples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Horizon: Scalable Dependency-driven Data CleaningEl Kindi Rezig, Mourad Ouzzani, Walid G. Aref, Ahmed K. Elmagarmid 等VLDB 2021 · 被引用 95 次
- RPT: Relational Pre-trained Transformer Is Almost All You Need towards Democratizing Data PreparationNan Tang, Ju Fan, Fangyi Li, Jianhong Tu 等VLDB 2021 · 被引用 92 次
- Annotating Columns with Pre-trained Language ModelsYoshihiko Suhara, Jinfeng Li, Yuliang Li, Dan Zhang 等SIGMOD 2022 · 被引用 81 次
- Table-GPT: Table Fine-tuned GPT for Diverse Table TasksPeng Li, Yeye He, Dror Yashar, Weiwei Cui 等SIGMOD 2024 · 被引用 63 次
- Rotom: A Meta-Learned Data Augmentation Framework for Entity Matching, Data Cleaning, Text Classification, and BeyondZhengjie Miao, Yuliang Li, Xiaolan WangSIGMOD 2021 · 被引用 63 次
相关 Paper
- A Zero-Training Error Correction System with Large Language ModelsYangyang Wu, Chen Yang, Mengying Zhu, Xiaoye Miao 等ICDE 2025 · 被引用 7 次
- Grammatical Error Correction in Low Error Density Domains: A New Benchmark and AnalysesSimon Flachs, Ophélie Lacroix, Helen Yannakoudakis, Marek Rei 等EMNLP 2020
- Editor: Multi-Resolution Cleaning of Multivariate Time Series Via Detect-Localize-RepairChenyang Li, Chaohong Ma, Xiaohui Yu, Cailong Li 等ICDE 2026
- Extracting Zero-shot Structured Information from Form-like Documents: Pretraining with Keys and TriggersRongyu Cao, Ping LuoAAAI 2021 · 被引用 6 次
- Learning Over Dirty Data Without CleaningJose Picado, John Davis, Arash Termehchy, Ga Young LeeSIGMOD 2020 · 被引用 13 次
