DataVinci: Learning Syntactic and Semantic String Repairs
Mukul Singh, José Cambronero, Sumit Gulwani, Vu Le, Carina Negreanu, Arjun Radhakrishna, Gust Verbruggen
摘要
String data is common in real-world datasets: 67.6% of values in a sample of 1.8 million real Excel spreadsheets from the web were represented as text. Systems that successfully clean such string data can have a significant impact on real users. While prior work has explored errors in string data, proposed approaches have often been limited to error detection or require that the user provide annotations, examples, or constraints to fix the errors. Furthermore, these systems have focused independently on syntactic errors or semantic errors in strings, but ignore that strings often contain both syntactic and semantic substrings. We introduce DataVinci, a fully unsupervised string data error detection and repair system. DataVinci learns regular-expression-based patterns that cover a majority of values in a column and reports values that do not satisfy such patterns as data errors. DataVinci can automatically derive edits to the data error based on the majority patterns and constraints learned over other columns without the need for further user interaction. To handle strings with both syntactic and semantic substrings, DataVinci uses an LLM to abstract (and reconcretize) portions of strings that are semantic prior to learning majority patterns and deriving edits. Because not all data can result in majority patterns, DataVinci leverages execution information from an existing program (which reads the target data) to identify and correct data repairs that would not otherwise be identified. DataVinci outperforms 7 baselines on both error detection and repair when evaluated on 4 existing and new benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- CodeT5: Identifier-aware Unified Pre-trained Encoder-Decoder Models for Code Understanding and GenerationYue Wang, Weishi Wang, Shafiq R. Joty, Steven C. H. HoiEMNLP 2021 · 被引用 1,224 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Pattern Functional Dependencies for Data CleaningAbdulhakim Ali Qahtan, Nan Tang, Mourad Ouzzani, Yang Cao 等VLDB 2020 · 被引用 42 次
- Semantic programming by example with pre-trained modelsGust Verbruggen, Vu Le, Sumit GulwaniOOPSLA 2021 · 被引用 26 次
相关 Paper
- Automated, Unsupervised, and Auto-Parameterized Inference of Data Patterns and Anomaly DetectionQiaolin Qin, Heng Li, Ettore Merlo, Maxime LamotheICSE 2025 · 被引用 1 次
- Human-in-the-loop Regular Expression Extraction for Single Column Format InconsistencyShaochen Yu, Lei Han, Marta Indulska, Shazia Sadiq 等WWW 2023 · 被引用 3 次
- Auto-Validate: Unsupervised Data Validation Using Data-Domain Patterns Inferred from Data LakesJie Song, Yeye HeSIGMOD 2021 · 被引用 26 次
- Can an LLM Find Its Way Around a Spreadsheet?Cho-Ting Lee, Andrew Neeser, Shengzhe Xu, Jay Katyan 等ICSE 2025 · 被引用 1 次
- A Zero-Training Error Correction System with Large Language ModelsYangyang Wu, Chen Yang, Mengying Zhu, Xiaoye Miao 等ICDE 2025 · 被引用 7 次
