Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness
Yuchuan Huang, Mohamed F. Mokbel
摘要
Though data cleaning systems have earned great success and wide spread in both academia and industry, they fall short when trying to clean spatial data. The main reason is that state-of-the-art data cleaning systems mainly rely on functional dependency rules where there is sufficient co-occurrence of value pairs to learn that a certain value of an attribute leads to a corresponding value of another attribute. However, for spatial attributes that represent locations on the form of <latitude, longitude>, there is very little chance that two records would have the same exact coordinates, and hence co-occurrence would unlikely to exist. This paper presents S (SPatially-AwaRe CLEaning); a novel framework that injects spatial awareness into the core engine of rule-based data cleaning systems as a means of boosting their accuracy. S injects two main spatial concepts into the core engine of data cleaning systems: (1) Spatial Neighborhood, where co-occurrence is relaxed to be within a certain spatial proximity rather than same exact value, and (2) Distance Weighting, where records are given different weights of whether they satisfy a dependency rule, based on their relative distance. Experimental results using a real deployment of S inside a state-of-the-art data cleaning system, and real and synthetic datasets, show that S significantly boosts the accuracy of data cleaning systems when dealing with spatial data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper2
相关 Paper
- MDedup: Duplicate Detection with Matching DependenciesIoannis K. Koumarelas, Thorsten Papenbrock, Felix NaumannVLDB 2020 · 被引用 18 次
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao 等SIGMOD 2025 · 被引用 12 次
- From Suspicious Errors to Valid Data: On Repairing Spatio-Temporal Data via Spatial and Temporal DependenciesWeiwei Deng, Yu Sun, Shaoxu Song, Xiaojie YuanSIGMOD 2026
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu 等VLDB 2023 · 被引用 24 次
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 被引用 3 次
