Sparcle: Boosting the Accuracy of Data Cleaning Systems through Spatial Awareness
Yuchuan Huang, Mohamed F. Mokbel
Abstract
Though data cleaning systems have earned great success and wide spread in both academia and industry, they fall short when trying to clean spatial data. The main reason is that state-of-the-art data cleaning systems mainly rely on functional dependency rules where there is sufficient co-occurrence of value pairs to learn that a certain value of an attribute leads to a corresponding value of another attribute. However, for spatial attributes that represent locations on the form of <latitude, longitude>, there is very little chance that two records would have the same exact coordinates, and hence co-occurrence would unlikely to exist. This paper presents S (SPatially-AwaRe CLEaning); a novel framework that injects spatial awareness into the core engine of rule-based data cleaning systems as a means of boosting their accuracy. S injects two main spatial concepts into the core engine of data cleaning systems: (1) Spatial Neighborhood, where co-occurrence is relaxed to be within a certain spatial proximity rather than same exact value, and (2) Distance Weighting, where records are given different weights of whether they satisfy a dependency rule, based on their relative distance. Experimental results using a real deployment of S inside a state-of-the-art data cleaning system, and real and synthetic datasets, show that S significantly boosts the accuracy of data cleaning systems when dealing with spatial data.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f41b4e7d-9cdc-45ee-904c-c4eae21243c2Cited by top-tier papers1
Ask how each one uses itBuilds on2
Related papers
- MDedup: Duplicate Detection with Matching DependenciesIoannis K. Koumarelas, Thorsten Papenbrock, Felix NaumannVLDB 2020 · 18 citations
- GIDCL: A Graph-Enhanced Interpretable Data Cleaning Framework with Large Language ModelsMengyi Yan, Yaoshu Wang, Yue Wang, Xiaoye Miao et al.SIGMOD 2025 · 12 citations
- From Suspicious Errors to Valid Data: On Repairing Spatio-Temporal Data via Spatial and Temporal DependenciesWeiwei Deng, Yu Sun, Shaoxu Song, Xiaojie YuanSIGMOD 2026
- Self-supervised and Interpretable Data Cleaning with Sequence Generative Adversarial NetworksJinfeng Peng, Derong Shen, Nan Tang, Tieying Liu et al.VLDB 2023 · 24 citations
- Generalizable Data Cleaning of Tabular Data in Latent SpaceEduardo Souza dos Reis, Mohamed Abdelaal, Carsten BinnigVLDB 2024 · 3 citations
