Progressive Entity Matching: A Design Space Exploration
Jakub Maciejewski, Konstantinos Nikoletos, George Papadakis, Yannis Velegrakis
Abstract
Entity Resolution (ER) is typically implemented as a batch task that processes all available data before identifying duplicate records. However, applications with time or computational constraints, e.g., those running in the cloud, require a progressive approach that produces results in a pay-as-you-go fashion. Numerous algorithms have been proposed for Progressive ER in the literature. In this work, we propose a novel framework for Progressive Entity Resolution that organizes relevant techniques into four consecutive steps: (i) filtering, which reduces the search space to the most likely candidate matches, (ii) weighting, which associates every pair of candidate matches with a similarity score, (iii) scheduling, which prioritizes the execution of the candidate matches so that the real duplicates precede the non-matching pairs, and (iv) matching, which applies a complex, matching function to the pairs in the order defined by the previous step. We associate each step with existing and novel techniques, illustrating that our framework overall generates a superset of the main existing works in the field. We select the most representative combinations resulting from our framework and fine-tune them over 10 established datasets for Record Linkage and 8 for Deduplication, with our results indicating that our taxonomy yields a wide range of high performing progressive techniques both in terms of effectiveness and time efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3b5c3079-95e2-45e1-8725-aaeb9e6fdf53Cited by top-tier papers1
Ask how each one uses itBuilds on10
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao et al.NeurIPS 2020 · 2,727 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
- Deep Learning for Blocking in Entity Matching: A Design Space ExplorationSaravanan Thirumuruganathan, Han Li, Nan Tang, Mourad Ouzzani et al.VLDB 2021 · 109 citations
- Pre-trained Embeddings for Entity Resolution: An Experimental AnalysisAlexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis KoubarakisVLDB 2023 · 63 citations
Related papers
- Entity Resolution On-DemandGiovanni Simonini, Luca Zecchini, Sonia Bergamaschi, Felix NaumannVLDB 2022 · 33 citations
- End-to-end Task Based Parallelization for Entity Resolution on Dynamic DataLeonardo Gazzarri, Melanie HerschelICDE 2021 · 14 citations
- A Critical Re-evaluation of Record Linkage Benchmarks for Learning-Based Matching AlgorithmsGeorge Papadakis, Nishadi Kirielle, Peter Christen, Themis PalpanasICDE 2024 · 8 citations
- Deep and Collective Entity Resolution in ParallelTing Deng, Wenfei Fan, Ping Lu, Xiaomeng Luo et al.ICDE 2022 · 6 citations
- HyperBlocker: Accelerating Rule-based Blocking in Entity Resolution using GPUsXiaoke Zhu, Min Xie, Ting Deng, Qi ZhangVLDB 2025 · 2 citations
