In-context Clustering-based Entity Resolution with Large Language Models: A Design Space Exploration
Jiajie Fu, Haitong Tang, Arijit Khan, Sharad Mehrotra, Xiangyu Ke, Yunjun Gao
摘要
Entity Resolution (ER) is a fundamental data quality improvement task that identifies and links records referring to the same real-world entity. Traditional ER approaches often rely on pairwise comparisons, which can be costly regarding both time and monetary resources, especially when large datasets are involved. Recently, Large Language Models (LLMs) have demonstrated promising results in ER tasks. Still, existing methods typically focus on pairwise matching, missing the potential of LLMs to directly perform clustering in a more cost-effective and scalable manner. In this paper, we propose a novel in-context clustering approach for ER, where LLMs are used to cluster records directly, reducing both time complexity and monetary costs. We systematically investigate the design space for in-context clustering, analyzing the impact of factors such as set size, diversity, variation, and ordering of records on clustering performance. Based on these insights, we develop LLM-CER (LLM-powered Clustering-based ER) that obtains high-quality ER results while minimizing LLM API calls. Our approach addresses key challenges, including efficient cluster merging and LLM's hallucination, providing a scalable and effective solution for ER. Extensive experiments on nine real-world datasets demonstrate that our method significantly improves result quality, achieving up to 150% higher accuracy, 10% increase in the FP-measure, and reducing API calls by up to 5X, while maintaining a comparable monetary cost to the most cost-effective baseline.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- Can Foundation Models Wrangle Your Data?Avanika Narayan, Ines Chami, Laurel J. Orr, Christopher RéVLDB 2023 · 被引用 325 次
- Deep Comprehensive Correlation Mining for Image ClusteringJianlong Wu, Keyu Long, Fei Wang, Chen Qian 等ICCV 2019 · 被引用 191 次
相关 Paper
- Cost-Effective In-Context Learning for Entity Resolution: A Design Space ExplorationMeihao Fan, Xiaoyue Han, Ju Fan, Chengliang Chai 等ICDE 2024 · 被引用 40 次
- Can we trust LLM Self-Explanations for Entity Resolution?Tommaso Teofili, Donatella Firmani, Nick Koudas, Paolo Merialdo 等VLDB 2026 · 被引用 2 次
- GraphER: Token-Centric Entity Resolution with Graph Convolutional Neural NetworksBing Li, Wei Wang, Yifang Sun, Linhan Zhang 等AAAI 2020 · 被引用 48 次
- GER-LLM: Efficient and Effective Geospatial Entity Resolution with Large Language ModelHaojia Zhu, Zhicheng Li, Jiahui JinEMNLP 2025 · 被引用 1 次
- Optimized Algorithms for Text Clustering with LLM-Generated ConstraintsChaoqi Jia, Weihong Wu, Longkun Guo, Zhigang Lu 等AAAI 2026
