Sudowoodo: Contrastive Self-supervised Learning for Multi-purpose Data Integration and Preparation
Runhui Wang, Yuliang Li, Jin Wang
摘要
Machine learning (ML) is playing an increasingly important role in data management tasks, particularly in Data Integration and Preparation (DI&P). The success of ML-based approaches, however, heavily relies on the availability of largescale, high-quality labeled datasets for different tasks. Moreover, the wide variety of DI&P tasks and pipelines oftentimes requires customizing ML solutions at a significant cost for model engineering and experimentation. These factors inevitably hold back the adoption of ML-based approaches to new domains and tasks.
In this paper, we propose Sudowoodo, a multi-purpose DI&P framework based on contrastive representation learning. Sudowoodo features a unified, matching-based problem definition capturing a wide range of DI&P tasks including Entity Matching (EM) in data integration, error correction in data cleaning, semantic type detection in data discovery, and more. Contrastive learning enables Sudowoodo to learn similarity-aware data representations from a large corpus of data items (e.g., entity entries, table columns) without using any labels. The learned representations can later be either directly used or facilitate fine-tuning with only a few labels to support different DI&P tasks. Our experiment results show that Sudowoodo achieves multiple state-of-the-art results on different levels of supervision and outperforms previous best specialized blocking or matching solutions for EM. Sudowoodo also achieves promising results in data cleaning and column matching tasks showing its versatility in DI&P applications.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Watchog: A Light-weight Contrastive Learning based Framework for Column AnnotationZhengjie Miao, Jin WangSIGMOD 2024 · 被引用 14 次
- Blocker and Matcher Can Mutually Benefit: A Co-Learning Framework for Low-Resource Entity ResolutionShiwen Wu, Qiyu Wu, Honghua Dong, Wen Hua 等VLDB 2024 · 被引用 10 次
- Data Imputation with Limited Data Redundancy Using Data LakesChenyu Yang, Yuyu Luo, Chuanxuan Cui, Ju Fan 等VLDB 2025 · 被引用 9 次
- KGLink: A Column Type Annotation Method that Combines Knowledge Graph and Pre-Trained Language ModelYubo Wang, Hao Xin, Lei ChenICDE 2024 · 被引用 7 次
- MultiEM: Efficient and Effective Unsupervised Multi-Table Entity MatchingXiaocan Zeng, Pengfei Wang, Yuren Mao, Lu Chen 等ICDE 2024 · 被引用 5 次
它引用的顶会 Paper19
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- SimCSE: Simple Contrastive Learning of Sentence EmbeddingsTianyu Gao, Xingcheng Yao, Danqi ChenEMNLP 2021 · 被引用 2,496 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
相关 Paper
- CrossEM: A Prompt Tuning Framework for Cross-Modal Entity MatchingQin Yuan, Ye Yuan, Zhenyu Wen, Chi Chen 等ICDE 2025 · 被引用 1 次
- Automating Entity Matching Model DevelopmentPei Wang, Weiling Zheng, Jiannan Wang, Jian PeiICDE 2021 · 被引用 13 次
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
- Bootstrapping Self-Improvement of Language Model Programs for Zero-Shot Schema MatchingNabeel Seedat, Mihaela van der SchaarICML 2025
- Guiding Cross-Modal Representations with MLLM Priors via Preference AlignmentPengfei Zhao, Rongbo Luan, Wei Zhang, Peng Wu 等NeurIPS 2025 · 被引用 3 次
