Unsupervised Matching of Data and Text
Naser Ahmadi, Hansjorg Sand, Paolo Papotti
摘要
Entity resolution is a widely studied problem with several proposals to match records across relations. Matching textual content is a widespread task in many applications, such as question answering and search. While recent methods achieve promising results for these two tasks, there is no clear solution for the more general problem of matching textual content and structured data. We introduce a framework that supports this new task in an unsupervised setting for any pair of corpora, being relational tables or text documents. Our method builds a fine-grained graph over the content of the corpora and derives word embeddings to represent the objects to match in a low dimensional space. The learned representation enables effective and efficient matching at different granularity, from relational tuples to text sentences and paragraphs. Our flexible framework can exploit pre-trained resources, but, differently from other solutions, it does not depends on their existence and achieves better quality performance in matching content when the vocab-ulary is domain specific. We also introduce optimizations in the graph creation process with an “expand and compress” approach that first identifies new valid relationships across elements, to improve matching, and then prunes nodes and edges, to reduce the graph size. Experiments on real use cases and public datasets show that our framework produces embeddings that outperform word embeddings and fine-tuned language models both in results' quality and in execution times.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PromptEM: Prompt-tuning for Low-resource Generalized Entity MatchingPengfei Wang, Xiaocan Zeng, Lu Chen, Fan Ye 等VLDB 2023 · 被引用 39 次
- ELEET: Efficient Learned Query Execution over Text and TablesMatthias Urban, Carsten BinnigVLDB 2024 · 被引用 15 次
- Cross Modal Data Discovery over Structured and Unstructured Data LakesMohamed Y. Eltabakh, Mayuresh Kunjir, Ahmed K. Elmagarmid, Mohammad Shahmeer AhmadVLDB 2023 · 被引用 11 次
- "The Data Says Otherwise" - Towards Automated Fact-checking and Communication of Data ClaimsYu Fu, Shunan Guo, Jane Hoffswell, Victor S. Bursztyn 等UIST 2024 · 被引用 6 次
- QueryArtisan: Generating Data Manipulation Codes for Ad-hoc Analysis in Data LakesXiu Tang, Wenhao Liu, Sai Wu, Chang Yao 等VLDB 2025 · 被引用 4 次
它引用的顶会 Paper12
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- TURL: Table Understanding through Representation LearningXiang Deng, Huan Sun, Alyssa Lees, You Wu 等VLDB 2021 · 被引用 2,406 次
- TabFact: A Large-scale Dataset for Table-based Fact VerificationWenhu Chen, Hongmin Wang, Jianshu Chen, Yunkai Zhang 等ICLR 2020 · 被引用 674 次
- Deep Entity Matching with Pre-Trained Language ModelsYuliang Li, Jinfeng Li, Yoshihiko Suhara, AnHai Doan 等VLDB 2021 · 被引用 484 次
- TaBERT: Pretraining for Joint Understanding of Textual and Tabular DataPengcheng Yin, Graham Neubig, Wen-tau Yih, Sebastian RiedelACL 2020 · 被引用 417 次
相关 Paper
- Creating Embeddings of Heterogeneous Relational Datasets for Data Integration TasksRiccardo Cappuzzo, Paolo Papotti, Saravanan ThirumuruganathanSIGMOD 2020 · 被引用 139 次
- MultiEM: Efficient and Effective Unsupervised Multi-Table Entity MatchingXiaocan Zeng, Pengfei Wang, Yuren Mao, Lu Chen 等ICDE 2024 · 被引用 5 次
- Unsupervised Entity Alignment for Temporal Knowledge GraphsXiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen 等WWW 2023 · 被引用 56 次
- GraphER: Token-Centric Entity Resolution with Graph Convolutional Neural NetworksBing Li, Wei Wang, Yifang Sun, Linhan Zhang 等AAAI 2020 · 被引用 48 次
- Pre-trained Embeddings for Entity Resolution: An Experimental AnalysisAlexandros Zeakis, George Papadakis, Dimitrios Skoutas, Manolis KoubarakisVLDB 2023 · 被引用 63 次
