CrossETR: A Semantic-Driven Framework for Entity Matching Across Images and Graph
Qin Yuan, Zhenyu Wen, Jiaxu Qian, Ye Yuan, Guoren Wang
Abstract
Entity matching (EM) aims to identify whether two entities from different data sources refer to the same real-world entity. Most existing cross-modal EM assume that images have simple scenes containing few objects, or do not fully consider the cross-modal knowledge associated with entities. To support more practical application scenarios such as multi-modal knowledge graph integration and visual question answering in data lakes, we introduce our problem of semantic-driven EM across graph and images in this paper. Current semantically matching solutions over cross-modal data face the obstacle of low training efficiency, since their time complexity quadratically grows with the number of entities. To alleviate this issue, we present a novel framework (namely CrossETR) that follows an exploration-then-refinement paradigm. Firstly, a candidate exploration policy is proposed to boost the training efficiency. It explores candidate pairs according to entity correlations and captures structural semantics by adaptive sampling the most informative neighborhood subgraphs. Secondly, the cross-modal entity representations are refined to break modality heterogeneity to support unsupervised matching prediction. Extensive experimental evaluations on three publicly available benchmarks demonstrate the superiority of CrossETR over state-of-the-art approaches in terms of effectiveness and efficiency. Furthermore, a case study highlights that our proposed semantic-driven EM is promising to improve the performance of downstream tasks such as multi-modal knowledge graph integration.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 78b85f79-79c3-4dc0-8312-205e45e79c19Related papers
- CrossEM: A Prompt Tuning Framework for Cross-Modal Entity MatchingQin Yuan, Ye Yuan, Zhenyu Wen, Chi Chen et al.ICDE 2025 · 1 citation
- CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity MatchingShiqi Zhang, Weixin Zeng, Ziheng Zhang, Wenzhe Hou et al.SIGIR 2026
- MEAformer: Multi-modal Entity Alignment Transformer for Meta Modality HybridZhuo Chen, Jiaoyan Chen, Wen Zhang, Lingbing Guo et al.ACM MM 2023 · 66 citations
- Unsupervised Entity Alignment for Temporal Knowledge GraphsXiaoze Liu, Junyang Wu, Tianyi Li, Lu Chen et al.WWW 2023 · 56 citations
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu et al.ACM MM 2020 · 35 citations
