CoMCo: Consistency-Aware Multi-Agent Coordination for Zero-Shot Cross-Modal Entity Matching
Shiqi Zhang, Weixin Zeng, Ziheng Zhang, Wenzhe Hou, Weidong Xiao, Xiang Zhao
摘要
Entity Matching (EM) is a fundamental task in data integration, traditionally studied over structured data such as tables and knowledge graphs. In modern repositories, real-world objects are often represented across multiple modalities, including structured entities with symbolic attributes and visual entities in image-centric collections. This motivates cross-modal entity matching, which aims to identify visual-structured entity pairs that refer to the same object. Existing methods typically rely on pretrained vision-language models to compute entity pair similarity and derive correspondence via local ranking, which, however, can be unreliable given noisy and ambiguous cross-modal data and may produce globally inconsistent correspondences across related entities. Multimodal large language models (MLLMs) offer richer cross-modal cues for matching, but exhaustive MLLM reasoning over large candidate spaces is prohibitively expensive. To address these limitations, in this work, we propose øurq, a blackboard-based multi-agent framework for zero-shot cross-modal entity matching that performs iterative, self-correcting refinement by fusing multiple matching signals, explicitly regulating global consistency, and selectively invoking an MLLM only for hard cases. We further construct two new benchmarks from real-world visual and structured data. Extensive experiments show that øurq consistently outperforms competitive baselines, providing an effective solution for zero-shot cross-modal entity matching. We release our data and code at https://github.com/Q-17/CoMCo.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Connecting the Dots: Training-Free Visual Grounding via Agentic ReasoningLiqin Luo, Guangyao Chen, Xiawu Zheng, Yongxing Dai 等AAAI 2026
- CrossETR: A Semantic-Driven Framework for Entity Matching Across Images and GraphQin Yuan, Zhenyu Wen, Jiaxu Qian, Ye Yuan 等ICDE 2025 · 被引用 1 次
- CrossEM: A Prompt Tuning Framework for Cross-Modal Entity MatchingQin Yuan, Ye Yuan, Zhenyu Wen, Chi Chen 等ICDE 2025 · 被引用 1 次
- U-MERE: Unconstrained Multimodal Entity and Relation Extraction with Collaborative Modeling and Order-Sensitive OptimizationWei Jia, Li Jin, Kaiwen Wei, Yuying Shang 等ACM MM 2025
- I2CR: Intra- and Inter-modal Collaborative Reflections for Multimodal Entity LinkingZiyan Liu, Junwen Li, Kaiwen Li, Tong Ruan 等ACM MM 2025 · 被引用 2 次
