Multi-Modal Knowledge Hypergraph for Diverse Image Retrieval
Yawen Zeng, Qin Jin, Tengfei Bao, Wenfeng Li
摘要
The task of keyword-based diverse image retrieval has received considerable attention due to its wide demand in real-world scenarios. Existing methods either rely on a multi-stage re-ranking strategy based on human design to diversify results, or extend sub-semantics via an implicit generator, which either relies on manual labor or lacks explainability. To learn more diverse and explainable representations, we capture sub-semantics in an explicit manner by leveraging the multi-modal knowledge graph (MMKG) that contains richer entities and relations. However, the huge domain gap between the off-the-shelf MMKG and retrieval datasets, as well as the semantic gap between images and texts, make the fusion of MMKG difficult. In this paper, we pioneer a degree-free hypergraph solution that models many-to-many relations to address the challenge of heterogeneous sources and heterogeneous modalities. Specifically, a hyperlink-based solution, Multi-Modal Knowledge Hyper Graph (MKHG) is proposed, which bridges heterogeneous data via various hyperlinks to diversify sub-semantics. Among them, a hypergraph construction module first customizes various hyperedges to link the heterogeneous MMKG and retrieval databases. A multi-modal instance bagging module then explicitly selects instances to diversify the semantics. Meanwhile, a diverse concept aggregator flexibly adapts key sub-semantics. Finally, several losses are adopted to optimize the semantic space. Extensive experiments on two real-world datasets have well verified the effectiveness and explainability of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- UniGraph2: Learning a Unified Embedding Space to Bind Multimodal GraphsYufei He, Yuan Sui, Xiaoxin He, Yue Liu 等WWW 2025 · 被引用 37 次
- Towards Semantic Consistency: Dirichlet Energy Driven Robust Multi-Modal Entity AlignmentYuanyi Wang, Haifeng Sun, Jiabo Wang, Jingyu Wang 等ICDE 2024 · 被引用 13 次
- Mitigating Modality Bias in Multi-modal Entity Alignment from a Causal PerspectiveTaoyu Su, Jiawei Sheng, Duohe Ma, Xiaodong Li 等SIGIR 2025 · 被引用 4 次
- Hypergraph-based Zero-shot Multi-modal Product Attribute Value ExtractionJiazhen Hu, Jiaying Gong, Hongda Shen, Hoda EldardiryWWW 2025 · 被引用 4 次
- Transferable Hypergraph Attack via Injecting Nodes into Pivotal HyperedgesMeixia He, Peican Zhu, Le Cheng, Yangming Guo 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Fine-grained Cross-modal Alignment Network for Text-Video RetrievalNing Han, Jingjing Chen, Guangyi Xiao, Hao Zhang 等ACM MM 2021 · 被引用 47 次
- Adversarial Video Moment Retrieval by Jointly Modeling Ranking and LocalizationDa Cao, Yawen Zeng, Xiaochi Wei, Liqiang Nie 等ACM MM 2020 · 被引用 36 次
- Interpretable Embedding for Ad-Hoc Video SearchJiaxin Wu, Chong-Wah NgoACM MM 2020 · 被引用 33 次
相关 Paper
- HL-CMR: Hypergraph Learning for Cross-Modal RetrievalGuohui Ding, Jing Li, Yimin Xu, Rui ZhouWWW 2026
- Multi-Granularity Multi-Modal Knowledge Graph Representation Learning via Subgraph-Aware Adaptive Fusion and Hierarchical Relation ModelingPeining Li, Meiyu Liang, Wei Huang, Junping Du 等WWW 2026
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 被引用 29 次
- MMKGR: Multi-hop Multi-modal Knowledge Graph ReasoningShangfei Zheng, Weiqing Wang, Jianfeng Qu, Hongzhi Yin 等ICDE 2023 · 被引用 40 次
- Hybrid Transformer with Multi-level Fusion for Multimodal Knowledge Graph CompletionXiang Chen, Ningyu Zhang, Lei Li, Shumin Deng 等SIGIR 2022 · 被引用 227 次
