Tran-GCN: Multi-label Pattern Image Retrieval via Transformer Driven Graph Convolutional Network
Ying Li, Chunming Guan, Rui Cai, Erwan Ye, Ding Yuxiang, Jiaquan Gao
摘要
Pattern images are artificially designed images that possess distinctiveness in their elements, styles, and arrangements. With the ever-growing number of pattern images, pattern image retrieval emerges as a promising technique with significant potential for commercial and industrial applications, such as fashion and home decoration, facilitating rapid identification of preferred print patterns by users. The main purpose of multi-label pattern image retrieval is to effectively represent and match images with their corresponding labels. Compared to conventional image retrieval, multi-label pattern image retrieval faces greater challenges due to the richer semantic information contained within the abstract print patterns and the complex relationships between multiple labels. To tackle these challenges, we propose a model specifically designed for multi-label pattern image retrieval, called Tran-GCN. Our proposed model is built upon a Transformer-based autoregressive architecture, which leverages image information to guide the exploration of correlations between different labels through the textual modality. By utilizing this correlation information, we construct a graph convolutional network (GCN) model to further enhance the correlations between image and label representations. To be more specific, our Tran-GCN model utilizes a cross-modal attention mechanisms at each layer to effectively aggregate visual features from the input image and update label semantics through residual connections. The GCN module is updated based on the correlation between textual features, as represented in a relationship matrix. Extensive experiments on two widely used public visual benchmarks, MS-COCO and NUS-WIDE, as well as a multi-label pattern image dataset, Pattern 2, consistently demonstrate the ability of our proposed Tran-GCN model for general use and its superior performance in multi-label pattern image retrieval tasks as well.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- Multi-label Pattern Image Retrieval via Attention Mechanism Driven Graph Convolutional NetworkYing Li, Hongwei Zhou, Yeyu Yin, Jiaquan GaoACM MM 2021 · 被引用 15 次
- Modular Graph Transformer Networks for Multi-Label Image ClassificationHoang D. Nguyen, Xuan-Son Vu, Duc-Trong LeAAAI 2021 · 被引用 78 次
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long 等AAAI 2020 · 被引用 221 次
- Two-Stream Transformer for Multi-Label Image ClassificationXuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu 等ACM MM 2022 · 被引用 37 次
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing 等ACM MM 2022 · 被引用 59 次
