Tran-GCN: Multi-label Pattern Image Retrieval via Transformer Driven Graph Convolutional Network
Ying Li, Chunming Guan, Rui Cai, Erwan Ye, Ding Yuxiang, Jiaquan Gao
Abstract
Pattern images are artificially designed images that possess distinctiveness in their elements, styles, and arrangements. With the ever-growing number of pattern images, pattern image retrieval emerges as a promising technique with significant potential for commercial and industrial applications, such as fashion and home decoration, facilitating rapid identification of preferred print patterns by users. The main purpose of multi-label pattern image retrieval is to effectively represent and match images with their corresponding labels. Compared to conventional image retrieval, multi-label pattern image retrieval faces greater challenges due to the richer semantic information contained within the abstract print patterns and the complex relationships between multiple labels. To tackle these challenges, we propose a model specifically designed for multi-label pattern image retrieval, called Tran-GCN. Our proposed model is built upon a Transformer-based autoregressive architecture, which leverages image information to guide the exploration of correlations between different labels through the textual modality. By utilizing this correlation information, we construct a graph convolutional network (GCN) model to further enhance the correlations between image and label representations. To be more specific, our Tran-GCN model utilizes a cross-modal attention mechanisms at each layer to effectively aggregate visual features from the input image and update label semantics through residual connections. The GCN module is updated based on the correlation between textual features, as represented in a relationship matrix. Extensive experiments on two widely used public visual benchmarks, MS-COCO and NUS-WIDE, as well as a multi-label pattern image dataset, Pattern 2, consistently demonstrate the ability of our proposed Tran-GCN model for general use and its superior performance in multi-label pattern image retrieval tasks as well.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 456360bd-dc92-4871-8792-2997bc5b2cd5Cited by top-tier papers1
Ask how each one uses itRelated papers
- Multi-label Pattern Image Retrieval via Attention Mechanism Driven Graph Convolutional NetworkYing Li, Hongwei Zhou, Yeyu Yin, Jiaquan GaoACM MM 2021 · 15 citations
- Modular Graph Transformer Networks for Multi-Label Image ClassificationHoang D. Nguyen, Xuan-Son Vu, Duc-Trong LeAAAI 2021 · 78 citations
- Cross-Modality Attention with Semantic Graph Embedding for Multi-Label ClassificationRenchun You, Zhiyao Guo, Lei Cui, Xiang Long et al.AAAI 2020 · 221 citations
- Two-Stream Transformer for Multi-Label Image ClassificationXuelin Zhu, Jiuxin Cao, Jiawei Ge, Weijia Liu et al.ACM MM 2022 · 37 citations
- Learning from Different text-image Pairs: A Relation-enhanced Graph Convolutional Network for Multimodal NERFei Zhao, Chunhui Li, Zhen Wu, Shangyu Xing et al.ACM MM 2022 · 59 citations
