Solving Mixed-Modal Jigsaw Puzzle for Fine-Grained Sketch-Based Image Retrieval
Kaiyue Pang, Yongxin Yang, Timothy M. Hospedales, Tao Xiang, Yi-Zhe Song
Abstract
ImageNet pre-training has long been considered crucial by the fine-grained sketch-based image retrieval (FG-SBIR) community due to the lack of large sketch-photo paired datasets for FG-SBIR training. In this paper, we propose a self-supervised alternative for representation pre-training. Specifically, we consider the jigsaw puzzle game of recomposing images from shuffled parts. We identify two key facets of jigsaw task design that are required for effective FG-SBIR pre-training. The first is formulating the puzzle in a mixed-modality fashion. Second we show that framing the optimisation as permutation matrix inference via Sinkhorn iterations is more effective than the common classifier formulation of Jigsaw self-supervision. Experiments show that this self-supervised pre-training strategy significantly outperforms the standard ImageNet-based pipeline across all four product-level FG-SBIR benchmarks. Interestingly it also leads to improved cross-category generalisation across both pre-train/fine-tune and fine-tune/testing stages.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f09e8439-bd8b-4344-8b91-b857dc90483cCited by top-tier papers29
- Sketch3T: Test-Time Training for Zero-Shot SBIRAneeshan Sain, Ayan Kumar Bhunia, Vaishnav Potlapalli, Pinaki Nath Chowdhury et al.CVPR 2022 · 55 citations
- Sketching without Worrying: Noise-Tolerant Sketch-Based Image RetrievalAyan Kumar Bhunia, Subhadeep Koley, Abdullah Faiz Ur Rahman Khilji, Aneeshan Sain et al.CVPR 2022 · 53 citations
- A-Net: Learning Attribute-Aware Hash Codes for Large-Scale Fine-Grained Image RetrievalXiu-Shen Wei, Yang Shen, Xuhao Sun, Han-Jia Ye et al.NeurIPS 2021 · 48 citations
- ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style SimilarityDan Ruta, Saeid Motiian, Baldo Faieta, Zhe Lin et al.ICCV 2021 · 44 citations
- Partially Does It: Towards Scene-Level FG-SBIR with Partial InputPinaki Nath Chowdhury, Ayan Kumar Bhunia, Viswanatha Reddy Gajjala, Aneeshan Sain et al.CVPR 2022 · 28 citations
Builds on1
Related papers
- Photo Pre-Training, But for SketchKe Li, Kaiyue Pang, Yi-Zhe SongCVPR 2023
- Self-Supervised Learning of Pretext-Invariant RepresentationsIshan Misra, Laurens van der MaatenCVPR 2020
- Sketch Less for More: On-the-Fly Fine-Grained Sketch-Based Image RetrievalAyan Kumar Bhunia, Yongxin Yang, Timothy M. Hospedales, Tao Xiang et al.CVPR 2020
- Exploiting Unlabelled Photos for Stronger Fine-Grained SBIRAneeshan Sain, Ayan Kumar Bhunia, Subhadeep Koley, Pinaki Nath Chowdhury et al.CVPR 2023
- Semantic-Aware Knowledge Preservation for Zero-Shot Sketch-Based Image RetrievalQing Liu, Lingxi Xie, Huiyu Wang, Alan L. YuilleICCV 2019 · 126 citations
