Intermediate Domain Alignment and Morphology Analogy for Patent-Product Image Retrieval
Haifan Gong, Xuanye Zhang, Ruifei Zhang, Yun Su, Zhuo Li, Yuhao Du, Anningzhe Gao, Xiang Wan, Haofeng Li
Abstract
Recent advances in artificial intelligence have significantly impacted image retrieval tasks, yet Patent-Product Image Retrieval (PPIR) has received limited attention. PPIR, which retrieves patent images based on product images to identify potential infringements, presents unique challenges: (1) both product and patent images often contain numerous categories of artificial objects, but models pre-trained on standard datasets exhibit limited discriminative power to recognize some of those unseen objects; and (2) the significant domain gap between binary patent line drawings and colorful RGB product images further complicates similarity comparisons for product-patent pairs. To address these challenges, we formulate it as an open-set image retrieval task and introduce a comprehensive Patent-Product Image Retrieval Dataset (PPIRD) including a test set with 439 product-patent pairs, a retrieval pool of 727,921 patents, and an unlabeled pre-training set of 3,799,695 images. We further propose a novel Intermediate Domain Alignment and Morphology Analogy (IDAMA) strategy. IDAMA maps both image types to an intermediate sketch domain using edge detection to minimize the domain discrepancy, and employs a Morphology Analogy Filter to select discriminative patent images based on visual features via analogical reasoning. Extensive experiments on PPIRD demonstrate that IDAMA significantly outperforms baseline methods (+7.58 mAR) and offers valuable insights into domain mapping and representation learning for PPIR. (The PPIRD dataset is available at: https://loslorien.github.io/idama-project/)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Image BERT Pre-training with Online TokenizerJinghao Zhou, Chen Wei, Huiyu Wang, Wei Shen et al.ICLR 2022 · 287 citations
Related papers
- DiDA: Disambiguated Domain Alignment for Cross-Domain Retrieval with Partial LabelsHaoran Liu, Ying Ma, Ming Yan, Yingke Chen et al.AAAI 2024 · 13 citations
- MCID: Multi-aspect Copyright Infringement Detection for Generated ImagesChuanwei Huang, Zexi Jia, Hongyan Fei, Yeshuang Zhu et al.ICCV 2025 · 1 citation
- PatentLMM: Large Multimodal Model for Generating Descriptions for Patent FiguresShreya Shukla, Nakul Sharma, Manish Gupta, Anand MishraAAAI 2025 · 6 citations
- Is It Novel and Why? Fine-Grained Patent Novelty Prediction Based on Passage RetrievalValentin Knappich, Anna Hätty, Simon Razniewski, Annemarie FriedrichSIGIR 2026 · 1 citation
- RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image RetrievalYijiang Li, Kunal Kotian, Ali Marjaninejad, Meir Friedenberg et al.CVPR 2026
