Patch-level Contrastive Learning via Positional Query for Visual Pre-training
Shaofeng Zhang, Qiang Zhou, Zhibin Wang, Fan Wang, Junchi Yan
Abstract
Dense contrastive learning (DCL) has been recently explored for learning localized information for dense prediction tasks (e.g., detection and segmentation). It still suffers the difficulty of mining pixels/patches correspondence between two views. A simple way is inputting the same view twice and aligning the pixel/patch representation. However, it would reduce the variance of inputs, and hurts the performance. We propose a plug-in method PQCL (Positional Query for patch-level Contrastive Learning), which allows performing patch-level contrasts between two views with exact patch correspondence. Besides, by using positional queries, PQCL increases the variance of inputs, to enhance training. We apply PQCL to popular transformer-based CL frameworks (DINO and iBOT, and evaluate them on classification, detection and segmentation tasks, where our method obtains stable improvements, especially for dense tasks. It achieves new state-of-the-art in most settings. Code is available at https://github. com/Sherrylone/Query_Contrastive .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based ApproachShaofeng Zhang, Jinfa Huang, Qiang Zhou, Zhibin Wang et al.ICLR 2024 · 24 citations
- Cross-view Masked Diffusion Transformers for Person Image SynthesisTrung X. Pham, Kang Zhang, Chang D. YooICML 2024 · 12 citations
- Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image GenerationJunyan Wang, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen et al.CVPR 2024 · 8 citations
- Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo MatchingYikun Miao, Meiqing Wu, Siew Kei Lam, Changsheng Li et al.NeurIPS 2024 · 5 citations
- SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingSitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen et al.CVPR 2024 · 5 citations
Builds on27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
Related papers
- Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation LearningShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023 · 8 citations
- CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation LearningShaofeng Zhang, Qiang Zhou, Sitong Wu, Haoru Tan et al.ICLR 2025
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong et al.CVPR 2021
- DisCo DETR: Distance-aware Multi-view Contrastive Learning for DETR Pre-trainingChao Ouyang, Yuyang Bai, Jun Zhang, Tianlu Gao et al.AAAI 2026
- Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic SegmentationMuxin Liao, Shishun Tian, Yuhang Zhang, Guoguang Hua et al.ACM MM 2023 · 14 citations
