Patch-level Contrastive Learning via Positional Query for Visual Pre-training
Shaofeng Zhang, Qiang Zhou, Zhibin Wang, Fan Wang, Junchi Yan
摘要
Dense contrastive learning (DCL) has been recently explored for learning localized information for dense prediction tasks (e.g., detection and segmentation). It still suffers the difficulty of mining pixels/patches correspondence between two views. A simple way is inputting the same view twice and aligning the pixel/patch representation. However, it would reduce the variance of inputs, and hurts the performance. We propose a plug-in method PQCL (Positional Query for patch-level Contrastive Learning), which allows performing patch-level contrasts between two views with exact patch correspondence. Besides, by using positional queries, PQCL increases the variance of inputs, to enhance training. We apply PQCL to popular transformer-based CL frameworks (DINO and iBOT, and evaluate them on classification, detection and segmentation tasks, where our method obtains stable improvements, especially for dense tasks. It achieves new state-of-the-art in most settings. Code is available at https://github. com/Sherrylone/Query_Contrastive .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Continuous-Multiple Image Outpainting in One-Step via Positional Query and A Diffusion-based ApproachShaofeng Zhang, Jinfa Huang, Qiang Zhou, Zhibin Wang 等ICLR 2024 · 被引用 24 次
- Cross-view Masked Diffusion Transformers for Person Image SynthesisTrung X. Pham, Kang Zhang, Chang D. YooICML 2024 · 被引用 12 次
- Towards Effective Usage of Human-Centric Priors in Diffusion Models for Text-based Human Image GenerationJunyan Wang, Zhenhong Sun, Zhiyu Tan, Xuanbai Chen 等CVPR 2024 · 被引用 8 次
- Hierarchical Object-Aware Dual-Level Contrastive Learning for Domain Generalized Stereo MatchingYikun Miao, Meiqing Wu, Siew Kei Lam, Changsheng Li 等NeurIPS 2024 · 被引用 5 次
- SaCo Loss: Sample-Wise Affinity Consistency for Vision-Language Pre-TrainingSitong Wu, Haoru Tan, Zhuotao Tian, Yukang Chen 等CVPR 2024 · 被引用 5 次
它引用的顶会 Paper27
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
相关 Paper
- Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation LearningShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023 · 被引用 8 次
- CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation LearningShaofeng Zhang, Qiang Zhou, Sitong Wu, Haoru Tan 等ICLR 2025
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong 等CVPR 2021
- DisCo DETR: Distance-aware Multi-view Contrastive Learning for DETR Pre-trainingChao Ouyang, Yuyang Bai, Jun Zhang, Tianlu Gao 等AAAI 2026
- Calibration-based Dual Prototypical Contrastive Learning Approach for Domain Generalization Semantic SegmentationMuxin Liao, Shishun Tian, Yuhang Zhang, Guoguang Hua 等ACM MM 2023 · 被引用 14 次
