CR2PQ: Continuous Relative Rotary Positional Query for Dense Visual Representation Learning
Shaofeng Zhang, Qiang Zhou, Sitong Wu, Haoru Tan, Zhibin Wang, Jinfa Huang, Junchi Yan
摘要
Dense visual representation learning (DRL) shows promise for learning localized information in dense prediction tasks, but struggles with establishing pixel/patch correspondence across different views (cross-contrasting). Existing methods primarily rely on self-contrasting the same view with variations, limiting input variance and hindering downstream performance. This paper delves into the mechanisms of selfcontrasting and cross-contrasting, identifying the crux of the issue: transforming discrete positional embeddings to continuous representations. To address the correspondence problem, we propose a Continuous Relative Rotary Positional Query (CR2PQ), enabling patch-level representation learning. Our extensive experiments on standard datasets demonstrate state-of-the-art (SOTA) results. Compared to the previous SOTA method (PQCL), our approach achieves significant improvements on COCO: with 300 epochs of pretraining, CR2PQ obtains 3.4% mAP bb and 2.1% mAP mk improvements for detection and segmentation tasks, respectively. Furthermore, CR2PQ exhibits faster convergence, achieving 10.4% mAP bb and 7.9% mAP mk improvements over SOTA with just 40 epochs of pretraining.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Mixture-of-Scores: Robust Image-Text Data Valuation via Three Lines of CodeSitong Wu, Haoru Tan, Yukang Chen, Shaofeng Zhang 等ICCV 2025 · 被引用 4 次
- Towards More Diverse and Challenging Pre-Training for Point Cloud Learning: Self-Supervised Cross Reconstruction with Decoupled ViewsXiangdong Zhang, Shaofeng Zhang, Junchi YanICCV 2025 · 被引用 4 次
- SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMsKoonting Yip, Qiyan Zhao, Wenhao Yu, Liangyu Yuan 等CVPR 2026 · 被引用 3 次
- PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language ModelsYuanhao Su, Shaofeng Zhang, Xiaosong Jia, Qi FanCVPR 2026 · 被引用 1 次
- Learning to See through Illumination Extremes with Event Streaming in Multimodal Large Language ModelsBaoheng Zhang, Jiahui Liu, Gui Zhao, Weizhou Zhang 等CVPR 2026
它引用的顶会 Paper34
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- Patch-level Contrastive Learning via Positional Query for Visual Pre-trainingShaofeng Zhang, Qiang Zhou, Zhibin Wang, Fan Wang 等ICML 2023 · 被引用 22 次
- DetCo: Unsupervised Contrastive Learning for Object DetectionEnze Xie, Jian Ding, Wenhai Wang, Xiaohang Zhan 等ICCV 2021 · 被引用 364 次
- Dense Contrastive Learning for Self-Supervised Visual Pre-TrainingXinlong Wang, Rufeng Zhang, Chunhua Shen, Tao Kong 等CVPR 2021
- Patch-Level Contrasting without Patch Correspondence for Accurate and Dense Contrastive Representation LearningShaofeng Zhang, Feng Zhu, Rui Zhao, Junchi YanICLR 2023 · 被引用 8 次
- Propagate Yourself: Exploring Pixel-Level Consistency for Unsupervised Visual Representation LearningZhenda Xie, Yutong Lin, Zheng Zhang, Yue Cao 等CVPR 2021
