OverLoCK: An Overview-first-Look-Closely-next ConvNet with Context-Mixing Dynamic Kernels
Meng Lou, Yizhou Yu
摘要
Top-down attention plays a crucial role in the human vision system, wherein the brain initially obtains a rough overview of a scene to discover salient cues (i.e., overview first), followed by a more careful finer-grained examination (i.e., look closely next). However, modern ConvNets remain confined to a pyramid structure that successively downsamples the feature map for receptive field expansion, neglecting this crucial biomimetic principle. We present OverLoCK, the first pure ConvNet backbone architecture that explicitly incorporates a top-down attention mechanism. Unlike pyramid backbone networks, our design features a branched architecture with three synergistic subnetworks: 1) a Base-Net that encodes low/mid-level features; 2) a lightweight Overview-Net that generates dynamic top-down attention through coarse global context modeling (i.e., overview first); and 3) a robust Focus-Net that performs finer-grained perception guided by top-down attention (i.e., look closely next). To fully unleash the power of top-down attention, we further propose a novel context-mixing dynamic convolution (ContMix) that effectively models long-range dependencies while preserving inherent local inductive biases even when the input resolution increases, addressing critical limitations in existing convolutions. Our OverLoCK exhibits a notable performance improvement over existing methods. For instance, OverLoCK-T achieves a Top-1 accuracy of 84.2%, significantly surpassing ConvNeXt-B while using only around one-third of the FLOPs/parameters. On object detection, our OverLoCK-S clearly surpasses MogaNet-B by 1% in AP b . On semantic segmentation, our OverLoCK-T remarkably improves UniRepLKNet-T by 1.7% in mIoU. Code is publicly available at https://rb.gy/wit4jh.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- Vision Transformers Need More Than RegistersCheng Shi, Yizhou Yu, Sibei YangCVPR 2026 · 被引用 17 次
- Frequency-Dynamic Attention Modulation for Dense PredictionLinwei Chen, Lin Gu, Ying FuICCV 2025 · 被引用 13 次
- D²-VPR: A Parameter-efficient Visual-foundation-model-based Visual Place Recognition Method via Knowledge Distillation and Deformable AggregationZheyuan Zhang, Jiwei Zhang, Boyu Zhou, Linzhimeng Duan 等AAAI 2026 · 被引用 2 次
- Vision Transformers Are Circulant Attention LearnersDongchen Han, Tianyu Li, Ziyi Wang, Gao HuangAAAI 2026 · 被引用 2 次
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 被引用 1 次
它引用的顶会 Paper34
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan 等ICCV 2021 · 被引用 4,909 次
相关 Paper
- TransNeXt: Robust Foveal Visual Perception for Vision TransformersDai ShiCVPR 2024 · 被引用 313 次
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu 等NeurIPS 2022 · 被引用 1,385 次
- Contextual Convolutional NetworksShuxian Liang, Xu Shen, Tongliang Liu, Xian-Sheng HuaICLR 2023 · 被引用 330 次
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 被引用 494 次
- MOAT: Alternating Mobile Convolution and Attention Brings Strong Vision ModelsChenglin Yang, Siyuan Qiao, Qihang Yu, Xiaoding Yuan 等ICLR 2023 · 被引用 22 次
