CARAFE: Content-Aware ReAssembly of FEatures
Jiaqi Wang, Kai Chen, Rui Xu, Ziwei Liu, Chen Change Loy, Dahua Lin
Abstract
Feature upsampling is a key operation in a number of modern convolutional network architectures, e.g. feature pyramids. Its design is critical for dense prediction tasks such as object detection and semantic/instance segmentation. In this work, we propose Content-Aware ReAssembly of FEatures (CARAFE), a universal, lightweight and highly effective operator to fulfill this goal. CARAFE has several appealing properties: (1) Large field of view. Unlike previous works (e.g. bilinear interpolation) that only exploit subpixel neighborhood, CARAFE can aggregate contextual information within a large receptive field. (2) Content-aware handling. Instead of using a fixed kernel for all samples (e.g. deconvolution), CARAFE enables instance-specific content-aware handling, which generates adaptive kernels on-the-fly. (3) Lightweight and fast to compute. CARAFE introduces little computational overhead and can be readily integrated into modern network architectures. We conduct comprehensive evaluations on standard benchmarks in object detection, instance/semantic segmentation and inpainting. CARAFE shows consistent and substantial gains across all the tasks (1.2% AP, 1.3% AP, 1.8% mIoU, 1.1dB respectively) with negligible computational overhead. It has great potential to serve as a strong building block for future research. Code and models are available at https: //github.com/open-mmlab/mmdetection .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0a541fd8-1271-43c5-96ff-4e297ab66563Cited by top-tier papers37
- FeatUp: A Model-Agnostic Framework for Features at Any ResolutionStephanie Fu, Mark Hamilton, Laura E. Brandt, Axel Feldmann et al.ICLR 2024 · 117 citations
- V3Det: Vast Vocabulary Visual Detection DatasetJiaqi Wang, Pan Zhang, Tao Chu, Yuhang Cao et al.ICCV 2023 · 86 citations
- Co-advise: Cross Inductive Bias DistillationSucheng Ren, Zhengqi Gao, Tianyu Hua, Zihui Xue et al.CVPR 2022 · 50 citations
- Rank & Sort Loss for Object Detection and Instance SegmentationKemal Oksuz, Baris Can Cam, Emre Akbas, Sinan KalkanICCV 2021 · 49 citations
- TVConv: Efficient Translation Variant Convolution for Layout-aware Visual ProcessingJierun Chen, Tianlang He, Weipeng Zhuo, Li Ma et al.CVPR 2022 · 39 citations
Builds on1
Related papers
- Learning to Upsample by Learning to SampleWenze Liu, Hao Lu, Hongtao Fu, Zhiguo CaoICCV 2023 · 518 citations
- LDA-AQU: Adaptive Query-guided Upsampling via Local Deformable AttentionZewen Du, Zhenjiang Hu, Guiyu Zhao, Ying Jin et al.ACM MM 2024 · 4 citations
- SAPA: Similarity-Aware Point Affiliation for Feature UpsamplingHao Lu, Wenze Liu, Zixuan Ye, Hongtao Fu et al.NeurIPS 2022 · 96 citations
- UPLiFT: Efficient Pixel-Dense Feature Upsampling with Local AttendersMatthew Walmer, Saksham Suri, Anirud Aggarwal, Abhinav ShrivastavaCVPR 2026 · 2 citations
- LPSNet: A Lightweight Solution for Fast Panoptic SegmentationWeixiang Hong, Qingpei Guo, Wei Zhang, Jingdong Chen et al.CVPR 2021
