MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
Yudong Hu, Yueju Han, Rui Sun, Jinke Ren
摘要
Capsule Network (CapsNet) has demonstrated significant potential in visual recognition by capturing spatial relationships and part-whole hierarchies for learning equivariant feature representations. However, existing CapsNet and variants often rely on a single high-level feature map, overlooking the rich complementary information provided by multi-scale features. Furthermore, conventional feature fusion strategies, such as addition and concatenation, struggle to reconcile multi-scale feature discrepancies, leading to suboptimal classification performance. To address these limitations, we propose the Multi-Scale Patchify Capsule Network (MSPCaps), a novel architecture that integrates multi-scale feature learning and efficient capsule routing. Specifically, MSPCaps consists of three key components: a Multi-Scale ResNet Backbone (MSRB), a Patchify Capsule Layer (PatchifyCaps), and a Cross-Agreement Routing (CAR) block. First, the MSRB extracts diverse multi-scale feature representations from input images, preserving both fine-grained details and global contextual information. Second, the PatchifyCaps partitions these multi-scale features into primary capsules using a uniform patch size, equipping the model with the ability to learn from diverse receptive fields. Finally, the CAR block adaptively routes the multi-scale capsules by identifying crossscale prediction pairs with maximum agreement. Unlike the simple concatenation of multiple self-routing blocks, CAR ensures that only the most coherent capsules (best part-towhole pairs) contribute to the final voting. Our proposed MSPCaps achieves remarkable scalability and superior robustness, consistently surpassing multiple baseline methods in terms of classification accuracy, with configurations ranging from a highly efficient Tiny model (344.3K parameters) to a powerful Large model (10.9M parameters), highlighting its potential in advancing feature representation learning. The code is available at https://github.com/abdn-hyd/MSPCaps .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Capsules with Inverted Dot-Product Attention RoutingYao-Hung Hubert Tsai, Nitish Srivastava, Hanlin Goh, Ruslan SalakhutdinovICLR 2020 · 被引用 91 次
- OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and PruningXinyu Geng, Jiaming Wang, Jiawei Gong, Yuerong Xue 等CVPR 2024
- ParseCaps: An Interpretable Parsing Capsule Network for Medical Image DiagnosisXinyu Geng, Jiaming Wang, Xiaolin Huang, Fanglin Chen 等AAAI 2025
相关 Paper
- PT-CapsNet: A Novel Prediction-Tuning Capsule Network Suitable for Deeper ArchitecturesChenbin Pan, Senem VelipasalarICCV 2021 · 被引用 11 次
- Linguistically Routing Capsule Network for Out-of-distribution Visual Question AnsweringQingxing Cao, Wentao Wan, Keze Wang, Xiaodan Liang 等ICCV 2021 · 被引用 16 次
- Pyramid Architecture for Multi-Scale Processing in Point Cloud SegmentationDong Nie, Rui Lan, Ling Wang, Xiaofeng RenCVPR 2022 · 被引用 38 次
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 被引用 339 次
- QuadTreeCapsule: QuadTree Capsules for Deep Regression TrackingDing Ma, Xiangqian WuACM MM 2022 · 被引用 2 次
