MSPCaps: A Multi-Scale Patchify Capsule Network with Cross-Agreement Routing for Visual Recognition
Yudong Hu, Yueju Han, Rui Sun, Jinke Ren
Abstract
Capsule Network (CapsNet) has demonstrated significant potential in visual recognition by capturing spatial relationships and part-whole hierarchies for learning equivariant feature representations. However, existing CapsNet and variants often rely on a single high-level feature map, overlooking the rich complementary information provided by multi-scale features. Furthermore, conventional feature fusion strategies, such as addition and concatenation, struggle to reconcile multi-scale feature discrepancies, leading to suboptimal classification performance. To address these limitations, we propose the Multi-Scale Patchify Capsule Network (MSPCaps), a novel architecture that integrates multi-scale feature learning and efficient capsule routing. Specifically, MSPCaps consists of three key components: a Multi-Scale ResNet Backbone (MSRB), a Patchify Capsule Layer (PatchifyCaps), and a Cross-Agreement Routing (CAR) block. First, the MSRB extracts diverse multi-scale feature representations from input images, preserving both fine-grained details and global contextual information. Second, the PatchifyCaps partitions these multi-scale features into primary capsules using a uniform patch size, equipping the model with the ability to learn from diverse receptive fields. Finally, the CAR block adaptively routes the multi-scale capsules by identifying crossscale prediction pairs with maximum agreement. Unlike the simple concatenation of multiple self-routing blocks, CAR ensures that only the most coherent capsules (best part-towhole pairs) contribute to the final voting. Our proposed MSPCaps achieves remarkable scalability and superior robustness, consistently surpassing multiple baseline methods in terms of classification accuracy, with configurations ranging from a highly efficient Tiny model (344.3K parameters) to a powerful Large model (10.9M parameters), highlighting its potential in advancing feature representation learning. The code is available at https://github.com/abdn-hyd/MSPCaps .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 53ae5961-a9f0-4cb1-ac09-10c3e2386cd7Builds on4
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Capsules with Inverted Dot-Product Attention RoutingYao-Hung Hubert Tsai, Nitish Srivastava, Hanlin Goh, Ruslan SalakhutdinovICLR 2020 · 91 citations
- OrthCaps: An Orthogonal CapsNet with Sparse Attention Routing and PruningXinyu Geng, Jiaming Wang, Jiawei Gong, Yuerong Xue et al.CVPR 2024
- ParseCaps: An Interpretable Parsing Capsule Network for Medical Image DiagnosisXinyu Geng, Jiaming Wang, Xiaolin Huang, Fanglin Chen et al.AAAI 2025
Related papers
- PT-CapsNet: A Novel Prediction-Tuning Capsule Network Suitable for Deeper ArchitecturesChenbin Pan, Senem VelipasalarICCV 2021 · 11 citations
- Linguistically Routing Capsule Network for Out-of-distribution Visual Question AnsweringQingxing Cao, Wentao Wan, Keze Wang, Xiaodan Liang et al.ICCV 2021 · 16 citations
- Pyramid Architecture for Multi-Scale Processing in Point Cloud SegmentationDong Nie, Rui Lan, Ling Wang, Xiaofeng RenCVPR 2022 · 38 citations
- MPViT: Multi-Path Vision Transformer for Dense PredictionYoungwan Lee, Jonghee Kim, Jeffrey Willette, Sung Ju HwangCVPR 2022 · 339 citations
- QuadTreeCapsule: QuadTree Capsules for Deep Regression TrackingDing Ma, Xiangqian WuACM MM 2022 · 2 citations
