Enabling Highly Efficient Capsule Networks Processing Through A PIM-Based Architecture Design
Xingyao Zhang, Shuaiwen Leon Song, Chenhao Xie, Jing Wang, Weigong Zhang, Xin Fu
Abstract
In recent years, the CNNs have achieved great successes in the image processing tasks, e.g., image recognition and object detection. Unfortunately, traditional CNN's classification is found to be easily misled by increasingly complex image features due to the usage of pooling operations, hence unable to preserve accurate position and pose information of the objects. To address this challenge, a novel neural network structure called Capsule Network has been proposed, which introduces equivariance through capsules to significantly enhance the learning ability for image segmentation and object detection. Due to its requirement of performing a high volume of matrix operations, CapsNets have been generally accelerated on modern GPU platforms that provide highly optimized software library for common deep learning tasks. However, based on our performance characterization on modern GPUs, CapsNets exhibit low efficiency due to the special program and execution features of their routing procedure, including massive unshareable intermediate variables and intensive synchronizations, which are very difficult to optimize at software level. To address these challenges, we propose a hybrid computing architecture design named PIM-CapsNet. It preserves GPU's on-chip computing capability for accelerating CNN types of layers in CapsNet, while pipelining with an off-chip in-memory acceleration solution that effectively tackles routing procedure's inefficiency by leveraging the processing-in-memory capability of today's 3D stacked memory. Using routing procedure's inherent parallellization feature, our design enables hierarchical improvements on CapsNet inference efficiency through minimizing data movement and maximizing parallel processing in memory. Evaluation results demonstrate that our proposed design can achieve substantial improvement on both performance and energy savings for CapsNet inference, with almost zero accuracy loss. The results also suggest good performance scalability in optimizing the routing procedure with increasing network size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7ae9b3c-4816-40b4-97f7-9d5c0a3ffb4eCited by top-tier papers3
- SpAtten: Efficient Sparse Attention Architecture with Cascade Token and Head PruningHanrui Wang, Zhekai Zhang, Song HanHPCA 2021 · 412 citations
- Accelerating applications using edge tensor processing unitsKuan-Chieh Hsu, Hung-Wei TsengSC 2021 · 33 citations
- η-LSTM: Co-Designing Highly-Efficient Large LSTM Training via Exploiting Memory-Saving and Architectural Design OpportunitiesXingyao Zhang, Haojun Xia, Donglin Zhuang, Hao Sun et al.ISCA 2021 · 7 citations
Related papers
- PT-CapsNet: A Novel Prediction-Tuning Capsule Network Suitable for Deeper ArchitecturesChenbin Pan, Senem VelipasalarICCV 2021 · 11 citations
- Lift: Exploiting Hybrid Stacked Memory for Energy-Efficient Processing of Graph Convolutional NetworksJiaxian Chen, Zhaoyu Zhong, Kaoyi Sun, Chenlin Ma et al.DAC 2023 · 10 citations
- Capsule: An Out-of-Core Training Mechanism for Colossal GNNsYongan Xiang, Zezhong Ding, Rui Guo, Shangyou Wang et al.SIGMOD 2025 · 6 citations
- Improving the Robustness of Capsule Networks to Image Affine TransformationsJindong Gu, Volker TrespCVPR 2020
- A Receptor Skeleton for Capsule Neural NetworksJintai Chen, Hongyun Yu, Chengde Qian, Danny Z. Chen et al.ICML 2021 · 5 citations
