Boosting Crowd Counting via Multifaceted Attention
Hui Lin, Zhiheng Ma, Rongrong Ji, Yaowei Wang, Xiaopeng Hong
Abstract
This paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size attention of recent vision transformers can well handle this kind of variations. To address this problem, we propose a Multifaceted Attention Network (MAN) to improve transformer models in local spatial relation encoding. MAN incorporates global attention from vanilla transformer, learnable local attention, and instance attention into a counting model. Firstly, the local Learnable Region Attention (LRA) is proposed to assign attention exclusive for each feature location dynamically. Secondly, we design the Local Attention Regularization to supervise the training of LRA by minimizing the deviation among the attention for different feature locations. Finally, we provide an Instance Attention mechanism to focus on the most important instances dynamically during training. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code: https://github.com/LoraLinH/Boosting- Crowd-Counting-via-Multifaceted-Attention.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3ae14558-a283-45af-ae0c-1dd59d8af403Cited by top-tier papers23
- Point-Query Quadtree for Crowd Counting, Localization, and MoreChengxin Liu, Hao Lu, Zhiguo Cao, Tongliang LiuICCV 2023 · 89 citations
- STEERER: Resolving Scale Variations for Counting and Localization via Selective Inheritance LearningTao Han, Lei Bai, Lingbo Liu, Wanli OuyangICCV 2023 · 74 citations
- Domain-General Crowd Counting in Unseen ScenariosZhipeng Du, Jiankang Deng, Miaojing ShiAAAI 2023 · 63 citations
- Semi-supervised Crowd Counting via Density AgencyHui Lin, Zhiheng Ma, Xiaopeng Hong, Yaowei Wang et al.ACM MM 2022 · 37 citations
- Single Domain Generalization for Crowd CountingZhuoxuan Peng, S.-H. Gary ChanCVPR 2024 · 27 citations
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos et al.ICML 2021 · 1,021 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
Related papers
- Relational Attention Network for Crowd CountingAnran Zhang, Jiayi Shen, Zehao Xiao, Fan Zhu et al.ICCV 2019 · 175 citations
- Attentional Neural Fields for Crowd CountingAnran Zhang, Lei Yue, Jiayi Shen, Fan Zhu et al.ICCV 2019 · 121 citations
- Attention Scaling for Crowd CountingXiaoheng Jiang, Li Zhang, Mingliang Xu, Tianzhu Zhang et al.CVPR 2020
- Vehicle Counting Network with Attention-based Mask Refinement and Spatial-awareness Block LossJi Zhang, Jian-Jun Qiao, Xiao Wu, Wei LiACM MM 2021 · 3 citations
- Shallow Feature Based Dense Attention Network for Crowd CountingYunqi Miao, Zijia Lin, Guiguang Ding, Jungong HanAAAI 2020 · 120 citations
