Boosting Crowd Counting via Multifaceted Attention
Hui Lin, Zhiheng Ma, Rongrong Ji, Yaowei Wang, Xiaopeng Hong
摘要
This paper focuses on the challenging crowd counting task. As large-scale variations often exist within crowd images, neither fixed-size convolution kernel of CNN nor fixed-size attention of recent vision transformers can well handle this kind of variations. To address this problem, we propose a Multifaceted Attention Network (MAN) to improve transformer models in local spatial relation encoding. MAN incorporates global attention from vanilla transformer, learnable local attention, and instance attention into a counting model. Firstly, the local Learnable Region Attention (LRA) is proposed to assign attention exclusive for each feature location dynamically. Secondly, we design the Local Attention Regularization to supervise the training of LRA by minimizing the deviation among the attention for different feature locations. Finally, we provide an Instance Attention mechanism to focus on the most important instances dynamically during training. Extensive experiments on four challenging crowd counting datasets namely ShanghaiTech, UCF-QNRF, JHU++, and NWPU have validated the proposed method. Code: https://github.com/LoraLinH/Boosting- Crowd-Counting-via-Multifaceted-Attention.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Point-Query Quadtree for Crowd Counting, Localization, and MoreChengxin Liu, Hao Lu, Zhiguo Cao, Tongliang LiuICCV 2023 · 被引用 89 次
- STEERER: Resolving Scale Variations for Counting and Localization via Selective Inheritance LearningTao Han, Lei Bai, Lingbo Liu, Wanli OuyangICCV 2023 · 被引用 74 次
- Domain-General Crowd Counting in Unseen ScenariosZhipeng Du, Jiankang Deng, Miaojing ShiAAAI 2023 · 被引用 63 次
- Semi-supervised Crowd Counting via Density AgencyHui Lin, Zhiheng Ma, Xiaopeng Hong, Yaowei Wang 等ACM MM 2022 · 被引用 37 次
- Single Domain Generalization for Crowd CountingZhuoxuan Peng, S.-H. Gary ChanCVPR 2024 · 被引用 27 次
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- ConViT: Improving Vision Transformers with Soft Convolutional Inductive BiasesStéphane d'Ascoli, Hugo Touvron, Matthew L. Leavitt, Ari S. Morcos 等ICML 2021 · 被引用 1,021 次
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 被引用 612 次
相关 Paper
- Relational Attention Network for Crowd CountingAnran Zhang, Jiayi Shen, Zehao Xiao, Fan Zhu 等ICCV 2019 · 被引用 175 次
- Attentional Neural Fields for Crowd CountingAnran Zhang, Lei Yue, Jiayi Shen, Fan Zhu 等ICCV 2019 · 被引用 121 次
- Attention Scaling for Crowd CountingXiaoheng Jiang, Li Zhang, Mingliang Xu, Tianzhu Zhang 等CVPR 2020
- Vehicle Counting Network with Attention-based Mask Refinement and Spatial-awareness Block LossJi Zhang, Jian-Jun Qiao, Xiao Wu, Wei LiACM MM 2021 · 被引用 3 次
- Shallow Feature Based Dense Attention Network for Crowd CountingYunqi Miao, Zijia Lin, Guiguang Ding, Jungong HanAAAI 2020 · 被引用 120 次
