Gramformer: Learning Crowd Counting via Graph-Modulated Transformer
Hui Lin, Zhiheng Ma, Xiaopeng Hong, Qinnan Shangguan, Deyu Meng
摘要
Transformer has been popular in recent crowd counting work since it breaks the limited receptive field of traditional CNNs. However, since crowd images always contain a large number of similar patches, the self-attention mechanism in Transformer tends to find a homogenized solution where the attention maps of almost all patches are identical. In this paper, we address this problem by proposing Gramformer: a graph-modulated transformer to enhance the network by adjusting the attention and input node features respectively on the basis of two different types of graphs. Firstly, an attention graph is proposed to diverse attention maps to attend to complementary information. The graph is building upon the dissimilarities between patches, modulating the attention in an anti-similarity fashion. Secondly, a feature-based centrality encoding is proposed to discover the centrality positions or importance of nodes. We encode them with a proposed centrality indices scheme to modulate the node features and similarity relationships. Extensive experiments on four challenging crowd counting datasets have validated the competitiveness of the proposed method. Code is available at https://github.com/LoraLinH/Gramformer .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Generative Adversarial Perturbations with Cross-paradigm Transferability on Localized Crowd CountingAlabi Mehzabin Anisha, Guangjing Wang, Sriram ChellappanCVPR 2026 · 被引用 1 次
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
- 2D Gaussians Spatial Transport for Point-supervised Density RegressionMiao Shang, Xiaopeng HongAAAI 2026
- Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image FusionYanglin Deng, Tianyang Xu, Chunyang Cheng, Hui Li 等CVPR 2026
- Dual Mamba for Node-Specific Representation Learning: Tackling Over-Smoothing with Selective State Space ModelingXin He, Yili Wang, Yiwei Dai, Xin WangAAAI 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Recipe for a General, Powerful, Scalable Graph TransformerLadislav Rampásek, Michael Galkin, Vijay Prakash Dwivedi, Anh Tuan Luu 等NeurIPS 2022 · 被引用 1,216 次
- Vision GNN: An Image is Worth Graph of NodesKai Han, Yunhe Wang, Jianyuan Guo, Yehui Tang 等NeurIPS 2022 · 被引用 668 次
- Graph Neural Networks with Learnable Structural and Positional RepresentationsVijay Prakash Dwivedi, Anh Tuan Luu, Thomas Laurent, Yoshua Bengio 等ICLR 2022 · 被引用 464 次
- Representing Long-Range Context for Graph Neural Networks with Global AttentionZhanghao Wu, Paras Jain, Matthew A. Wright, Azalia Mirhoseini 等NeurIPS 2021 · 被引用 450 次
相关 Paper
- Clusformer: A Transformer Based Clustering Approach to Unsupervised Large-Scale Face and Visual Landmark RecognitionXuan-Bac Nguyen, Duc Toan Bui, Chi Nhan Duong, Tien D. Bui 等CVPR 2021
- Learning Action-guided Spatio-temporal Transformer for Group Activity RecognitionWei Li, Tianzhao Yang, Xiao Wu, Xian-Jun Du 等ACM MM 2022 · 被引用 21 次
- Boosting Crowd Counting via Multifaceted AttentionHui Lin, Zhiheng Ma, Rongrong Ji, Yaowei Wang 等CVPR 2022 · 被引用 229 次
- Leveraging Contrastive Learning for Enhanced Node Representations in Tokenized Graph TransformersJinsong Chen, Hanpeng Liu, John E. Hopcroft, Kun HeNeurIPS 2024 · 被引用 23 次
- T-former: An Efficient Transformer for Image InpaintingYe Deng, Siqi Hui, Sanping Zhou, Deyu Meng 等ACM MM 2022 · 被引用 58 次
