Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd Counting
Lingbo Liu, Jiaqi Chen, Hefeng Wu, Guanbin Li, Chenglong Li, Liang Lin
摘要
Crowd counting is a fundamental yet challenging task, which desires rich information to generate pixel-wise crowd density maps. However, most previous methods only used the limited information of RGB images and cannot well discover potential pedestrians in unconstrained scenarios. In this work, we find that incorporating optical and thermal information can greatly help to recognize pedestrians. To promote future researches in this field, we introduce a large-scale RGBT Crowd Counting (RGBT-CC) benchmark, which contains 2,030 pairs of RGB-thermal images with 138,389 annotated people. Furthermore, to facilitate the multimodal crowd counting, we propose a crossmodal collaborative representation learning framework, which consists of multiple modality-specific branches, a modality-shared branch, and an Information Aggregation-Distribution Module (IADM) to capture the complementary information of different modalities fully. Specifically, our IADM incorporates two collaborative information transfers to dynamically enhance the modality-shared and modalityspecific representations with a dual information propagation mechanism. Extensive experiments conducted on the RGBT-CC benchmark demonstrate the effectiveness of our framework for RGBT crowd counting. Moreover, the proposed approach is universal for multimodal crowd counting and is also capable to achieve superior performance on the ShanghaiTechRGBD [22] dataset. Finally, our source code and benchmark are released at http://lingboliu. com/RGBT_Crowd_Counting.html.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Rethinking Spatial Invariance of Convolutional Networks for Object CountingZhi-Qi Cheng, Qi Dai, Hong Li, Jingkuan Song 等CVPR 2022 · 被引用 119 次
- CLIP-Count: Towards Text-Guided Zero-Shot Object CountingRuixiang Jiang, Lingbo Liu, Changwen ChenACM MM 2023 · 被引用 78 次
- Coarse to Fine: Domain Adaptive Crowd Counting via Adversarial Scoring NetworkZhikang Zou, Xiaoye Qu, Pan Zhou, Shuangjie Xu 等ACM MM 2021 · 被引用 37 次
- UniRGB-IR: A Unified Framework for Visible-Infrared Semantic Tasks via Adapter TuningMaoxun Yuan, Bo Cui, Tianyi Zhao, Jiayi Wang 等ACM MM 2025 · 被引用 21 次
- DR.VIC: Decomposition and Reasoning for Video Individual CountingTao Han, Lei Bai, Junyu Gao, Qi Wang 等CVPR 2022 · 被引用 18 次
它引用的顶会 Paper14
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 被引用 612 次
- Depth-Induced Multi-Scale Recurrent Attention Network for Saliency DetectionYongri Piao, Wei Ji, Jingjing Li, Miao Zhang 等ICCV 2019 · 被引用 450 次
- Crowd Counting With Deep Structured Scale Integration NetworkLingbo Liu, Zhilin Qiu, Guanbin Li, Shufan Liu 等ICCV 2019 · 被引用 254 次
- Multi-Level Bottom-Top and Top-Bottom Feature Fusion for Crowd CountingVishwanath Sindagi, Vishal M. PatelICCV 2019 · 被引用 194 次
- Relational Attention Network for Crowd CountingAnran Zhang, Jiayi Shen, Zehao Xiao, Fan Zhu 等ICCV 2019 · 被引用 175 次
相关 Paper
- Robust Multi-Modality Person Re-identificationAihua Zheng, Zi Wang, Zi-Han Chen, Chenglong Li 等AAAI 2021 · 被引用 79 次
- SemanticRT: A Large-Scale Dataset and Method for Robust Semantic Segmentation in Multispectral ImagesWei Ji, Jingjing Li, Cheng Bian, Zhicheng Zhang 等ACM MM 2023 · 被引用 22 次
- Simplifying Cross-modal Interaction via Modality-Shared Features for RGBT TrackingLiqiu Chen, Yuqing Huang, Hengyu Li, Zikun Zhou 等ACM MM 2024 · 被引用 2 次
- Free Lunch Enhancements for Multi-modal Crowd CountingHaoliang Meng, Xiaopeng Hong, Zhengqin Lai, Miao ShangCVPR 2025
- DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action RecognitionYuanjun Tan, Aoran Xiao, Liqian Deng, Zhigang TuCVPR 2026 · 被引用 1 次
