Free Lunch Enhancements for Multi-modal Crowd Counting
Haoliang Meng, Xiaopeng Hong, Zhengqin Lai, Miao Shang
Abstract
This paper addresses multi-modal crowd counting with a novel 'free lunch' training enhancement strategy that requires no additional data, parameters, or increased inference complexity. First, we introduce a cross-modal alignment technique as a plug-in post-processing step for the pre-trained backbone network, enhancing the model's ability to capture shared information across modalities. Second, we incorporate a regional density supervision mechanism during the fine-tuning stage, which differentiates features in regions with varying crowd densities. Extensive experiments on three multi-modal crowd counting datasets validate our approach, making it the first to achieve an MAE below 10 on RGBT-CC. The code is available at https://github.com/HenryCilence/Free-Lunch-Multimodal-Counting .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ef0cbc10-effe-4924-9ec8-beb736fa54e0Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
Related papers
- Semi-supervised Crowd Counting via Density AgencyHui Lin, Zhiheng Ma, Xiaopeng Hong, Yaowei Wang et al.ACM MM 2022 · 37 citations
- CrossNet: Boosting Crowd Counting with LocalizationJi Zhang, Zhi-Qi Cheng, Xiao Wu, Wei Li et al.ACM MM 2022 · 20 citations
- Cross-Modal Collaborative Representation Learning and a Large-Scale RGBT Benchmark for Crowd CountingLingbo Liu, Jiaqi Chen, Hefeng Wu, Guanbin Li et al.CVPR 2021
- Cross-View Cross-Scene Multi-View Crowd CountingQi Zhang, Wei Lin, Antoni B. ChanCVPR 2021
- 3D Crowd Counting via Multi-View Fusion with 3D Gaussian KernelsQi Zhang, Antoni B. ChanAAAI 2020 · 41 citations
