GigaHumanDet: Exploring Full-Body Detection on Gigapixel-Level Images
Chenglong Liu, Haoran Wei, Jinze Yang, Jintao Liu, Wenxi Li, Yuchen Guo, Lu Fang
摘要
Performing person detection in super-high-resolution images has been a challenging task. For such a task, modern detectors, which usually encode a box using center and width/height, struggle with accuracy due to two factors: 1) Human characteristic: people come in various postures and the center with high freedom is difficult to capture robust visual pattern; 2) Image characteristic: due to vast scale diversity of input (gigapixel-level), distance regression (for width and height) is hard to pinpoint, especially for a person, with substantial scale, who is near the camera. To address these challenges, we propose GigaHumanDet, an innovative solution aimed at further enhancing detection accuracy for gigapixel-level images. GigaHumanDet employs the corner modeling method to avoid the potential issues of a high degree of freedom in center pinpointing. To better distinguish similar-looking persons and enforce instance consistency of corner pairs, an instance-guided learning approach is designed to capture discriminative individual semantics. Further, we devise reliable shape-aware bodyness equipped with a multi-precision strategy as the human corner matching guidance to be appropriately adapted to the single-view large scene. Experimental results on PANDA and STCrowd datasets show the superiority and strong applicability of our design. Notably, our model achieves 82.4% in term of AP, outperforming current state-of-the-arts by more than 10%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionXiang Li, Wenxi Li, Yuetong Wang, Chenyang Lyu 等AAAI 2026 · 被引用 1 次
- 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information PropagationLonglong Yu, Wenxi Li, Yaoqi Sun, Hang Xu 等AAAI 2026
它引用的顶会 Paper11
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi 等ICCV 2019 · 被引用 3,348 次
- DINO: DETR with Improved DeNoising Anchor Boxes for End-to-End Object DetectionHao Zhang, Feng Li, Shilong Liu, Lei Zhang 等ICLR 2023 · 被引用 753 次
- Clustered Object Detection in Aerial ImagesFan Yang, Heng Fan, Peng Chu, Erik Blasch 等ICCV 2019 · 被引用 384 次
- Mask-Guided Attention Network for Occluded Pedestrian DetectionYanwei Pang, Jin Xie, Muhammad Haris Khan, Rao Muhammad Anwer 等ICCV 2019 · 被引用 216 次
- STCrowd: A Multimodal Dataset for Pedestrian Perception in Crowded ScenesPeishan Cong, Xinge Zhu, Feng Qiao, Yiming Ren 等CVPR 2022 · 被引用 43 次
相关 Paper
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 等CVPR 2020
- HumanLiker: A Human-like Object Detector to Model the Manual Labeling ProcessHaoran Wei, Ping Guo, Yangguang Zhu, Chenglong Liu 等NeurIPS 2022 · 被引用 4 次
- HigherHRNet: Scale-Aware Representation Learning for Bottom-Up Human Pose EstimationBowen Cheng, Bin Xiao, Jingdong Wang, Honghui Shi 等CVPR 2020
- When Visual Grounding Meets Gigapixel-Level Large-Scale Scenes: Benchmark and ApproachM. Tao, Bing Bai, Haozhe Lin, Heyuan Wang 等CVPR 2024 · 被引用 4 次
- Speed up Object Detection on Gigapixel-level Images with Patch ArrangementJiahao Fan, Huabin Liu, Wenjie Yang, John See 等CVPR 2022 · 被引用 13 次
