Speed up Object Detection on Gigapixel-level Images with Patch Arrangement
Jiahao Fan, Huabin Liu, Wenjie Yang, John See, Aixin Zhang, Weiyao Lin
摘要
With the appearance of super high-resolution (e.g., gigapixel-level) images, performing efficient object detection on such images becomes an important issue. Most ex-isting works for efficient object detection on high-resolution images focus on generating local patches where objects may exist, and then every patch is detected independently. How-ever, when the image resolution reaches gigapixel-level, they will suffer from a huge time cost for detecting numerous patches. Different from them, we devise a novel patch ar-rangement frameworkfor fast object detection on gigapixel-level images. Under this framework, a Patch Arrangement Network (PAN) is proposed to accelerate the detection by determining which patches could be packed together into a compact canvas. Specifically, PAN consists of (1) a Patch Filter Module (PFM) (2) a Patch Packing Module (PPM). PFM filters patch candidates by learning to select patches between two granularities. Subsequently, from the remaining patches, PPM determines how to pack these patches to-gether into a smaller number of canvases. Meanwhile, it generates an ideal layout of patches on canvas. These can-vases are fed to the detector to get final results. Experiments show that our method could improve the inference speed on gigapixel-level images by 5 x while maintaining great performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- GigaHumanDet: Exploring Full-Body Detection on Gigapixel-Level ImagesChenglong Liu, Haoran Wei, Jinze Yang, Jintao Liu 等AAAI 2024 · 被引用 6 次
- SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerWenxi Li, Yuchen Guo, Jilai Zheng, Haozhe Lin 等ACM MM 2024 · 被引用 3 次
- GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionXiang Li, Wenxi Li, Yuetong Wang, Chenyang Lyu 等AAAI 2026 · 被引用 1 次
- ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision TransformerWenxi Li, Jingchen Huang, Chenyang Lyu, Moran Liu 等CVPR 2026
- 2D-CrossScan Mamba: Enhancing State Space Models with Spatially Consistent Multi-Path 2D Information PropagationLonglong Yu, Wenxi Li, Yaoqi Sun, Hang Xu 等AAAI 2026
它引用的顶会 Paper3
- Clustered Object Detection in Aerial ImagesFan Yang, Heng Fan, Peng Chu, Erik Blasch 等ICCV 2019 · 被引用 384 次
- AutoFocus: Efficient Multi-Scale InferenceMahyar Najibi, Bharat Singh, Larry DavisICCV 2019 · 被引用 143 次
- PANDA: A Gigapixel-Level Human-Centric Video DatasetXueyang Wang, Xiya Zhang, Yinheng Zhu, Yuchen Guo 等CVPR 2020
相关 Paper
- Patch Proposal Network for Fast Semantic Segmentation of High-Resolution ImagesTong Wu, Zhenzhen Lei, Bingqian Lin, Cuihua Li 等AAAI 2020 · 被引用 42 次
- Efficient and Accurate Arbitrary-Shaped Text Detection With Pixel Aggregation NetworkWenhai Wang, Enze Xie, Xiaoge Song, Yuhang Zang 等ICCV 2019 · 被引用 490 次
- UFPMP-Det: Toward Accurate and Efficient Object Detection on Drone ImageryYecheng Huang, Jiaxin Chen, Di HuangAAAI 2022 · 被引用 162 次
- Revisiting Spatial-Frequency Information Integration from a Hierarchical Perspective for Panchromatic and Multi-Spectral Image FusionJiangtong Tan, Jie Huang, Naishan Zheng, Man Zhou 等CVPR 2024 · 被引用 27 次
- Faster-PPN: Towards Real-Time Semantic Segmentation with Dual Mutual Learning for Ultra-High Resolution ImagesBicheng Dai, Kaisheng Wu, Tong Wu, Kai Li 等ACM MM 2021 · 被引用 3 次
