ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer
Wenxi Li, Jingchen Huang, Chenyang Lyu, Moran Liu, Haozhe Lin, Guiguang Ding, Yuchen Guo
Abstract
Recent advances in gigapixel-level imaging have brought High-Resolution Wide shots to the forefront of research. However, these images present significant challenges: extreme sparsity of foreground, gigapixel-level resolutions and diverse target counts. This makes traditional closeup detectors inaccurate and slow as they are overwhelmed by the background. Although previous research has explored sparse backbones, their fixed sparsity patterns lack the adaptability required to handle diverse target numbers. To address this, we introduce ElasticFormer, a sparse backbone that dynamically allocates computational resources based on foreground proportion. After scoring windows based on variance, proposed ElasticSelector module will predict the foreground proportion for top-k selection. The mechanism guides the model to select target-containing windows, scaling resources in areas where objects are clustered. We introduce a novel loss function combined with the 3-phase training strategy for ElasticSelector, allowing it to function properly when bounding box annotations are missing. A WSOD study is carried on PASCAL VOC 2007 to evaluate its extensibility. Further, ElasticNet is created to verify its backbone-agnostic nature. In experiments on the PANDA gigapixel benchmark, ElasticFormer reduces backbone FLOPs by 80% while achieving a significant improvement in AP 50 when compared to fixed-ratio sparse methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without ConvolutionsWenhai Wang, Enze Xie, Xiang Li, Deng-Ping Fan et al.ICCV 2021 · 4,909 citations
- DynamicViT: Efficient Vision Transformers with Dynamic Token SparsificationYongming Rao, Wenliang Zhao, Benlin Liu, Jiwen Lu et al.NeurIPS 2021 · 1,343 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
Related papers
- SparseFormer: Detecting Objects in HRW Shots via Sparse Vision TransformerWenxi Li, Yuchen Guo, Jilai Zheng, Haozhe Lin et al.ACM MM 2024 · 3 citations
- GigaMoE: Sparsity-Guided Mixture of Experts for Efficient Gigapixel Object DetectionXiang Li, Wenxi Li, Yuetong Wang, Chenyang Lyu et al.AAAI 2026 · 1 citation
- FSHNet: Fully Sparse Hybrid Network for 3D Object DetectionShuai Liu, Mingyue Cui, Boyang Li, Quanmin Liang et al.CVPR 2025
- Sparse DETR: Efficient End-to-End Object Detection with Learnable SparsityByungseok Roh, Jaewoong Shin, Wuhyun Shin, Saehoon KimICLR 2022 · 256 citations
- Elastic Attention: Test-time Adaptive Sparsity Ratios for Efficient TransformersZecheng Tang, Quantong Qiu, Yi Yang, Zhiyi Hong et al.ICML 2026 · 4 citations
