Boltzmann Attention Sampling for Image Analysis with Small Objects
Theodore Zhao, Sid Kiblawi, Naoto Usuyama, Ho Hin Lee, Sam Preston, Hoifung Poon, Mu Wei
摘要
Detecting and segmenting small objects, such as lung nodules and tumor lesions, remains a critical challenge in image analysis. These objects often occupy less than 0.1% of an image, making traditional transformer architectures inefficient and prone to performance degradation due to redundant attention computations on irrelevant regions. Existing sparse attention mechanisms rely on rigid hierarchical structures, which are poorly suited for detecting small, variable, and uncertain object locations. In this paper, we propose BoltzFormer, a novel transformer-based architecture designed to address these challenges through dynamic sparse attention. BoltzFormer identifies and focuses attention on relevant areas by modeling uncertainty using a Boltzmann distribution with an annealing schedule. Initially, a higher temperature allows broader area sampling in early layers, when object location uncertainty is greatest. As the temperature decreases in later layers, attention becomes more focused, enhancing efficiency and accuracy. BoltzFormer seamlessly integrates into existing transformer architectures via a modular Boltzmann attention sampling mechanism. Comprehensive evaluations on benchmark datasets demonstrate that BoltzFormer significantly improves segmentation performance for small objects while reducing attention computation by an order of magnitude compared to previous state-of-the-art methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- VoxTell: Free-Text Promptable Universal 3D Medical Image SegmentationMaximilian Rokuss, Moritz Langenberg, Yannick Kirchhoff, Fabian Isensee 等CVPR 2026 · 被引用 22 次
- Beyond Predictive Resampling: Learning Input-Agnostic Downsampling for Efficient Aligned Vision RecognitionKai Zhao, Liting Ruan, Haoran Jiang, Xiaoqiang Zhu 等AAAI 2026
- Masked-Diffusion Autoencoders for 3D Medical Vision Representation LearningJiachen Tu, Guanghui Qin, Theodore Zhengde Zhao, Jeya Maria Jose Valanarasu 等CVPR 2026
它引用的顶会 Paper13
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li 等NeurIPS 2023 · 被引用 889 次
- Focal Modulation NetworksJianwei Yang, Chunyuan Li, Xiyang Dai, Jianfeng GaoNeurIPS 2022 · 被引用 494 次
相关 Paper
- AutoFocusFormer: Image Segmentation off the GridZiwen Chen, Kaushik Patnaik, Shuangfei Zhai, Alvin Wan 等CVPR 2023
- BiFormer: Vision Transformer with Bi-Level Routing AttentionLei Zhu, Xinjiang Wang, Zhanghan Ke, Wayne Zhang 等CVPR 2023
- RKformer: Runge-Kutta Transformer with Random-Connection Attention for Infrared Small Target DetectionMingjin Zhang, Haichen Bai, Jing Zhang, Rui Zhang 等ACM MM 2022 · 被引用 227 次
- DTMFormer: Dynamic Token Merging for Boosting Transformer-Based Medical Image SegmentationZhehao Wang, Xian Lin, Nannan Wu, Li Yu 等AAAI 2024 · 被引用 14 次
- SSTVOS: Sparse Spatiotemporal Transformers for Video Object SegmentationBrendan Duke, Abdalla Ahmed, Christian Wolf, Parham Aarabi 等CVPR 2021
