SC2024Top-tier venue
Adaptive Patching for High-resolution Image Segmentation with Transformers
Enzhi Zhang, Isaac Lyngaas, Peng Chen, Xiao Wang, Jun Igarashi, Yuankai Huo, Masaharu Munetomo, Mohamed Wahib
Abstract
Attention-based models are proliferating in the space of image analytics, including segmentation. The standard method of feeding images to transformer encoders is to divide the images into patches and then feed the patches to the model as a linear sequence of tokens. For high-resolution images, e.g. microscopic pathology images, the quadratic compute and memory cost prohibits the use of an attention-based model, if we are to use smaller patch sizes that are favorable in segmentation. The solution is to either use custom complex multi-resolution models or approximate attention schemes. We take inspiration from Adapative Mesh Refinement (AMR) methods in HPC by adaptively patching the images, as a pre-processing step, based on the image details to reduce the number of patches being fed to the model, by orders of magnitude. This method has a negligible overhead, and works seamlessly with any attention-based model, i.e. it is a pre-processing step that can be adopted by any attention-based model without friction. We demonstrate superior segmentation quality over SoTA segmentation models for real-world pathology datasets while gaining a geomean speedup of for resolutions up to , on up to 2,048 GPUs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ecb2019-da63-46e0-833f-c5c4e5625c5dCited by top-tier papers1
Ask how each one uses itBuilds on23
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 2,665 citations
Related papers
- Faster Vision Transformers with Adaptive PatchesRohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang et al.ICLR 2026 · 8 citations
- Iterative Patch Selection for High-Resolution Image RecognitionBenjamin Bergner, Christoph Lippert, Aravindh MahendranICLR 2023 · 4 citations
- MeshTok: Efficient Multi-Scale Tokenization for Scalable PDE TransformersZhao Yanshun, Xiaoyu Peng, Jiamin Jiang, Congcong Zhu et al.ICML 2026
- Act Like a Pathologist: Tissue-Aware Whole Slide Image ReasoningWentao Huang, Weimin Lyu, Peiliang Lou, Qingqiao Hu et al.CVPR 2026 · 3 citations
- One Patch Doesn't Fit All: Adaptive Patching for Native-Resolution Multimodal Large Language ModelsWenzhuo Liu, Weijie Yin, Fei Zhu, Shijie Ma et al.ICLR 2026
