Context Matters: Query-aware Dynamic Long Sequence Modeling of Gigapixel Images
Zhengrui Guo, Qichen Sun, Jiabo Ma, Lishuang Feng, Jinzhuo Wang, Hao Chen
Abstract
Whole slide image (WSI) analysis presents significant computational challenges due to the massive number of patches in gigapixel images. While transformer architectures excel at modeling longrange correlations through self-attention, their quadratic computational complexity makes them impractical for computational pathology applications. Existing solutions like local-global or linear self-attention reduce computational costs but compromise the strong modeling capabilities of full self-attention. In this work, we propose Querent, i.e., the query-aware long contextual dynamic modeling framework, which achieves a theoretically bounded approximation of full selfattention while delivering practical efficiency. Our method adaptively predicts which surrounding regions are most relevant for each patch, enabling focused yet unrestricted attention computation only with potentially important contexts. By using efficient region-wise metadata computation and importance estimation, our approach dramatically reduces computational overhead while preserving global perception to model fine-grained patch correlations. Through comprehensive experiments on biomarker prediction, gene mutation prediction, cancer subtyping, and survival analysis across over 10 WSI datasets, our method demonstrates superior performance compared to the state-of-the-art approaches. Codes are here.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance LearningDaniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson et al.ICLR 2026 · 3 citations
- Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image ClassificationConghao Xiong, Zhengrui Guo, Zhe Xu, Yifei Zhang et al.ICLR 2026 · 1 citation
- FOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image ClassificationZhengrui Guo, Conghao Xiong, Jiabo Ma, Qichen Sun et al.CVPR 2025
Builds on20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang et al.NeurIPS 2021 · 1,163 citations
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan et al.AAAI 2021 · 675 citations
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen et al.CVPR 2022 · 490 citations
Related papers
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen et al.ICCV 2021 · 369 citations
- Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisHonglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui et al.NeurIPS 2024 · 27 citations
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous ReasoningJiusong Ge, Yingkang Zhan, Wenjie Zhao, Di Zhang et al.ICML 2026
- Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationYingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning WangAAAI 2024 · 10 citations
- TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation LearningShahira Abousamra, Asmita Sood, Sylvia PlevritisCVPR 2026
