Context Matters: Query-aware Dynamic Long Sequence Modeling of Gigapixel Images
Zhengrui Guo, Qichen Sun, Jiabo Ma, Lishuang Feng, Jinzhuo Wang, Hao Chen
摘要
Whole slide image (WSI) analysis presents significant computational challenges due to the massive number of patches in gigapixel images. While transformer architectures excel at modeling longrange correlations through self-attention, their quadratic computational complexity makes them impractical for computational pathology applications. Existing solutions like local-global or linear self-attention reduce computational costs but compromise the strong modeling capabilities of full self-attention. In this work, we propose Querent, i.e., the query-aware long contextual dynamic modeling framework, which achieves a theoretically bounded approximation of full selfattention while delivering practical efficiency. Our method adaptively predicts which surrounding regions are most relevant for each patch, enabling focused yet unrestricted attention computation only with potentially important contexts. By using efficient region-wise metadata computation and importance estimation, our approach dramatically reduces computational overhead while preserving global perception to model fine-grained patch correlations. Through comprehensive experiments on biomarker prediction, gene mutation prediction, cancer subtyping, and survival analysis across over 10 WSI datasets, our method demonstrates superior performance compared to the state-of-the-art approaches. Codes are here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Mixture of Mini Experts: Overcoming the Linear Layer Bottleneck in Multiple Instance LearningDaniel Shao, Joel Runevic, Richard J. Chen, Drew F. K. Williamson 等ICLR 2026 · 被引用 3 次
- Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image ClassificationConghao Xiong, Zhengrui Guo, Zhe Xu, Yifei Zhang 等ICLR 2026 · 被引用 1 次
- FOCUS: Knowledge-enhanced Adaptive Visual Compression for Few-shot Whole Slide Image ClassificationZhengrui Guo, Conghao Xiong, Jiabo Ma, Qichen Sun 等CVPR 2025
它引用的顶会 Paper20
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image ClassificationZhuchen Shao, Hao Bian, Yang Chen, Yifeng Wang 等NeurIPS 2021 · 被引用 1,163 次
- Nyströmformer: A Nyström-based Algorithm for Approximating Self-AttentionYunyang Xiong, Zhanpeng Zeng, Rudrasis Chakraborty, Mingxing Tan 等AAAI 2021 · 被引用 675 次
- Scaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningRichard J. Chen, Chengkuan Chen, Yicong Li, Tiffany Y. Chen 等CVPR 2022 · 被引用 490 次
相关 Paper
- Multimodal Co-Attention Transformer for Survival Prediction in Gigapixel Whole Slide ImagesRichard J. Chen, Ming Y. Lu, Wei-Hung Weng, Tiffany Y. Chen 等ICCV 2021 · 被引用 369 次
- Rethinking Transformer for Long Contextual Histopathology Whole Slide Image AnalysisHonglin Li, Yunlong Zhang, Pingyi Chen, Zhongyi Shui 等NeurIPS 2024 · 被引用 27 次
- Thinking in Scales: Accelerating Gigapixel Pathology Image Analysis via Adaptive Continuous ReasoningJiusong Ge, Yingkang Zhan, Wenjie Zhao, Di Zhang 等ICML 2026
- Transformer-Based Video-Structure Multi-Instance Learning for Whole Slide Image ClassificationYingfan Ma, Xiaoyuan Luo, Kexue Fu, Manning WangAAAI 2024 · 被引用 10 次
- TopoSlide: Topologically-Informed Histopathology Whole Slide Image Representation LearningShahira Abousamra, Asmita Sood, Sylvia PlevritisCVPR 2026
