Toward Real Ultra Image Segmentation: Leveraging Surrounding Context to Cultivate General Segmentation Model
Sai Wang, Yutian Lin, Yu Wu, Bo Du
Abstract
Existing ultra image segmentation methods suffer from two major challenges, namely the scalability issue (i.e. they lack the stability and generality of standard segmentation models, as they are tailored to specific datasets), and the architectural issue (i.e. they are incompatible with real-world ultra image scenes, as they compromise between image size and computing resources). To tackle these issues, we revisit the classic sliding inference framework, upon which we propose a Sur-rounding Guided Segmentation framework (SGNet) for ultra image segmentation. The SGNet leverages a larger area around each image patch to refine the general segmentation results of local patches. Specifically, we propose a surrounding context integration module to absorb surrounding context information and extract specific features that are beneficial to local patches. Note that, SGNet can be seamlessly integrated to any general segmentation model. Extensive experiments on five datasets demonstrate that SGNet achieves competitive performance and consistent improvements across a variety of general segmentation models, surpassing the traditional ultra image segmentation methods by a large margin.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing SegmentationHengzhi Chen, Liqian Feng, Wenhua Wu, Xiaogang Zhu et al.CVPR 2026 · 9 citations
- SSR-SAM: Retrieval-Style Segment Anything Model for Semi-Supervised Ultra-High-Resolution Image SegmentationShijie Li, Yiming Chen, Zhineng Chen, Kai Hu et al.AAAI 2026
Builds on21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- Dynamic Snake Convolution based on Topological Geometric Constraints for Tubular Structure SegmentationYaolei Qi, Yuting He, Xiaoming Qi, Yuan Zhang et al.ICCV 2023 · 467 citations
- RegionViT: Regional-to-Local Attention for Vision TransformersChun-Fu Chen, Rameswar Panda, Quanfu FanICLR 2022 · 246 citations
Related papers
- From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual CorrelationQi Li, Weixiang Yang, Wenxi Liu, Yuanlong Yu et al.ICCV 2021 · 55 citations
- ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-high Resolution SegmentationShaohua Guo, Liang Liu, Zhenye Gan, Yabiao Wang et al.CVPR 2022 · 66 citations
- Patch Proposal Network for Fast Semantic Segmentation of High-Resolution ImagesTong Wu, Zhenzhen Lei, Bingqian Lin, Cuihua Li et al.AAAI 2020 · 42 citations
- Show and Segment: Universal Medical Image Segmentation via In-Context LearningYunhe Gao, Di Liu, Zhuowei Li, Yunsheng Li et al.CVPR 2025
- SPGNet: Semantic Prediction Guidance for Scene ParsingBowen Cheng, Liang-Chieh Chen, Yunchao Wei, Yukun Zhu et al.ICCV 2019 · 117 citations
