Ultra-High Resolution Segmentation via Boundary-Enhanced Patch-Merging Transformer
Haopeng Sun, Yingwei Zhang, Lumin Xu, Sheng Jin, Yiqiang Chen
Abstract
Segmentation of ultra-high resolution (UHR) images is a critical task with numerous applications, yet it poses significant challenges due to high spatial resolution and rich fine details. Recent approaches adopt a dual-branch architecture, where a global branch learns long-range contextual information and a local branch captures fine details. However, they struggle to handle the conflict between global and local information while adding significant extra computational cost. Inspired by the human visual system's ability to rapidly orient attention to important areas with fine details and filter out irrelevant information, we propose a novel UHR segmentation method called Boundary-enhanced Patch-merging Transformer (BPT). BPT consists of two key components: (1) Patch-Merging Transformer (PMT) for dynamically allocating tokens to informative regions to acquire global and local representations, and (2) Boundary-Enhanced Module (BEM) that leverages boundary information to enrich fine details. Extensive experiments on multiple UHR image segmentation benchmarks demonstrate that our BPT outperforms previous state-of-the-art methods without introducing extra computational overhead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 08bd9c7e-c029-4674-bf10-9b9e21839952Cited by top-tier papers9
- ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language TasksRuixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li et al.CVPR 2026 · 19 citations
- Self-Prompting Analogical Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Song Tang et al.AAAI 2025 · 10 citations
- Conditional Latent Coding with Learnable Synthesized Reference for Deep Image CompressionSiqi Wu, Yinda Chen, Dong Liu, Zhihai HeAAAI 2025 · 9 citations
- Exploring Salient Object Detection with Adder Neural NetworksBo-Wen Yin, Zheng LinAAAI 2025 · 4 citations
- Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesLing Li, Ruiwen Gu, Chongyang Wang, Junliang Xing et al.AAAI 2025 · 3 citations
Builds on24
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian et al.CVPR 2022 · 132 citations
- Pedestrian-specific Bipartite-aware Similarity Learning for Text-based Person RetrievalFei Shen, Xiangbo Shu, Xiaoyu Du, Jinhui TangACM MM 2023 · 101 citations
- ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-high Resolution SegmentationShaohua Guo, Liang Liu, Zhenye Gan, Yabiao Wang et al.CVPR 2022 · 66 citations
Related papers
- From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual CorrelationQi Li, Weixiang Yang, Wenxi Liu, Yuanlong Yu et al.ICCV 2021 · 55 citations
- EDTER: Edge Detection with TransformerMengyang Pu, Yaping Huang, Yuming Liu, Qingji Guan et al.CVPR 2022 · 224 citations
- F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing SegmentationHengzhi Chen, Liqian Feng, Wenhua Wu, Xiaogang Zhu et al.CVPR 2026 · 9 citations
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao et al.ICCV 2021 · 473 citations
- Object Part Parsing with Hierarchical Dual TransformerJiamin Chen, Jianlou Si, Naihao Liu, Yao Wu et al.ACM MM 2023 · 1 citation
