Ultra-High Resolution Segmentation via Boundary-Enhanced Patch-Merging Transformer
Haopeng Sun, Yingwei Zhang, Lumin Xu, Sheng Jin, Yiqiang Chen
摘要
Segmentation of ultra-high resolution (UHR) images is a critical task with numerous applications, yet it poses significant challenges due to high spatial resolution and rich fine details. Recent approaches adopt a dual-branch architecture, where a global branch learns long-range contextual information and a local branch captures fine details. However, they struggle to handle the conflict between global and local information while adding significant extra computational cost. Inspired by the human visual system's ability to rapidly orient attention to important areas with fine details and filter out irrelevant information, we propose a novel UHR segmentation method called Boundary-enhanced Patch-merging Transformer (BPT). BPT consists of two key components: (1) Patch-Merging Transformer (PMT) for dynamically allocating tokens to informative regions to acquire global and local representations, and (2) Boundary-Enhanced Module (BEM) that leverages boundary information to enrich fine details. Extensive experiments on multiple UHR image segmentation benchmarks demonstrate that our BPT outperforms previous state-of-the-art methods without introducing extra computational overhead.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- ZoomEarth: Active Perception for Ultra-High-Resolution Geospatial Vision-Language TasksRuixun Liu, Bowen Fu, Jiayi Song, Kaiyu Li 等CVPR 2026 · 被引用 19 次
- Self-Prompting Analogical Reasoning for UAV Object DetectionNianxin Li, Mao Ye, Lihua Zhou, Song Tang 等AAAI 2025 · 被引用 10 次
- Conditional Latent Coding with Learnable Synthesized Reference for Deep Image CompressionSiqi Wu, Yinda Chen, Dong Liu, Zhihai HeAAAI 2025 · 被引用 9 次
- Exploring Salient Object Detection with Adder Neural NetworksBo-Wen Yin, Zheng LinAAAI 2025 · 被引用 4 次
- Multi-View 3D Human Pose Estimation with Weakly Synchronized ImagesLing Li, Ruiwen Gu, Chongyang Wang, Junliang Xing 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper24
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar 等NeurIPS 2021 · 被引用 9,661 次
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 被引用 2,196 次
- Not All Tokens Are Equal: Human-centric Visual Analysis via Token Clustering TransformerWang Zeng, Sheng Jin, Wentao Liu, Chen Qian 等CVPR 2022 · 被引用 132 次
- Pedestrian-specific Bipartite-aware Similarity Learning for Text-based Person RetrievalFei Shen, Xiangbo Shu, Xiaoyu Du, Jinhui TangACM MM 2023 · 被引用 101 次
- ISDNet: Integrating Shallow and Deep Networks for Efficient Ultra-high Resolution SegmentationShaohua Guo, Liang Liu, Zhenye Gan, Yabiao Wang 等CVPR 2022 · 被引用 66 次
相关 Paper
- From Contexts to Locality: Ultra-high Resolution Image Segmentation via Locality-aware Contextual CorrelationQi Li, Weixiang Yang, Wenxi Liu, Yuanlong Yu 等ICCV 2021 · 被引用 55 次
- EDTER: Edge Detection with TransformerMengyang Pu, Yaping Huang, Yuming Liu, Qingji Guan 等CVPR 2022 · 被引用 224 次
- F2Net: A Frequency-Fused Network for Ultra-High Resolution Remote Sensing SegmentationHengzhi Chen, Liqian Feng, Wenhua Wu, Xiaogang Zhu 等CVPR 2026 · 被引用 9 次
- Visual Saliency TransformerNian Liu, Ni Zhang, Kaiyuan Wan, Ling Shao 等ICCV 2021 · 被引用 473 次
- Object Part Parsing with Hierarchical Dual TransformerJiamin Chen, Jianlou Si, Naihao Liu, Yao Wu 等ACM MM 2023 · 被引用 1 次
