MaSS13K: A Matting-level Semantic Segmentation Benchmark
Chenxi Xie, Minghan Li, Hui Zeng, Jun Luo, Lei Zhang
Abstract
High-resolution semantic segmentation is essential for applications such as image editing, bokeh imaging, AR/VR, etc. Unfortunately, existing datasets often have limited resolution and lack precise mask details and boundaries. In this work, we build a large-scale, matting-level semantic segmentation dataset, named MaSS13K, which consists of 13,348 real-world images, all at 4K resolution. MaSS13K provides high-quality mask annotations of a number of objects, which are categorized into seven categories: human, vegetation, ground, sky, water, building, and others. MaSS13K features precise masks, with an average mask complexity 20-50 times higher than existing semantic segmentation datasets. We consequently present a method specifically designed for high-resolution semantic segmentation, namely MaSSFormer, which employs an efficient pixel decoder that aggregates high-level semantic features and low-level texture features across three stages, aiming to produce high-resolution masks with minimal computational cost. Finally, we propose a new learning paradigm, which integrates the high-quality masks of the seven given categories with pseudo labels from new classes, enabling MaSSFormer to transfer its accurate segmentation capability to other classes of objects. Our proposed MaSSFormer is comprehensively evaluated on the MaSS13K benchmark together with 14 representative segmentation models. We expect that our meticulously annotated MaSS13K dataset and the MaSSFormer model can facilitate the research of high-resolution and high-quality semantic segmentation. Datasets and codes can be found at https://github.com/xiechenxi99/MaSS13K.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ab59fa27-655d-4915-a94e-48c8161fe772Builds on20
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Per-Pixel Classification is Not All You Need for Semantic SegmentationBowen Cheng, Alexander G. Schwing, Alexander KirillovNeurIPS 2021 · 2,196 citations
- SegNeXt: Rethinking Convolutional Attention Design for Semantic SegmentationMeng-Hao Guo, Cheng-Ze Lu, Qibin Hou, Zhengning Liu et al.NeurIPS 2022 · 1,385 citations
- Towards High-Resolution Salient Object DetectionYi Zeng, Pingping Zhang, Zhe Lin, Jianming Zhang et al.ICCV 2019 · 232 citations
Related papers
- High Quality Entity SegmentationLu Qi, Jason Kuen, Tiancheng Shen, Jiuxiang Gu et al.ICCV 2023 · 91 citations
- αMatte4K & µMatting: Dataset and Model for Ultra-Micro Precision Alpha Video MattingXinyi Chen, Hang Dong, Baowei Jiang, Shenkun Xu et al.CVPR 2026
- EFormer: Enhanced Transformer Towards Semantic-Contour Features of Foreground for Portraits MattingZitao Wang, Qiguang Miao, Yue Xi, Peipei ZhaoCVPR 2024
- Scaling up Image Segmentation across Data and TasksPei Wang, Zhaowei Cai, Hao Yang, Ashwin Swaminathan et al.CVPR 2025
- CascadePSP: Toward Class-Agnostic and Very High-Resolution Segmentation via Global and Local RefinementHo Kei Cheng, Jihoon Chung, Yu-Wing Tai, Chi-Keung TangCVPR 2020
