MatAnyone 2: Scaling Video Matting via a Learned Quality Evaluator
Peiqing Yang, Shangchen Zhou, Kai Hao, Qingyi Tao
Abstract
Video matting remains limited by the scale and realism of existing datasets. While leveraging segmentation data can enhance semantic stability, the lack of effective boundary supervision often leads to segmentation-like mattes lacking fine details. To this end, we introduce a learned Matting Quality Evaluator (MQE) that assesses semantic and boundary quality of alpha mattes without ground truth. It produces a pixel-wise evaluation map that identifies reliable and erroneous regions, enabling fine-grained quality assessment. The MQE scales up video matting in two ways: (1) as an online matting-quality feedback during training to suppress erroneous regions, providing comprehensive supervision, and (2) as an offline selection module for data curation, improving annotation quality by combining the strengths of leading video and image matting models. This process allows us to build a large-scale real-world video matting dataset, VMReal, containing 28K clips and 2.4M frames. To handle large appearance variations in long videos, we introduce a reference-frame training strategy that incorporates long-range frames beyond the local window for effective training. Our MatAnyone 2 achieves state-of-the-art performance on both synthetic and real-world benchmarks, surpassing prior methods across all metrics.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d815be1-36d9-412f-aeef-fb260c5a9875Builds on26
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Segment Everything Everywhere All at OnceXueyan Zou, Jianwei Yang, Hao Zhang, Feng Li et al.NeurIPS 2023 · 889 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- MatAnyone: Stable Video Matting with Consistent Memory PropagationPeiqing Yang, Shangchen Zhou, Jixin Zhao, Qingyi Tao et al.CVPR 2025
- VideoMaMa: Mask-Guided Video Matting via Generative PriorSangbeom Lim, Seoung Wug Oh, Gabriel Huang, Heeji Yoon et al.CVPR 2026 · 3 citations
- Generative Video MattingYongtao Ge, Kangyang Xie, Guangkai Xu, Li Ke et al.SIGGRAPH 2025 · 1 citation
- αMatte4K & µMatting: Dataset and Model for Ultra-Micro Precision Alpha Video MattingXinyi Chen, Hang Dong, Baowei Jiang, Shenkun Xu et al.CVPR 2026
- Mask-Guided Matting in the WildKwanyong Park, Sanghyun Woo, Seoung Wug Oh, In So Kweon et al.CVPR 2023
