Attention-guided Temporally Coherent Video Object Matting
Yunke Zhang, Chi Wang, Miaomiao Cui, Peiran Ren, Xuansong Xie, Xian-Sheng Hua, Hujun Bao, Qixing Huang, Weiwei Xu
Abstract
This paper proposes a novel deep learning-based video object matting method that can achieve temporally coherent matting results.
Its key component is an attention-based temporal aggregation module that maximizes image matting networks' strength for video matting networks. This module computes temporal correlations for pixels adjacent to each other along the time axis in feature space, which is robust against motion noises. We also design a novel loss term to train the attention weights, which drastically boosts the video matting performance. Besides, we show how to effectively solve the trimap generation problem by fine-tuning a state-of-the-art video object segmentation network with a sparse set of user-annotated keyframes. To facilitate video matting and trimap generation networks' training, we construct a large-scale video matting dataset with 80 training and 28 validation foreground video clips with ground-truth alpha mattes. Experimental results show that our method can generate high-quality alpha mattes for various videos featuring appearance change, occlusion, and fast motion. Our code and dataset can be found at: https://github.com/yunkezhang/TCVOM
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers15
- Tripartite Information Mining and Integration for Image MattingYuhao Liu, Jiake Xie, Xiao Shi, Yu Qiao et al.ICCV 2021 · 66 citations
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan et al.NeurIPS 2023 · 28 citations
- MatAnyone 2: Scaling Video Matting via a Learned Quality EvaluatorPeiqing Yang, Shangchen Zhou, Kai Hao, Qingyi TaoCVPR 2026 · 7 citations
- Video Generation with Stable Transparency via Shiftable RGB-A Distribution LearnerHaotian Dong, Wenjing Wang, Chen Li, Jing LYU et al.CVPR 2026 · 7 citations
- VideoMaMa: Mask-Guided Video Matting via Generative PriorSangbeom Lim, Seoung Wug Oh, Gabriel Huang, Heeji Yoon et al.CVPR 2026 · 3 citations
Builds on13
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 845 citations
- MODNet: Real-Time Trimap-Free Portrait Matting via Objective DecompositionZhanghan Ke, Jiayu Sun, Kaican Li, Qiong Yan et al.AAAI 2022 · 220 citations
- Indices Matter: Learning to Index for Deep Image MattingHao Lu, Yutong Dai, Chunhua Shen, Songcen XuICCV 2019 · 206 citations
- Natural Image Matting via Guided Contextual AttentionYaoyi Li, Hongtao LuAAAI 2020 · 189 citations
- Context-Aware Image Matting for Simultaneous Foreground and Alpha EstimationQiqi Hou, Feng LiuICCV 2019 · 171 citations
Related papers
- Deep Video Matting via Spatio-Temporal Alignment and AggregationYanan Sun, Guanzhi Wang, Qiao Gu, Chi-Keung Tang et al.CVPR 2021
- End-to-end Video Matting with Trimap PropagationWei-Lun Huang, Ming-Sui LeeCVPR 2023
- Adaptive Human Matting for Dynamic VideosChung-Ching Lin, Jiang Wang, Kun Luo, Kevin Lin et al.CVPR 2023
- Video Matting via Consistency-Regularized Graph Neural NetworksTiantian Wang, Sifei Liu, Yapeng Tian, Kai Li et al.ICCV 2021 · 31 citations
- αMatte4K & µMatting: Dataset and Model for Ultra-Micro Precision Alpha Video MattingXinyi Chen, Hang Dong, Baowei Jiang, Shenkun Xu et al.CVPR 2026
