Deep Video Matting via Spatio-Temporal Alignment and Aggregation
Yanan Sun, Guanzhi Wang, Qiao Gu, Chi-Keung Tang, Yu-Wing Tai
Abstract
Despite the significant progress made by deep learning in natural image matting, there has been so far no representative work on deep learning for video matting due to the inherent technical challenges in reasoning temporal domain and lack of large-scale video matting datasets. In this paper, we propose a deep learning-based video matting framework which employs a novel and effective spatio-temporal feature aggregation module (ST-FAM). As optical flow estimation can be very unreliable within matting regions, ST-FAM is designed to effectively align and aggregate information across different spatial scales and temporal frames within the network decoder. To eliminate frame-by-frame trimap annotations, a lightweight interactive trimap propagation network is also introduced. The other contribution consists of a large-scale video matting dataset with groundtruth alpha mattes for quantitative evaluation and real-world highresolution videos with trimaps for qualitative evaluation. Quantitative and qualitative experimental results show that our framework significantly outperforms conventional video matting and deep image matting methods applied to video in presence of multi-frame temporal information. Our dataset is available at https://github.com/nowsyn/DVM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Tripartite Information Mining and Integration for Image MattingYuhao Liu, Jiake Xie, Xiao Shi, Yu Qiao et al.ICCV 2021 · 66 citations
- Attention-guided Temporally Coherent Video Object MattingYunke Zhang, Chi Wang, Miaomiao Cui, Peiran Ren et al.ACM MM 2021 · 31 citations
- BiMatting: Efficient Video Matting via BinarizationHaotong Qin, Lei Ke, Xudong Ma, Martin Danelljan et al.NeurIPS 2023 · 28 citations
- Human Instance Matting via Mutual Guidance and Multi-Instance RefinementYanan Sun, Chi-Keung Tang, Yu-Wing TaiCVPR 2022 · 23 citations
- FactorMatte: Redefining Video Matting for Re-Composition TasksZeqi Gu, Wenqi Xian, Noah Snavely, Abe DavisSIGGRAPH 2023 · 8 citations
Builds on7
- Indices Matter: Learning to Index for Deep Image MattingHao Lu, Yutong Dai, Chunhua Shen, Songcen XuICCV 2019 · 206 citations
- Natural Image Matting via Guided Contextual AttentionYaoyi Li, Hongtao LuAAAI 2020 · 189 citations
- Context-Aware Image Matting for Simultaneous Foreground and Alpha EstimationQiqi Hou, Feng LiuICCV 2019 · 171 citations
- MaskFlownet: Asymmetric Feature Matching With Learnable Occlusion MaskShengyu Zhao, Yilun Sheng, Yue Dong, Eric I-Chao Chang et al.CVPR 2020
- Boosting Semantic Human Matting With Coarse AnnotationsJinlin Liu, Yuan Yao, Wendi Hou, Miaomiao Cui et al.CVPR 2020
Related papers
- End-to-end Video Matting with Trimap PropagationWei-Lun Huang, Ming-Sui LeeCVPR 2023
- Long-Range Feature Propagating for Natural Image MattingQinglin Liu, Haozhe Xie, Shengping Zhang, Bineng Zhong et al.ACM MM 2021 · 31 citations
- High-Resolution Deep Image MattingHaichao Yu, Ning Xu, Zilong Huang, Yuqian Zhou et al.AAAI 2021 · 61 citations
- Semantic Image MattingYanan Sun, Chi-Keung Tang, Yu-Wing TaiCVPR 2021
- Video Matting via Consistency-Regularized Graph Neural NetworksTiantian Wang, Sifei Liu, Yapeng Tian, Kai Li et al.ICCV 2021 · 31 citations
