Generative Video Matting
Yongtao Ge, Kangyang Xie, Guangkai Xu, Li Ke, Mingyu Liu, Longtao Huang, Hui Xue, Hao Chen, Chunhua Shen
摘要
Video matting has traditionally been limited by the lack of high-quality ground-truth data. Most existing video matting datasets provide only human-annotated imperfect alpha and foreground annotations, which must be composited to background images or videos during the training stage. Thus, the generalization capability of previous methods in real-world scenarios is typically poor. In this work, we propose to solve the problem from two perspectives. First, we emphasize the importance of large-scale pre-training by pursuing diverse synthetic and pseudo-labeled segmentation datasets. We also develop a scalable synthetic data generation pipeline that can render diverse human bodies and fine-grained hairs, yielding around 200 video clips with a 3-second duration for fine-tuning. Second, we introduce a novel video matting approach that can effectively leverage the rich priors from pre-trained video diffusion models. This architecture offers two key advantages. First, strong priors play a critical role in bridging the domain gap between synthetic and real-world scenes. Second, unlike most existing methods that process video matting frame-by-frame and use an independent decoder to aggregate temporal information, our model is inherently designed for video, ensuring strong temporal consistency. We provide a comprehensive quantitative evaluation across three benchmark datasets, demonstrating our approach's superior performance, and present comprehensive qualitative results in diverse real-world scenes, illustrating the strong generalization capability of our method. The code is available at https://github.com/aim-uofa/GVM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- MatAnyone 2: Scaling Video Matting via a Learned Quality EvaluatorPeiqing Yang, Shangchen Zhou, Kai Hao, Qingyi TaoCVPR 2026 · 被引用 7 次
- VideoMaMa: Mask-Guided Video Matting via Generative PriorSangbeom Lim, Seoung Wug Oh, Gabriel Huang, Heeji Yoon 等CVPR 2026 · 被引用 3 次
- Matting Anything 2: Towards Video Matting for AnythingChenyi Zhang, Yiheng Lin, Yunchao Wei, Hongsong Wang 等ICLR 2026
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- ActAnywhere: Subject-Aware Video Background GenerationBoxiao Pan, Zhan Xu, Chun-Hao Paul Huang, Krishna Kumar Singh 等NeurIPS 2024 · 被引用 10 次
- αMatte4K & µMatting: Dataset and Model for Ultra-Micro Precision Alpha Video MattingXinyi Chen, Hang Dong, Baowei Jiang, Shenkun Xu 等CVPR 2026
- Matting by GenerationZhixiang Wang, Baiang Li, Jian Wang, Yu-Lun Liu 等SIGGRAPH 2024 · 被引用 5 次
- Vivid-ZOO: Multi-View Video Generation with Diffusion ModelBing Li, Cheng Zheng, Wenxuan Zhu, Jinjie Mai 等NeurIPS 2024 · 被引用 48 次
- MAGICK: A Large-Scale Captioned Dataset from Matting Generated Images Using Chroma KeyingRyan D. Burgert, Brian L. Price, Jason Kuen, Yijun Li 等CVPR 2024
