Associating Objects and Their Effects in Video through Coordination Games
Erika Lu, Forrester Cole, Weidi Xie, Tali Dekel, Bill Freeman, Andrew Zisserman, Michael Rubinstein
摘要
We explore a feed-forward approach for decomposing a video into layers, where each layer contains an object of interest along with its associated shadows, reflections, and other visual effects. This problem is challenging since associated effects vary widely with the 3D geometry and lighting conditions in the scene, and ground-truth labels for visual effects are difficult (and in some cases impractical) to collect. We take a self-supervised approach and train a neural network to produce a foreground image and alpha matte from a rough object segmentation mask under a reconstruction and sparsity loss. Under reconstruction loss, the layer decomposition problem is underdetermined: many combinations of layers may reconstruct the input video. Inspired by the game theory concept of focal points—or Schelling points —we pose the problem as a coordination game, where each player (network) predicts the effects for a single object without knowledge of the other players’ choices. The players learn to converge on the “natural” layer decomposition in order to maximize the likelihood of their choices aligning with the other players’. We train the network to play this game with itself, and show how to design the rules of this game so that the focal point lies at the correct layer decomposition. We demonstrate feed-forward results on a challenging synthetic dataset, then show that pretraining on this dataset significantly reduces optimization time for real videos.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- StableVideo: Text-driven Consistency-aware Diffusion Video EditingWenhao Chai, Xun Guo, Gaoang Wang, Yan LuICCV 2023 · 被引用 219 次
- CoDeF: Content Deformation Fields for Temporally Consistent Video ProcessingHao Ouyang, Qiuyu Wang, Yuxi Xiao, Qingyan Bai 等CVPR 2024 · 被引用 43 次
- Generative Omnimatte: Learning to Decompose Video into LayersYao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer 等CVPR 2025
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered DecompositioYihan Hu, Xuelin Chen, Xiaodong CunCVPR 2026
它引用的顶会 Paper5
- Video Object Segmentation Using Space-Time Memory NetworksSeoung Wug Oh, Joon-Young Lee, Ning Xu, Seon Joo KimICCV 2019 · 被引用 845 次
- Context-Aware Image Matting for Simultaneous Foreground and Alpha EstimationQiqi Hou, Feng LiuICCV 2019 · 被引用 171 次
- Controllable Attention for Structured Layered Video DecompositionJean-Baptiste Alayrac, João Carreira, Relja Arandjelovic, Andrew ZissermanICCV 2019 · 被引用 10 次
- Instance Shadow DetectionTianyu Wang, Xiaowei Hu, Qiong Wang, Pheng-Ann Heng 等CVPR 2020
- Background Matting: The World Is Your Green ScreenSoumyadip Sengupta, Vivek Jayaram, Brian Curless, Steven M. Seitz 等CVPR 2020
相关 Paper
- FactorMatte: Redefining Video Matting for Re-Composition TasksZeqi Gu, Wenqi Xian, Noah Snavely, Abe DavisSIGGRAPH 2023 · 被引用 8 次
- Omnimatte: Associating Objects and Their Effects in VideoErika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman 等CVPR 2021
- EyeIR: Single Eye Image Inverse Rendering In the WildShijun Liang, Haofei Wang, Feng LuSIGGRAPH 2024 · 被引用 1 次
- Separate in Latent Space: Unsupervised Single Image Layer SeparationYunfei Liu, Feng LuAAAI 2020 · 被引用 16 次
- Neural Spline Fields for Burst Image Fusion and Layer SeparationIlya Chugunov, David Shustin, Ruyu Yan, Chenyang Lei 等CVPR 2024
