Controllable Attention for Structured Layered Video Decomposition
Jean-Baptiste Alayrac, João Carreira, Relja Arandjelovic, Andrew Zisserman
Abstract
The objective of this paper is to be able to separate a video into its natural layers, and to control which of the separated layers to attend to. For example, to be able to separate reflections, transparency or object motion. We make the following three contributions: (i) we introduce a new structured neural network architecture that explicitly incorporates layers (as spatial masks) into its design. This improves separation performance over previous general purpose networks for this task; (ii) we demonstrate that we can augment the architecture to leverage external cues such as audio for controllability and to help disambiguation; and (iii) we experimentally demonstrate the effectiveness of our approach and training procedure with controlled experiments while also showing that the proposed model can be successfully applied to real-word applications such as reflection removal and action recognition in cluttered scenes.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman et al.ICCV 2021 · 188 citations
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 74 citations
- Associating Objects and Their Effects in Video through Coordination GamesErika Lu, Forrester Cole, Weidi Xie, Tali Dekel et al.NeurIPS 2022 · 8 citations
- Hashing Neural Video Decomposition with Multiplicative Residuals in Space-TimeCheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun, Hwann-Tzong ChenICCV 2023 · 5 citations
- Omnimatte: Associating Objects and Their Effects in VideoErika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman et al.CVPR 2021
Related papers
- Separate in Latent Space: Unsupervised Single Image Layer SeparationYunfei Liu, Feng LuAAAI 2020 · 16 citations
- Language-guided Image Reflection SeparationHaofeng Zhong, Yuchen Hong, Shuchen Weng, Jinxiu Liang et al.CVPR 2024 · 14 citations
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
- Generative Omnimatte: Learning to Decompose Video into LayersYao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer et al.CVPR 2025
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang et al.ACM MM 2020 · 1 citation
