Controllable Attention for Structured Layered Video Decomposition
Jean-Baptiste Alayrac, João Carreira, Relja Arandjelovic, Andrew Zisserman
摘要
The objective of this paper is to be able to separate a video into its natural layers, and to control which of the separated layers to attend to. For example, to be able to separate reflections, transparency or object motion. We make the following three contributions: (i) we introduce a new structured neural network architecture that explicitly incorporates layers (as spatial masks) into its design. This improves separation performance over previous general purpose networks for this task; (ii) we demonstrate that we can augment the architecture to leverage external cues such as audio for controllability and to help disambiguation; and (iii) we experimentally demonstrate the effectiveness of our approach and training procedure with controlled experiments while also showing that the proposed model can be successfully applied to real-word applications such as reflection removal and action recognition in cluttered scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Self-supervised Video Object Segmentation by Motion GroupingCharig Yang, Hala Lamdouar, Erika Lu, Andrew Zisserman 等ICCV 2021 · 被引用 188 次
- Segmenting Moving Objects via an Object-Centric Layered RepresentationJunyu Xie, Weidi Xie, Andrew ZissermanNeurIPS 2022 · 被引用 74 次
- Associating Objects and Their Effects in Video through Coordination GamesErika Lu, Forrester Cole, Weidi Xie, Tali Dekel 等NeurIPS 2022 · 被引用 8 次
- Hashing Neural Video Decomposition with Multiplicative Residuals in Space-TimeCheng-Hung Chan, Cheng-Yang Yuan, Cheng Sun, Hwann-Tzong ChenICCV 2023 · 被引用 5 次
- Omnimatte: Associating Objects and Their Effects in VideoErika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman 等CVPR 2021
相关 Paper
- Separate in Latent Space: Unsupervised Single Image Layer SeparationYunfei Liu, Feng LuAAAI 2020 · 被引用 16 次
- Language-guided Image Reflection SeparationHaofeng Zhong, Yuchen Hong, Shuchen Weng, Jinxiu Liang 等CVPR 2024 · 被引用 14 次
- Learning Semantic-Aware Dynamics for Video PredictionXinzhu Bei, Yanchao Yang, Stefano SoattoCVPR 2021
- Generative Omnimatte: Learning to Decompose Video into LayersYao-Chih Lee, Erika Lu, Sarah Rumbley, Michal Geyer 等CVPR 2025
- Multi-Group Multi-Attention: Towards Discriminative Spatiotemporal RepresentationZhensheng Shi, Liangjie Cao, Cheng Guan, Ju Liang 等ACM MM 2020 · 被引用 1 次
