Compositional Video Synthesis with Action Graphs
Amir Bar, Roei Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, Amir Globerson
摘要
Videos of actions are complex signals containing rich compositional structure in space and time. Current video generation methods lack the ability to condition the generation on multiple coordinated and potentially simultaneous timed actions. To address this challenge, we propose to represent the actions in a graph structure called Action Graph and present the new "Action Graph To Video" synthesis task. Our generative model for this task (AG2Vid) disentangles motion and appearance features, and by incorporating a scheduling mechanism for actions facilitates a timely and coordinated video generation. We train and evaluate AG2Vid on the CATER and Something-Something V2 datasets, and show that the resulting videos have better visual quality and semantic consistency compared to baselines. Finally, our model demonstrates zero-shot abilities by synthesizing novel compositions of the learned actions. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper20
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone 等ICLR 2022 · 被引用 290 次
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 被引用 182 次
- RoboDreamer: Learning Compositional World Models for Robot ImaginationSiyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li 等ICML 2024 · 被引用 140 次
- Learning to Compose Visual RelationsNan Liu, Shuang Li, Yilun Du, Josh Tenenbaum 等NeurIPS 2021 · 被引用 98 次
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig 等NeurIPS 2023 · 被引用 93 次
它引用的顶会 Paper6
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 被引用 840 次
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 被引用 190 次
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn 等ICLR 2020 · 被引用 142 次
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 被引用 84 次
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu 等CVPR 2020
相关 Paper
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang 等NeurIPS 2025 · 被引用 1 次
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao 等ACM MM 2021 · 被引用 6 次
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer 等ACL 2020 · 被引用 24 次
- DisMo: Disentangled Motion Representations for Open-World Motion TransferThomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann 等NeurIPS 2025 · 被引用 11 次
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
