Compositional Video Synthesis with Action Graphs
Amir Bar, Roei Herzig, Xiaolong Wang, Anna Rohrbach, Gal Chechik, Trevor Darrell, Amir Globerson
Abstract
Videos of actions are complex signals containing rich compositional structure in space and time. Current video generation methods lack the ability to condition the generation on multiple coordinated and potentially simultaneous timed actions. To address this challenge, we propose to represent the actions in a graph structure called Action Graph and present the new "Action Graph To Video" synthesis task. Our generative model for this task (AG2Vid) disentangles motion and appearance features, and by incorporating a scheduling mechanism for actions facilitates a timely and coordinated video generation. We train and evaluate AG2Vid on the CATER and Something-Something V2 datasets, and show that the resulting videos have better visual quality and semantic consistency compared to baselines. Finally, our model demonstrates zero-shot abilities by synthesizing novel compositions of the learned actions. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7d40236b-51dd-4069-80fb-74d43902c73eCited by top-tier papers20
- Conditional Object-Centric Learning from VideoThomas Kipf, Gamaleldin Fathy Elsayed, Aravindh Mahendran, Austin Stone et al.ICLR 2022 · 290 citations
- Illiterate DALL-E Learns to ComposeGautam Singh, Fei Deng, Sungjin AhnICLR 2022 · 182 citations
- RoboDreamer: Learning Compositional World Models for Robot ImaginationSiyuan Zhou, Yilun Du, Jiaben Chen, Yandong Li et al.ICML 2024 · 140 citations
- Learning to Compose Visual RelationsNan Liu, Shuang Li, Yilun Du, Josh Tenenbaum et al.NeurIPS 2021 · 98 citations
- Dense and Aligned Captions (DAC) Promote Compositional Reasoning in VL ModelsSivan Doveh, Assaf Arbelle, Sivan Harary, Roei Herzig et al.NeurIPS 2023 · 93 citations
Builds on6
- Everybody Dance NowCaroline Chan, Shiry Ginosar, Tinghui Zhou, Alexei A. EfrosICCV 2019 · 840 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
- VideoFlow: A Conditional Flow-Based Model for Stochastic Video GenerationManoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn et al.ICLR 2020 · 142 citations
- Compositional Video PredictionYufei Ye, Maneesh Singh, Abhinav Gupta, Shubham TulsianiICCV 2019 · 84 citations
- Something-Else: Compositional Action Recognition With Spatial-Temporal Interaction NetworksJoanna Materzynska, Tete Xiao, Roei Herzig, Huijuan Xu et al.CVPR 2020
Related papers
- Prompt-guided Disentangled Representation for Action RecognitionTianci Wu, Guangming Zhu, Jiang Lu, Siyuan Wang et al.NeurIPS 2025 · 1 citation
- Video Representation Learning with Graph Contrastive AugmentationJingran Zhang, Xing Xu, Fumin Shen, Yazhou Yao et al.ACM MM 2021 · 6 citations
- Learning to Segment Actions from Observation and NarrationDaniel Fried, Jean-Baptiste Alayrac, Phil Blunsom, Chris Dyer et al.ACL 2020 · 24 citations
- DisMo: Disentangled Motion Representations for Open-World Motion TransferThomas Ressler-Antal, Frank Fundel, Malek Ben Alaya, Stefan Andreas Baumann et al.NeurIPS 2025 · 11 citations
- Self-Supervised Video GANs: Learning for Appearance Consistency and Motion CoherencySangeek Hyun, Jihwan Kim, Jae-Pil HeoCVPR 2021
