GenHowTo: Learning to Generate Actions and State Transformations from Instructional Videos
Tomás Soucek, Dima Damen, Michael Wray, Ivan Laptev, Josef Sivic
摘要
We address the task of generating temporally consistent and physically plausible images of actions and object state transformations. Given an input image and a text prompt describing the targeted transformation, our generated images preserve the environment and transform objects in the initial image. Our contributions are threefold. First, we leverage a large body of instructional videos and automatically mine a dataset of triplets of consecutive frames corresponding to initial object states, actions, and resulting object transformations. Second, equipped with this data, we develop and train a conditioned diffusion model dubbed GenHowTo. Third, we evaluate GenHowTo on a variety of objects and actions and show superior performance compared to existing methods. In particular, we introduce a quantitative evaluation where GenHowTo achieves 88% and 74% on seen and unseen interaction categories, respectively, outperforming prior work by a large margin.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Mask2IV: Interaction-Centric Video Generation via Mask TrajectoriesGen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-LaraAAAI 2026 · 被引用 6 次
- The Promise of RL for Autoregressive Image EditingSaba Ahmadi, Rabiul Awal, Ankur Sikarwar, Amirhossein Kazemnejad 等NeurIPS 2025 · 被引用 6 次
- VG-TVP: Multimodal Procedural Planning via Visually Grounded Text-Video PromptingMuhammet Furkan Ilaslan, Ali Köksal, Kevin Qinghong Lin, Burak Satar 等AAAI 2025 · 被引用 3 次
- OSCBench: Benchmarking Object State Change in Text-to-Video GenerationXianjing Han, Bin Zhu, Shiqi Hu, Franklin Mingzhe Li 等ACL 2026 · 被引用 2 次
- Learning Procedural-Aware Video Representations Through State-Grounded Hierarchy UnfoldingJinghan Zhao, Yifei Huang, Feng LuAAAI 2026
它引用的顶会 Paper45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- Target-Aware Video Diffusion ModelsTaeksoo Kim, Hanbyul JooICLR 2026 · 被引用 7 次
- ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video GenerationMingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing 等CVPR 2026 · 被引用 8 次
- Auto-Regressive Diffusion for Generating 3D Human-Object InteractionsZichen Geng, Zeeshan Hayder, Wei Liu, Ajmal Saeed MianAAAI 2025 · 被引用 8 次
- GameGen-X: Interactive Open-world Game Video GenerationHaoxuan Che, Xuanhua He, Quande Liu, Cheng Jin 等ICLR 2025 · 被引用 2 次
- Flexible Motion In-betweening with Diffusion ModelsSetareh Cohan, Guy Tevet, Daniele Reda, Xue Bin Peng 等SIGGRAPH 2024 · 被引用 40 次
