Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
Gen Li, Bo Zhao, Jianfei Yang, Laura Sevilla-Lara
摘要
Generating interaction-centric videos, such as those depicting humans or robots interacting with objects, is crucial for embodied intelligence, as they provide rich and diverse visual priors for robot learning, manipulation policy training, and affordance reasoning. However, existing methods often struggle to model such complex and dynamic interactions. While recent studies show that masks can serve as effective control signals and enhance generation quality, obtaining dense and precise mask annotations remains a major challenge for real-world use. To overcome this limitation, we introduce Mask2IV, a novel framework specifically designed for interaction-centric video generation. It adopts a decoupled two-stage pipeline that first predicts plausible motion trajectories for both actor and object, then generates a video conditioned on these trajectories. This design eliminates the need for dense mask inputs from users while preserving the flexibility to manipulate the interaction process. Furthermore, Mask2IV supports versatile and intuitive control, allowing users to specify the target object of interaction and guide the motion trajectory through action descriptions or spatial position cues. To support systematic training and evaluation, we curate two benchmarks covering diverse action and object categories across both human-object interaction and robotic manipulation scenarios. Extensive experiments demonstrate that our method achieves superior visual realism and controllability compared to existing baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- ORV: 4D Occupancy-centric Robot Video GenerationXiuyu Yang, Bohan Li, Shaocong Xu, Nan Wang 等CVPR 2026 · 被引用 19 次
- Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware RepresentationHaodong Yan, Hang Yu, Zhide Zhong, Weilin Yuan 等CVPR 2026 · 被引用 5 次
它引用的顶会 Paper24
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu 等AAAI 2024 · 被引用 1,641 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
相关 Paper
- MIMIC: Mask-Injected Manipulation Video Generation with Interaction ControlTianxiao Chen, Jintao Rong, Huajin Chen, Jingya Wang 等ICLR 2026
- Motion Prompting: Controlling Video Generation with Motion TrajectoriesDaniel Geng, Charles Herrmann, Junhwa Hur, Forrester Cole 等CVPR 2025
- Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video GenerationGuy Yariv, Yuval Kirstain, Amit Zohar, Shelly Sheynin 等CVPR 2025
- Precise Action-to-Video Generation Through Visual Action PromptsYuang Wang, Chao Wen, Haoyu Guo, Sida Peng 等ICCV 2025 · 被引用 1 次
- MotionPro: A Precise Motion Controller for Image-to-Video GenerationZhongwei Zhang, Fuchen Long, Zhaofan Qiu, Yingwei Pan 等CVPR 2025
