ROSE: Remove Objects with Side Effects in Videos
Chenxuan Miao, Yutong Feng, Jianshu Zeng, Zixiang Gao, Hantang Liu, Yunfeng Yan, Donglian Qi, Xi Chen, Bin Wang, Hengshuang Zhao
摘要
Video object removal has achieved advanced performance due to the recent success of video generative models. However, when addressing the side effects of objects, e.g., their shadows and reflections, existing works struggle to eliminate these effects for the scarcity of paired video data as supervision. This paper presents ROSE, termed Remove Objects with Side Effects, a framework that systematically studies the object's effects on environment, which can be categorized into five common cases: shadows, reflections, light, translucency and mirror. Given the challenges of curating paired videos exhibiting the aforementioned effects, we leverage a 3D rendering engine for synthetic data generation. We carefully construct a fully-automatic pipeline for data preparation, which simulates a largescale paired dataset with diverse scenes, objects, shooting angles, and camera trajectories. ROSE is implemented as an video inpainting model built on diffusion transformer. To localize all object-correlated areas, the entire video is fed into the model for reference-based erasing. Moreover, additional supervision is introduced to explicitly predict the areas affected by side effects, which can be revealed through the differential mask between the paired videos. To fully investigate the model performance on various side effect removal, we presents a new benchmark, dubbed ROSE-Bench, incorporating both common scenarios and the five special side effects for comprehensive evaluation. Experimental results demonstrate that ROSE achieves superior performance compared to existing video object erasing models and generalizes well to real-world video scenarios. The project page is https://rose2025-inpaint.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect ErasingYANG FU, Yike Zheng, Ziyun Dai, Henghui DingCVPR 2026 · 被引用 14 次
- Precise Object and Effect Removal with Adaptive Target-Aware AttentionJixin Zhao, Zhouxia Wang, Peiqing Yang, Shangchen ZhouCVPR 2026 · 被引用 13 次
- SynMotion: Semantic-Visual Adaptation for Motion Customized Video GenerationShuai Tan, Biao Gong, Yujie Wei, Shiwei Zhang 等CVPR 2026 · 被引用 9 次
- NOVA: Sparse Control, Dense Synthesis for Pair-Free Video EditingTianlin Pan, Jiayi Dai, Chenpu Yuan, Zhengyao Lv 等CVPR 2026 · 被引用 3 次
- MiVE: Multiscale Vision-language features for reference-guided video EditingTong Wang, Meng Zou, WU CHENGJING, Xiaochao Qu 等ICML 2026
它引用的顶会 Paper18
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Object-WIPER: Training-Free Object and Associated Effect Removal in VideosSaksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep KulkarniCVPR 2026 · 被引用 5 次
- Omnimatte: Associating Objects and Their Effects in VideoErika Lu, Forrester Cole, Tali Dekel, Andrew Zisserman 等CVPR 2021
- Multi-subject Open-set Personalization in Video GenerationTsai-Shien Chen, Aliaksandr Siarohin, Willi Menapace, Yuwei Fang 等CVPR 2025
- VideoPainter: Any-length Video Inpainting and Editing with Plug-and-Play Context ControlYuxuan Bian, Zhaoyang Zhang, Xuan Ju, Mingdeng Cao 等SIGGRAPH 2025 · 被引用 11 次
- Generative Video PropagationShaoteng Liu, Tianyu Wang, Jui-Hsien Wang, Qing Liu 等CVPR 2025
