MiniMax-Remover: Taming Bad Noise Helps Video Object Removal
Bojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang, Shihao Zhao, Rong Xiao, Kam-Fai Wong
Abstract
Recent advances in video diffusion models have driven rapid progress in video editing techniques. However, video object removal, a critical subtask of video editing, remains challenging due to issues such as hallucinated objects and visual artifacts. Furthermore, existing methods often rely on computationally expensive sampling procedures and classifier-free guidance (CFG), resulting in slow inference. To address these limitations, we propose MiniMax-Remover, a novel two-stage video object removal approach. Motivated by the observation that text condition is not best suited for this task, we simplify the pretrained video generation model by removing textual input and cross-attention layers, resulting in a more lightweight and efficient model architecture in the first stage. In the second stage, we distilled our remover on successful videos produced by the stage-1 model and curated by human annotators, using a minimax optimization strategy to further improve editing quality and inference speed. Specifically, the inner maximization identifies adversarial input noise ("bad noise") that makes failure removals, while the outer minimization step trains the model to generate high-quality removal results even under such challenging conditions. As a result, our method achieves a state-of-the-art video object removal results with as few as 6 sampling steps and doesn't rely on CFG, significantly improving inference efficiency. Extensive experiments demonstrate the effectiveness and superiority of MiniMax-Remover compared to existing methods. Codes and Videos are available at: https://minimax-remover.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35608cb2-b40d-4560-9de4-40b4de18c889Cited by top-tier papers10
- EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect ErasingYANG FU, Yike Zheng, Ziyun Dai, Henghui DingCVPR 2026 · 14 citations
- EasyV2V: A High-quality Instruction-based Video Editing FrameworkJinjie Mai, Chaoyang Wang, Gordon Guocheng Qian, Willi Menapace et al.CVPR 2026 · 12 citations
- O-DisCo-Edit: Object Distortion Control for Unified Realistic Video EditingYuqing Chen, Junjie Wang, Lin Liu, Ruihang Chu et al.AAAI 2026 · 6 citations
- VideoCoF: Unified Video Editing with Temporal ReasonerXiangpeng Yang, Ji Xie, Yiyuan Yang, Yue Ma et al.CVPR 2026 · 6 citations
- YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object RemovalChenyang Wu, Lina Lei, Fan Li, Chunle Guo et al.CVPR 2026 · 4 citations
Builds on28
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
Related papers
- FasterCache: Training-Free Video Diffusion Model Acceleration with High QualityZhengyao Lv, Chenyang Si, Junhao Song, Zhenyu Yang et al.ICLR 2025
- ZeroPatcher: Training-free Sampler for Video Inpainting and EditingShaoshu Yang, Yingya Zhang, Ran HeNeurIPS 2025 · 2 citations
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 370 citations
- Object-WIPER: Training-Free Object and Associated Effect Removal in VideosSaksham Singh Kushwaha, Sayan Nag, Yapeng Tian, Kuldeep KulkarniCVPR 2026 · 5 citations
- On Distillation of Guided Diffusion ModelsChenlin Meng, Robin Rombach, Ruiqi Gao, Diederik P. Kingma et al.CVPR 2023
