YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
Chenyang Wu, Lina Lei, Fan Li, Chunle Guo, Dehong Kong, Xinran Qin, Zhixin Wang, Ming-Ming Cheng, Chongyi Li
摘要
Recent advances in Diffusion Transformer (DiT)-based video generation technologies have shown impressive results for video object removal. However, these methods still suffer from substantial inference latency. For instance, although MiniMax Remover achieves state-of-the-art visual quality, it operates at only around 10FPS, primarily due to dense computations over the entire spatiotemporal token space, even when only a small masked region actually requires processing. In this paper, we present YOSE, You Only Select Essential Tokens, an efficient fine-tuning framework. YOSE introduces two key components: Batch Variable-length Indexing (BVI) and Diffusion Process Simulator (DiffSim) Module. BVI is a differentiable dynamic indexing operator that adaptively selects essential tokens based on mask information, enabling variable-length token processing across samples. DiffSim provides a diffusion process approximation mechanism for unmasked tokens, which simulates the influence of unmasked regions within DiT self-attention to maintain semantic consistency for masked tokens. With these designs, YOSE achieves mask-aware acceleration, where the inference time scales approximately linearly with the masked regions, in contrast to full-token diffusion methods whose computation remains constant regardless of the mask size. Extensive experiments demonstrate that YOSE achieves up to 2.5X speedup in 70% of cases while maintaining visual quality comparable to the baseline. Code is available at: https://github.com/Wucy0519/YOSE-CVPR26.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
- DiT4Edit: Diffusion Transformer for Image EditingKunyu Feng, Yue Ma, Bingyuan Wang, Chenyang Qi 等AAAI 2025 · 被引用 92 次
- VACE: All-in-One Video Creation and EditingZeyinzi Jiang, Zhen Han, Chaojie Mao, Jingfeng Zhang 等ICCV 2025 · 被引用 58 次
- MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalBojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang 等NeurIPS 2025 · 被引用 43 次
相关 Paper
- Accelerating Diffusion-based Video Editing via Heterogeneous Caching: Beyond Full Computing at Sampled Denoising TimestepTianyi Liu, Ye Lu, Linfeng Zhang, Chen Cai 等CVPR 2026 · 被引用 2 次
- Content-Aware Dynamic Patchification for Efficient Video DiffusionSheng Li, Connelly Barnes, Mamshad Nayeem Rizve, Hongwu Peng 等CVPR 2026
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang 等ICCV 2025 · 被引用 6 次
- VFRTok: Variable Frame Rates Video Tokenizer with Duration-Proportional Information AssumptionTianxiong Zhong, Xingye Tian, Boyuan Jiang, Xuebo Wang 等NeurIPS 2025 · 被引用 4 次
- Astraea: A Token-wise Acceleration Framework for Video Diffusion TransformersHaosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu 等ICLR 2026 · 被引用 9 次
