CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Subtitle Removal
Qingdong He, Chaoyi Wang, Peng TANG, Yifan Yang, Xiaobin Hu
摘要
Video subtitle removal is essential for content localization and media re-editing, yet existing mask-guided diffusion methods face critical limitations: training inefficiency requiring extensive annotations and full model fine-tuning, inference complexity demanding explicit mask sequences, and static prior utilization unable to adapt to quality variations. We present CLEAR (Context-aware Learning for End-to-end Adaptive subtitle Removal), a lightweight adapter-based framework addressing these challenges through three technical innovations. First, self-supervised prior learning (Stage I) extracts occlusion guidance from video pairs using pixel differences as weak supervision, eliminating annotation dependency while learning generalizable subtitle features across languages. Second, LoRA-based adaptive refinement (Stage II) enables parameter-efficient training that preserves pre-trained visual priors while achieving true mask-free end-to-end inference without external detection modules. Third, adaptive focal weighting dynamically adjusts prior influence based on local quality assessment, effectively handling diverse subtitle styles and noisy guidance signals. Extensive experiments demonstrate CLEAR's superior performance in multilingual subtitle removal while requiring only 0.77% trainable parameters, establishing a new paradigm for efficient video text removal without inference-time mask dependencies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- ProPainter: Improving Propagation and Transformer for Video InpaintingShangchen Zhou, Chongyi Li, Kelvin C. K. Chan, Chen Change LoyICCV 2023 · 被引用 205 次
- MiniMax-Remover: Taming Bad Noise Helps Video Object RemovalBojia Zi, Weixuan Peng, Xianbiao Qi, Jianan Wang 等NeurIPS 2025 · 被引用 43 次
- Omni-Effects: Unified and Spatially-Controllable Visual Effects GenerationFangyuan Mao, Aiming Hao, Jintao Chen, Dongxia Liu 等AAAI 2026 · 被引用 20 次
相关 Paper
- Bilevel Layer-Positioning LoRA for Real Image DehazingYan Zhang, Long Ma, Yuxin Feng, Zhe Huang 等CVPR 2026 · 被引用 13 次
- EasyOmnimatte: Taming Pretrained Inpainting Diffusion Models for End-to-End Video Layered DecompositioYihan Hu, Xuelin Chen, Xiaodong CunCVPR 2026
- Just-Dub-It: Video dubbing via Joint Audio-Visual DiffusionAnthony Chen, Naomi Ken Korem, Tavi Halperin, Matan Ben-Yosef 等SIGGRAPH 2026 · 被引用 2 次
- Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-TuningChenjian Gao, Lihe Ding, Xin Cai, Zhanpeng Huang 等ICLR 2026 · 被引用 24 次
- One-Step Specular Highlight Removal with Adapted Diffusion ModelsMahir Atmis, Levent Karacan, Mehmet SarigülICCV 2025 · 被引用 1 次
