STRIVE: Scene Text Replacement In Videos
Vijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart, Shaobo Fang, Kelly Guan, Raja Bala
Abstract
We propose replacing scene text in videos using deep style transfer and learned photometric transformations. Building on recent progress on still image text replacement, we present extensions that alter text while preserving the appearance and motion characteristics of the original video. Compared to the problem of still image text replacement, our method addresses additional challenges introduced by video, namely effects induced by changing lighting, motion blur, diverse variations in camera-object pose over time, and preservation of temporal consistency. We parse the problem into three steps. First, the text in all frames is normalized to a frontal pose using a spatio-temporal transformer network. Second, the text is replaced in a single reference frame using a state-of-art still-image text replacement method. Finally, the new text is transferred from the reference to remaining frames using a novel learned image transformation network that captures lighting and blur effects in a temporally consistent manner. Results on synthetic and challenging real videos show realistic text transfer, competitive quantitative and qualitative performance, and superior inference speed relative to alternatives. We introduce new synthetic and real-world datasets with paired text objects. To the best of our knowledge this is the first attempt at deep video text replacement.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Exploring Stroke-Level Modifications for Scene Text EditingYadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu et al.AAAI 2023 · 51 citations
- Detect Any AI-Counterfeited Text ImageChenfan Qu, Yiwu Zhong, Xuekang Zhu, Junchi Li et al.CVPR 2026
Builds on3
- Controllable Artistic Text Style Transfer via Shape-Matching GANShuai Yang, Zhangyang Wang, Zhaowen Wang, Ning Xu et al.ICCV 2019 · 110 citations
- SwapText: Image Based Texts Transfer in ScenesQiangpeng Yang, Jun Huang, Wei LinCVPR 2020
- STEFANN: Scene Text Editor Using Font Adaptive Neural NetworkPrasun Roy, Saumik Bhattacharya, Subhankar Ghosh, Umapada PalCVPR 2020
Related papers
- StyleMaster: Stylize Your Video with Artistic Generation and TranslationZixuan Ye, Huijuan Huang, Xintao Wang, Pengfei Wan et al.CVPR 2025
- Arbitrary Video Style Transfer via Multi-Channel CorrelationYingying Deng, Fan Tang, Weiming Dong, Haibin Huang et al.AAAI 2021 · 197 citations
- Unsupervised Coherent Video Cartoonization with Perceptual Motion ConsistencyZhenhuan Liu, Liang Li, Huajie Jiang, Xin Jin et al.AAAI 2022 · 7 citations
- Reenact Anything: Semantic Video Motion Transfer Using Motion-Textual InversionManuel Kansy, Jacek Naruniec, Christopher Schroers, Markus Gross et al.SIGGRAPH 2025 · 5 citations
- Preserving Global and Local Temporal Consistency for Arbitrary Video Style TransferXinxiao Wu, Jialu ChenACM MM 2020 · 14 citations
