Exploring Stroke-Level Modifications for Scene Text Editing
Yadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu, YuXin Wang, Yongdong Zhang
Abstract
Scene text editing (STE) aims to replace text with the desired one while preserving background and styles of the original text. However, due to the complicated background textures and various text styles, existing methods fall short in generating clear and legible edited text images. In this study, we attribute the poor editing performance to two problems: 1) Implicit decoupling structure. Previous methods of editing the whole image have to learn different translation rules of background and text regions simultaneously. 2) Domain gap. Due to the lack of edited real scene text images, the network can only be well trained on synthetic pairs and performs poorly on real-world images. To handle the above problems, we propose a novel network by MOdifying Scene Text image at strokE Level (MOSTEL). Firstly, we generate stroke guidance maps to explicitly indicate regions to be edited. Different from the implicit one by directly modifying all the pixels at image level, such explicit instructions filter out the distractions from background and guide the network to focus on editing rules of text regions. Secondly, we propose a Semi-supervised Hybrid Learning to train the network with both labeled synthetic images and unpaired real scene text images. Thus, the STE model is adapted to real-world datasets distributions. Moreover, two new datasets (Tamper-Syn2k and Tamper-Scene) are proposed to fill the blank of public evaluation datasets. Extensive experiments demonstrate that our MOSTEL outperforms previous methods both qualitatively and quantitatively. Datasets and code will be available at https://github.com/qqqyd/MOSTEL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cee2ea6b-5c3b-4d23-9d15-13e8b15f76d9Cited by top-tier papers21
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan et al.NeurIPS 2023 · 61 citations
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Text Image Inpainting via Global Structure-Guided Diffusion ModelsShipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao et al.AAAI 2024 · 25 citations
- Revisiting Tampered Scene Text Detection in the Era of Generative AIChenfan Qu, Yiwu Zhong, Fengjun Guo, Lianwen JinAAAI 2025 · 20 citations
Builds on6
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park et al.ICCV 2019 · 551 citations
- Controllable Artistic Text Style Transfer via Shape-Matching GANShuai Yang, Zhangyang Wang, Zhaowen Wang, Ning Xu et al.ICCV 2019 · 110 citations
- Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text DetectionXugong Qin, Yu Zhou, Youhui Guo, Dayan Wu et al.ACM MM 2021 · 33 citations
- STRIVE: Scene Text Replacement In VideosVijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart et al.ICCV 2021 · 14 citations
- SwapText: Image Based Texts Transfer in ScenesQiangpeng Yang, Jun Huang, Wei LinCVPR 2020
Related papers
- Self-Supervised Cross-Language Scene Text EditingFuxiang Yang, Tonghua Su, Xiang Zhou, Donglin Di et al.ACM MM 2023 · 2 citations
- Self-Supervised Text Erasing with Controllable Image SynthesisGangwei Jiang, Shiyao Wang, Tiezheng Ge, Yuning Jiang et al.ACM MM 2022 · 10 citations
- Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context LearningHongxi Li, Tong Wang, WU CHENGJING, Tianbao Liu et al.ICML 2026
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu et al.ACM MM 2022 · 2 citations
- Recognition-Synergistic Scene Text EditingZhengyao Fang, Pengyuan Lyu, Jingjing Wu, Chengquan Zhang et al.CVPR 2025
