Exploring Stroke-Level Modifications for Scene Text Editing
Yadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu, YuXin Wang, Yongdong Zhang
摘要
Scene text editing (STE) aims to replace text with the desired one while preserving background and styles of the original text. However, due to the complicated background textures and various text styles, existing methods fall short in generating clear and legible edited text images. In this study, we attribute the poor editing performance to two problems: 1) Implicit decoupling structure. Previous methods of editing the whole image have to learn different translation rules of background and text regions simultaneously. 2) Domain gap. Due to the lack of edited real scene text images, the network can only be well trained on synthetic pairs and performs poorly on real-world images. To handle the above problems, we propose a novel network by MOdifying Scene Text image at strokE Level (MOSTEL). Firstly, we generate stroke guidance maps to explicitly indicate regions to be edited. Different from the implicit one by directly modifying all the pixels at image level, such explicit instructions filter out the distractions from background and guide the network to focus on editing rules of text regions. Secondly, we propose a Semi-supervised Hybrid Learning to train the network with both labeled synthetic images and unpaired real scene text images. Thus, the STE model is adapted to real-world datasets distributions. Moreover, two new datasets (Tamper-Syn2k and Tamper-Scene) are proposed to fill the blank of public evaluation datasets. Extensive experiments demonstrate that our MOSTEL outperforms previous methods both qualitatively and quantitatively. Datasets and code will be available at https://github.com/qqqyd/MOSTEL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper21
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 等NeurIPS 2023 · 被引用 290 次
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan 等NeurIPS 2023 · 被引用 61 次
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang 等NeurIPS 2024 · 被引用 55 次
- Text Image Inpainting via Global Structure-Guided Diffusion ModelsShipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao 等AAAI 2024 · 被引用 25 次
- Revisiting Tampered Scene Text Detection in the Era of Generative AIChenfan Qu, Yiwu Zhong, Fengjun Guo, Lianwen JinAAAI 2025 · 被引用 20 次
它引用的顶会 Paper6
- What Is Wrong With Scene Text Recognition Model Comparisons? Dataset and Model AnalysisJeonghun Baek, Geewook Kim, Junyeop Lee, Sungrae Park 等ICCV 2019 · 被引用 551 次
- Controllable Artistic Text Style Transfer via Shape-Matching GANShuai Yang, Zhangyang Wang, Zhaowen Wang, Ning Xu 等ICCV 2019 · 被引用 110 次
- Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text DetectionXugong Qin, Yu Zhou, Youhui Guo, Dayan Wu 等ACM MM 2021 · 被引用 33 次
- STRIVE: Scene Text Replacement In VideosVijay Kumar B. G, Jeyasri Subramanian, Varnith Chordia, Eugene Bart 等ICCV 2021 · 被引用 14 次
- SwapText: Image Based Texts Transfer in ScenesQiangpeng Yang, Jun Huang, Wei LinCVPR 2020
相关 Paper
- Self-Supervised Cross-Language Scene Text EditingFuxiang Yang, Tonghua Su, Xiang Zhou, Donglin Di 等ACM MM 2023 · 被引用 2 次
- Self-Supervised Text Erasing with Controllable Image SynthesisGangwei Jiang, Shiyao Wang, Tiezheng Ge, Yuning Jiang 等ACM MM 2022 · 被引用 10 次
- Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context LearningHongxi Li, Tong Wang, WU CHENGJING, Tianbao Liu 等ICML 2026
- Self-supervised Scene Text Segmentation with Object-centric Layered Representations Augmented by Text RegionsYibo Wang, Yunhu Ye, Yuanpeng Mao, Yanwei Yu 等ACM MM 2022 · 被引用 2 次
- Recognition-Synergistic Scene Text EditingZhengyao Fang, Pengyuan Lyu, Jingjing Wu, Chengquan Zhang 等CVPR 2025
