Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Edit via In-Context Learning
Hongxi Li, Tong Wang, WU CHENGJING, Tianbao Liu, Jiangtao Yao, Xiaochao Qu, Xinxiao Wu, Luoqi Liu, Ting Liu
Abstract
Scene text editing aims to modify text in a target region of an image while preserving surrounding background style and texture. Existing methods rely solely on image background information while neglecting the visual details of target regions, which discards stylistic features in the original text and essentially degrades the task to text rendering. Moreover, the conditions imposed by pre-trained glyph encoder limit the scope of editable text. To address these issues, this paper proposes a self-prompting scene text editing method that constructs style and glyph prompts directly from the original image, without introducing additional style or glyph encoders. We employ a twostage training strategy: the diffusion transformer is first trained on large-scale self-supervised data and then refined using a small set of paired images. By leveraging the in-context learning capability of the Multi-Modal Diffusion Transformer (MM-DiT), it achieves open-vocabulary and styleconsistent text editing. Experimental results on various languages demonstrate that our method achieves the state-of-the-art performance in both text accuracy and style consistency. Our project page: hongxiii.github.io/mstedit.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1aaeaf68-40e2-46bb-9b2b-7416ad67d79dBuilds on11
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Free-Form Image Inpainting With Gated ConvolutionJiahui Yu, Zhe Lin, Jimei Yang, Xiaohui Shen et al.ICCV 2019 · 1,990 citations
- Blended Diffusion for Text-driven Editing of Natural ImagesOmri Avrahami, Dani Lischinski, Ohad FriedCVPR 2022 · 670 citations
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
Related papers
- DiffUTE: Universal Text Editing Diffusion ModelHaoxing Chen, Zhuoer Xu, Zhangxuan Gu, Jun Lan et al.NeurIPS 2023 · 61 citations
- Recognition-Synergistic Scene Text EditingZhengyao Fang, Pengyuan Lyu, Jingjing Wu, Chengquan Zhang et al.CVPR 2025
- GlyphMastero: A Glyph Encoder for High-Fidelity Scene Text EditingTong Wang, Ting Liu, Xiaochao Qu, Chengjing Wu et al.CVPR 2025
- Exploring Stroke-Level Modifications for Scene Text EditingYadong Qu, Qingfeng Tan, Hongtao Xie, Jianjun Xu et al.AAAI 2023 · 51 citations
- QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video EditingTiancheng Shen, Zilong Huang, Xiangtai Li, Zhijie Lin et al.ICCV 2025 · 2 citations
