RealText: Realistic Text Image Generation based on Glyph and Scene Aware Inpainting
Zihou Liu, Dongming Zhang, Jing Zhang, Jun Li, Yongdong Zhang
Abstract
Text-to-image generation models can create diverse, high-quality images, but they frequently encounter challenges in accurately rendering text within those images due to the insufficient representation of desired text. In this study, we introduce RealText, a method for generating scene text images that excels in producing precise and realistic scene text images in any language. We disentangle scene text images generation into three stages: background and glyph image generation, text deformation, and whole image generation. Initially, we utilize prompts to guide the creation of well-organized background images. By identifying optimal text placements on these backgrounds, we render the glyph images of target text using user-specified font, effectively eliminating incorrect characters. In the next stage, we propose scene sensing to perceive text carrier surfaces and viewpoints through 3D scene reconstruction using depth and normal map to apply text deformation, thereby enhancing the realism of generated images. The final stage involves generating complete image with the aid of background and glyph guidance. Thanks to glyph disentangling, scene sensing, and text inpainting, we can exert more precise control over scene text image generation process. We have developed a unified framework which supports major generation models. Extensive experiments illustrate the exceptional performance of our method in generating images with multilingual text. The codes will soon be available at https://github.com/cccvl/RealText.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 01eed94d-cc93-4620-b8e9-c526949e6cb3Related papers
- Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelLingjun Zhang, Xinyuan Chen, Yaohui Wang, Yue Lu et al.AAAI 2024 · 54 citations
- StyleTextGen: Style-Conditioned Multilingual Scene Text GenerationZeyu Chen, Fangmin Zhao, Yan Shu, Yichao Liu et al.CVPR 2026 · 4 citations
- Exploring Font-independent Features for Scene Text RecognitionYizhi Wang, Zhouhui LianACM MM 2020 · 19 citations
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Text2Scene: Text-driven Indoor Scene Stylization with Part-Aware DetailsInwoo Hwang, Hyeonwoo Kim, Young Min KimCVPR 2023
