TextSSR: Diffusion-Based Data Synthesis for Scene Text Recognition
Xingsong Ye, Yongkun Du, Yunbo Tao, Zhineng Chen
摘要
Scene text recognition (STR) suffers from challenges of either less realistic synthetic training data or the difficulty of collecting sufficient high-quality real-world data, limiting the effectiveness of trained models. Meanwhile, despite producing holistically appealing text images, diffusion-based visual text generation methods struggle to synthesize accurate and realistic instance-level text at scale. To tackle this, we introduce TextSSR: a novel pipeline for Synthesizing Scene Text Recognition training data. TextSSR targets three key synthesizing characteristics: accuracy, realism, and scalability. It achieves accuracy through a proposed region-centric text generation with position-glyph enhancement, ensuring proper character placement. It maintains realism by guiding style and appearance generation using contextual hints from surrounding text or background. This character-aware diffusion architecture enjoys precise character-level control and semantic coherence preservation, without relying on natural language prompts. Therefore, TextSSR supports large-scale generation through combinatorial text permutations. Based on these, we present TextSSR-F, a dataset of 3.55 million quality-screened text instances. Extensive experiments show that STR models trained on TextSSR-F outperform those trained on existing synthetic datasets by clear margins on common benchmarks, and further improvements are observed when mixed with real-world training data. Code is available at https: //github.com/YesianRohn/TextSSR.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- OmniText: A Training-Free Generalist for Controllable Text-Image ManipulationAgus Gunawan, Samuel Teodoro, Yun Chen, Soo Ye Kim 等ICLR 2026 · 被引用 3 次
- What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-EvolutionXingsong Ye, Yongkun Du, JiaXin Zhang, Chen Li 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper19
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui 等NeurIPS 2023 · 被引用 290 次
- GlyphControl: Glyph Conditional Control for Visual Text GenerationYukang Yang, Dongnan Gui, Yuhui Yuan, Weicong Liang 等NeurIPS 2023 · 被引用 163 次
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng 等ICLR 2024 · 被引用 148 次
相关 Paper
- Layout-Agnostic Scene Text Image Synthesis with Diffusion ModelsQilong Zhangli, Jindong Jiang, Di Liu, Licheng Yu 等CVPR 2024 · 被引用 7 次
- TextNeRF: A Novel Scene-Text Image Synthesis Method Based on Neural Radiance FieldsJialei Cui, Jianwei Du, Wenzhuo Liu, Zhouhui LianCVPR 2024
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang 等NeurIPS 2024 · 被引用 55 次
- Revisiting Scene Text Recognition: A Data PerspectiveQing Jiang, Jiapeng Wang, Dezhi Peng, Chongyu Liu 等ICCV 2023 · 被引用 70 次
- Brush Your Text: Synthesize Any Scene Text on Images via Diffusion ModelLingjun Zhang, Xinyuan Chen, Yaohui Wang, Yue Lu 等AAAI 2024 · 被引用 54 次
