Choose What You Need: Disentangled Representation Learning for Scene Text Recognition, Removal and Editing
Boqiang Zhang, Hongtao Xie, Zuan Gao, Yuxin Wang
Abstract
Scene text images contain not only style information (font, background) but also content information (character, texture). Different scene text tasks need different information, but previous representation learning methods use tightly coupled features for all tasks, resulting in sub-optimal performance. We propose a Disentangled Representation Learning framework (DARLING) aimed at disentangling these two types of features for improved adaptability in better addressing various downstream tasks (choose what you really need). Specifically, we synthesize a dataset of image pairs with identical style but different content. Based on the dataset, we decouple the two types of features by the supervision design. Clearly, we directly split the visual representation into style and content features, the content features are supervised by a text recognition loss, while an alignment loss aligns the style features in the image pairs. Then, style features are employed in reconstructing the counterpart image via an image decoder with a prompt that indicates the counterpart's content. Such an operation effectively decouples the features based on their distinctive properties. To the best of our knowledge, this is the first time in the field of scene text that disentangles the inherent properties of the text images. Our method achieves state-of-the-art performance in Scene Text Recognition, Removal, and Editing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8c20fb56-cbd6-42cd-9b68-1a375ab84d56Cited by top-tier papers13
- Graph-based Unsupervised Disentangled Representation Learning via Multimodal Large Language ModelsBaao Xie, Qiuyu Chen, Yunnan Wang, Zequn Zhang et al.NeurIPS 2024 · 15 citations
- How Control Information Influences Multilingual Text Image Generation and Editing?Boqiang Zhang, Zuan Gao, Yadong Qu, Hongtao XieNeurIPS 2024 · 11 citations
- OmniText: A Training-Free Generalist for Controllable Text-Image ManipulationAgus Gunawan, Samuel Teodoro, Yun Chen, Soo Ye Kim et al.ICLR 2026 · 3 citations
- CLEAR: Context-Aware Learning with End-to-End Mask-Free Inference for Adaptive Subtitle RemovalQingdong He, Chaoyi Wang, Peng TANG, Yifan Yang et al.ICML 2026 · 1 citation
- Mitigating Translationese Bias in Multilingual LLM-as-a-Judge via Disentangled Information BottleneckHongbin Zhang, Kehai Chen, Xuefeng Bai, Youcheng Pan et al.ICML 2026 · 1 citation
Builds on20
- LayoutLMv3: Pre-training for Document AI with Unified Text and Image MaskingYupan Huang, Tengchao Lv, Lei Cui, Yutong Lu et al.ACM MM 2022 · 606 citations
- TextDiffuser: Diffusion Models as Text PaintersJingye Chen, Yupan Huang, Tengchao Lv, Lei Cui et al.NeurIPS 2023 · 290 citations
- From Two to One: A New Scene Text Recognizer with Visual Language Modeling NetworkYuxin Wang, Hongtao Xie, Shancheng Fang, Jing Wang et al.ICCV 2021 · 184 citations
- AnyText: Multilingual Visual Text Generation and EditingYuxiang Tuo, Wangmeng Xiang, Jun-Yan He, Yifeng Geng et al.ICLR 2024 · 148 citations
- MomentDiff: Generative Video Moment Retrieval from Random to RealPandeng Li, Chen-Wei Xie, Hongtao Xie, Liming Zhao et al.NeurIPS 2023 · 113 citations
Related papers
- Text Style Transfer based on Multi-factor Disentanglement and MixtureAnna Zhu, Zhanhui Yin, Brian Kenji Iwana, Xinyu Zhou et al.ACM MM 2022 · 5 citations
- TripleFDS: Triple Feature Disentanglement and Synthesis for Scene Text EditingYuchen Bao, Yiting Wang, Wenjian Huang, Haowei Wang et al.AAAI 2026
- TextCtrl: Diffusion-based Scene Text Editing with Prior Guidance ControlWeichao Zeng, Yan Shu, Zhenhang Li, Dongbao Yang et al.NeurIPS 2024 · 55 citations
- Self-Supervised Cross-Language Scene Text EditingFuxiang Yang, Tonghua Su, Xiang Zhou, Donglin Di et al.ACM MM 2023 · 2 citations
- Recognition-Synergistic Scene Text EditingZhengyao Fang, Pengyuan Lyu, Jingjing Wu, Chengquan Zhang et al.CVPR 2025
