TextNeRF: A Novel Scene-Text Image Synthesis Method Based on Neural Radiance Fields
Jialei Cui, Jianwei Du, Wenzhuo Liu, Zhouhui Lian
Abstract
Acquiring large-scale, well-annotated datasets is essential for training robust scene text detectors, yet the process is often resource-intensive and time-consuming. While some efforts have been made to explore the synthesis of scene text images, a notable gap remains between syn-thetic and authentic data. In this paper, we introduce a novel method that utilizes Neural Radiance Fields (NeRF) to model real-world scenes and emulate the data collection process by rendering images from diverse camera per-spectives, enriching the variability and realism of the synthesized data. A semi-supervised learning framework is proposed to categorize semantic regions within 3D scenes, ensuring consistent labeling of text regions across various viewpoints. Our method also models the pose, and view-dependent appearance of text regions, thereby offering precise control over camera poses and significantly improving the realism of text insertion and editing within scenes. Employing our technique on real-world scenes has led to the creation of a novel scene text image dataset (https://github.com/cuijl-ai/TextNeRF). Compared to other existing benchmarks, the proposed dataset is distinctive in providing not only standard annotations such as bounding boxes and transcriptions but also the information of 3D pose attributes for text regions, enabling a more detailed evaluation of the robustness of text detection algorithms. Through extensive experiments, we demonstrate the effectiveness of our proposed method in enhancing the performance of scene text detectors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85611852-783b-477a-ba4b-4ac0300aecafCited by top-tier papers1
Ask how each one uses itBuilds on9
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 4,089 citations
- Plenoxels: Radiance Fields without Neural NetworksSara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen et al.CVPR 2022 · 1,237 citations
- Direct Voxel Grid Optimization: Super-fast Convergence for Radiance Fields ReconstructionCheng Sun, Min Sun, Hwann-Tzong ChenCVPR 2022 · 859 citations
- Real-Time Scene Text Detection with Differentiable BinarizationMinghui Liao, Zhaoyi Wan, Cong Yao, Kai Chen et al.AAAI 2020 · 818 citations
Related papers
- Language-driven Object Fusion into Neural Radiance Fields with Pose-Conditioned Dataset UpdatesKa-Chun Shum, Jaeyeon Kim, Binh-Son Hua, Duc Thanh Nguyen et al.CVPR 2024 · 7 citations
- TextSSR: Diffusion-Based Data Synthesis for Scene Text RecognitionXingsong Ye, Yongkun Du, Yunbo Tao, Zhineng ChenICCV 2025 · 4 citations
- Empowering Sparse-Input Neural Radiance Fields with Dual-Level Semantic Guidance from Dense Novel ViewsYingji Zhong, Kaichen Zhou, Zhihao Li, Lanqing Hong et al.AAAI 2026 · 4 citations
- Pose-Free Neural Radiance Fields via Implicit Pose RegularizationJiahui Zhang, Fangneng Zhan, Yingchen Yu, Kunhao Liu et al.ICCV 2023 · 17 citations
- NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo CollectionsRicardo Martin-Brualla, Noha Radwan, Mehdi S. M. Sajjadi, Jonathan T. Barron et al.CVPR 2021
