DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations
Ruilu Wang, Yang Xue, Lianwen Jin
Abstract
Document Image Enhancement (DIE) remains challenging due to the prevalence of multiple degradations in document images captured by cameras. In this paper, we respond an interesting question: can the performance of pre-trained models and downstream DIE models be improved if they are bootstrapped using different degradation types of the same semantic samples and their high-dimensional features with ambiguous inter-class distance? To this end, we propose an effective contrastive learning paradigm for DIE — a Document image enhancement framework with Normalization and Latent Contrast (DocNLC). While existing DIE methods focus on eliminating one type of degradation, DocNLC considers the relationship between different types of degradation while utilizing both direct and latent contrasts to constrain content consistency, thus achieving a unified treatment of multiple types of degradation. Specifically, we devise a latent contrastive learning module to enforce explicit decorrelation of the normalized representations of different degradation types and to minimize the redundancy between them. Comprehensive experiments show that our method outperforms state-of-the-art DIE models in both pre-training and fine-tuning stages on four publicly available independent datasets. In addition, we discuss the potential benefits of DocNLC for downstream tasks. Our code is released at https://github.com/RylonW/DocNLC
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 955e5e79-7884-47d9-8014-05753ddcb170Cited by top-tier papers2
- Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionFangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang et al.ACM MM 2025 · 1 citation
- Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text RecognitionShihao Zou, Wei Wei, Leyang Xu, Kaihe Xu et al.AAAI 2026
Builds on9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun et al.ICML 2021 · 2,942 citations
- From Canonical Correlation Analysis to Self-supervised Graph Neural NetworksHengrui Zhang, Qitian Wu, Junchi Yan, David Wipf et al.NeurIPS 2021 · 319 citations
- Semantically Contrastive Learning for Low-Light Image EnhancementDong Liang, Ling Li, Mingqiang Wei, Shuo Yang et al.AAAI 2022 · 132 citations
Related papers
- Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language ModelsXin Li, Yunfei Wu, Xinghua Jiang, Zhihao Guo et al.CVPR 2024
- Iterative Prompt Learning for Unsupervised Backlit Image EnhancementZhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng et al.ICCV 2023 · 196 citations
- Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document EnhancementMohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten et al.AAAI 2023 · 31 citations
- SCOB: Universal Text Understanding via Character-wise Supervised Contrastive Learning with Online Text Rendering for Bridging Domain GapDaehee Kim, Yoonsik Kim, Donghyun Kim, Yumin Lim et al.ICCV 2023 · 4 citations
- Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image ModelsLei Wang, Jiabang He, Xing Xu, Ning Liu et al.AAAI 2023 · 3 citations
