DocNLC: A Document Image Enhancement Framework with Normalized and Latent Contrastive Representation for Multiple Degradations
Ruilu Wang, Yang Xue, Lianwen Jin
摘要
Document Image Enhancement (DIE) remains challenging due to the prevalence of multiple degradations in document images captured by cameras. In this paper, we respond an interesting question: can the performance of pre-trained models and downstream DIE models be improved if they are bootstrapped using different degradation types of the same semantic samples and their high-dimensional features with ambiguous inter-class distance? To this end, we propose an effective contrastive learning paradigm for DIE — a Document image enhancement framework with Normalization and Latent Contrast (DocNLC). While existing DIE methods focus on eliminating one type of degradation, DocNLC considers the relationship between different types of degradation while utilizing both direct and latent contrasts to constrain content consistency, thus achieving a unified treatment of multiple types of degradation. Specifically, we devise a latent contrastive learning module to enforce explicit decorrelation of the normalized representations of different degradation types and to minimize the redundancy between them. Comprehensive experiments show that our method outperforms state-of-the-art DIE models in both pre-training and fine-tuning stages on four publicly available independent datasets. In addition, we discuss the potential benefits of DocNLC for downstream tasks. Our code is released at https://github.com/RylonW/DocNLC
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Uni-DocDiff: A Unified Document Restoration Model Based on DiffusionFangmin Zhao, Weichao Zeng, Zhenhang Li, Dongbao Yang 等ACM MM 2025 · 被引用 1 次
- Appearance Discrepancy-guided Sequence Hybrid Masking for Robust Scene Text RecognitionShihao Zou, Wei Wei, Leyang Xu, Kaihe Xu 等AAAI 2026
它引用的顶会 Paper9
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Barlow Twins: Self-Supervised Learning via Redundancy ReductionJure Zbontar, Li Jing, Ishan Misra, Yann LeCun 等ICML 2021 · 被引用 2,942 次
- From Canonical Correlation Analysis to Self-supervised Graph Neural NetworksHengrui Zhang, Qitian Wu, Junchi Yan, David Wipf 等NeurIPS 2021 · 被引用 319 次
- Semantically Contrastive Learning for Low-Light Image EnhancementDong Liang, Ling Li, Mingqiang Wei, Shuo Yang 等AAAI 2022 · 被引用 132 次
相关 Paper
- Enhancing Visual Document Understanding with Contrastive Learning in Large Visual-Language ModelsXin Li, Yunfei Wu, Xinghua Jiang, Zhihao Guo 等CVPR 2024
- Iterative Prompt Learning for Unsupervised Backlit Image EnhancementZhexin Liang, Chongyi Li, Shangchen Zhou, Ruicheng Feng 等ICCV 2023 · 被引用 196 次
- Text-DIAE: A Self-Supervised Degradation Invariant Autoencoder for Text Recognition and Document EnhancementMohamed Ali Souibgui, Sanket Biswas, Andrés Mafla, Ali Furkan Biten 等AAAI 2023 · 被引用 31 次
- SCOB: Universal Text Understanding via Character-wise Supervised Contrastive Learning with Online Text Rendering for Bridging Domain GapDaehee Kim, Yoonsik Kim, Donghyun Kim, Yumin Lim 等ICCV 2023 · 被引用 4 次
- Alignment-Enriched Tuning for Patch-Level Pre-trained Document Image ModelsLei Wang, Jiabang He, Xing Xu, Ning Liu 等AAAI 2023 · 被引用 3 次
