From Enhancement to Understanding: Build a Generalized Bridge for Low-Light Vision via Semantically Consistent Unsupervised Fine-Tuning
Sen Wang, Shao Zeng, Tianjun Gu, Zhizhong Zhang, Ruixin Zhang, Shouhong Ding, Jingyun Zhang, Jun Wang, Xin Tan, Yuan Xie, Lizhuang Ma
Abstract
Low-level enhancement and high-level visual understanding in low-light vision have traditionally been treated separately. Low-light enhancement improves image quality for downstream tasks, but existing methods rely on physical or geometric priors, limiting generalization. Evaluation mainly focuses on visual quality rather than downstream performance. Low-light visual understanding, constrained by scarce labeled data, primarily uses task-specific domain adaptation, which lacks scalability. To address these challenges, we build a generalized bridge between low-light enhancement and low-light understanding, which we term Generalized Enhancement For Understanding (GEFU). This paradigm improves both generalization and scalability. To address the diverse causes of low-light degradation, we leverage pretrained generative diffusion models to optimize images, achieving zero-shot generalization performance. Building on this, we propose Semantically Consistent Unsupervised Fine-tuning (SCUF). Specifically, to overcome text prompt limitations, we introduce an illumination-aware image prompt to explicitly guide image generation and propose a cycle-attention adapter to maximize its semantic potential. To mitigate semantic degradation in unsupervised training, we propose caption and reflectance consistency to learn high-level semantics and image-level spatial semantics. Extensive experiments demonstrate that our proposed method outperforms current state-of-the-art methods in traditional image quality and GEFU tasks including classification, detection, and semantic segmentation. The code is available at GEFU.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a8089c40-0bd0-4250-8af0-4da3ba3dc7d0Cited by top-tier papers5
- Explore with Long-term Memory: A Benchmark and Multimodal LLM-based Reinforcement Learning Framework for Embodied ExplorationSen Wang, Bangwei Liu, Zhenkun Gao, Lizhuang Ma et al.CVPR 2026 · 14 citations
- FAPE-IR: Frequency-Aware Planning and Execution Framework for All-in-One Image RestorationJingren Liu, Shuning Xu, Qirui Yang, Yun Wang et al.CVPR 2026 · 4 citations
- MR. Illuminate: Zero-Shot Low-Light Image Enhancement with Diffusion PriorJoshua Cho, Sara Aghajanzadeh, Zhen Zhu, David ForsythCVPR 2026
- Towards Generalized Representations for Low-Light Understanding: When Signal Constancy Meets Semantic EnrichmentYifan Li, Haofeng Huang, Wenhan Yang, Jiaying LiuCVPR 2026
- L2DGS: Low-Light Dynamic Gaussian SplattingAshish Kumar, Rajagopalan AmbasamudramCVPR 2026
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
Related papers
- Zero-Reference Low-Light Enhancement via Physical Quadruple PriorsWenjing Wang, Huan Yang, Jianlong Fu, Jiaying LiuCVPR 2024 · 90 citations
- Zero-Shot Low-Light Image Enhancement via Latent Diffusion ModelsYan Huang, Xiaoshan Liao, Jinxiu Liang, Yuhui Quan et al.AAAI 2025 · 14 citations
- DLDA: Unified Dual-Level Domain Adaptation for Low-Light Object DetectionJiayi Hu, Qian Zhao, Gang LiAAAI 2026
- Cycle-Interactive Generative Adversarial Network for Robust Unsupervised Low-Light EnhancementZhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang et al.ACM MM 2022 · 41 citations
- TorchAdapt: Towards Light-Agnostic Real-Time Visual PerceptionKhurram Azeem Hashmi, Karthik Palyakere Suresh, Didier Stricker, Muhammad Zeshan AfzalICCV 2025 · 2 citations
