End-to-End Low-Light Enhancement for Object Detection with Learned Metadata from RAWs
Xuelin Shen, Haifeng Jiao, Yitong Wang, Yulin He, Wenhan Yang
Abstract
Although RAW images offer advantages over sRGB by avoiding ISP-induced distortion and preserving more information in low-light conditions, their widespread use is limited due to high storage costs, transmission burdens, and the need for significant architectural changes for downstream tasks. To address the issues, this paper explores a new raw-based machine vision paradigm, termed Compact RAW Metadata-guided Image Refinement (CRM-IR). In particular, we propose a Machine Vision-oriented Image Refinement (MV-IR) module that refines sRGB images to better suit machine vision preferences, guided by learned raw metadata. In detail, we propose a Cross-Modal Contextual Entropy (CMCE) network for raw metadata extraction and compression. It builds upon the latent representation and entropy modeling framework of learned image compression methods, and uniquely exploits the contextual correspondence between raw images and their sRGB counterparts to achieve more efficient and compact metadata representation. Additionally, we integrate priors derived from the ISP pipeline to simplify the refinement process, enabling a more efficient design. Such a design allows the CRM-IR to focus on extracting the most essential metadata from raw images to support downstream machine vision tasks, while remaining plug-and-play and fully compatible with existing imaging pipelines, without any changes to model architectures or ISP modules. We implement our CRM-IR scheme on various object detection networks, and extensive experiments under low-light conditions demonstrate that it can significantly improve performance with an additional bitrate cost of less than 10 -3 bits per pixel. Code is available at https://github.com/haifengjiao001/CRM-IR.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 10a3695b-fa1d-4f37-8db3-e4965eecd1e0Builds on15
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- FCOS: Fully Convolutional One-Stage Object DetectionZhi Tian, Chunhua Shen, Hao Chen, Tong HeICCV 2019 · 6,042 citations
- CenterNet: Keypoint Triplets for Object DetectionKaiwen Duan, Song Bai, Lingxi Xie, Honggang Qi et al.ICCV 2019 · 3,348 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- Enhanced Invertible Encoding for Learned Image CompressionYueqi Xie, Ka Leong Cheng, Qifeng ChenACM MM 2021 · 195 citations
Related papers
- Raw Image Reconstruction with Learned Compact MetadataYufei Wang, Yi Yu, Wenhan Yang, Lanqing Guo et al.CVPR 2023
- Prior Metadata-Driven RAW Reconstruction: Eliminating the Need for Per-Image MetadataWencheng Han, Chen Zhang, Yang Zhou, Wentao Liu et al.ACM MM 2024
- Learning sRGB-to-Raw-RGB De-rendering with Content-Aware MetadataSeonghyeon Nam, Abhijith Punnappurath, Marcus A. Brubaker, Michael S. BrownCVPR 2022 · 16 citations
- Task-Aware Image Signal Processor for Advanced Visual PerceptionKai Chen, Jin Xiao, Leheng Zhang, Kexuan Shi et al.CVPR 2026 · 3 citations
- Metadata-Based RAW Reconstruction via Implicit Neural FunctionsLeyi Li, Huijie Qiao, Qi Ye, Qinmin YangCVPR 2023
