Logarithmic Lenses: Exploring Log RGB Data for Image Classification
Bruce A. Maxwell, Sumegha Singhania, Avnish Patel, Rahul Kumar, Heather Fryling, Sihan Li, Haonan Sun, Ping He, Zewen Li
Abstract
The design of deep network architectures and training methods in computer vision has been well-explored. However, in almost all cases the images have been used as provided, with little exploration of pre-processing steps beyond normalization and data augmentation. Virtually all images posted on the web or captured by devices are processed for viewing by humans. Is the pipeline used for humans also best for use by computers and deep networks? The human visual system uses logarithmic sensors; differences and sums correspond to ratios and products. Features in log space will be invariant to intensity changes and robust to color balance changes. Log RGB space also reveals structure that is corrupted by typical pre-processing. We explore using linear and log RGB data for training standard backbone architectures on an image classification task using data derived directly from RAW images to guarantee its integrity. We found that networks trained on log RGB data exhibit improved performance on an unmodified test set and invariance to intensity and color balance modifications without additional training or data augmentation. Furthermore, we found that the gains from using high quality log data could also be partially or fully realized from data in 8-bit sRGB-JPG format by inverting the sRGB transform and taking the log. These results imply existing databases may benefit from this type of pre-processing. While working with log data, we found it was critical to retain the integrity of the log relationships and that networks using log data train best with meta-parameters different than those used for sRGB or linear data. Finally, we introduce a new 10-category 10k RAW image data set (RAW10) for image classification and other purposes to enable further the exploration of log RGB as an input format for deep networks in computer vision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46d72ced-577a-443c-bd0b-d8e1db070057Cited by top-tier papers2
- Beyond RGB: Adaptive Parallel Processing for RAW Object DetectionShani Gamrian, Hila Barel, Feiran Li, Masakazu Yoshimura et al.ICCV 2025 · 4 citations
- ReRAW: RGB-to-RAW Image Reconstruction via Stratified Sampling for Efficient Object Detection on the EdgeRadu Berdan, Beril Besbinar, Christoph Reinders, Junji Otsuka et al.CVPR 2025
Builds on1
Related papers
- Learning sRGB-to-Raw-RGB De-rendering with Content-Aware MetadataSeonghyeon Nam, Abhijith Punnappurath, Marcus A. Brubaker, Michael S. BrownCVPR 2022 · 16 citations
- Invertible Image Signal ProcessingYazhou Xing, Zian Qian, Qifeng ChenCVPR 2021
- Metadata-Based RAW Reconstruction via Implicit Neural FunctionsLeyi Li, Huijie Qiao, Qi Ye, Qinmin YangCVPR 2023
- Prior Metadata-Driven RAW Reconstruction: Eliminating the Need for Per-Image MetadataWencheng Han, Chen Zhang, Yang Zhou, Wentao Liu et al.ACM MM 2024
- RawHDR: High Dynamic Range Image Reconstruction from a Single Raw ImageYunhao Zou, Chenggang Yan, Ying FuICCV 2023 · 36 citations
