STAR: A Structure-aware Lightweight Transformer for Real-time Image Enhancement
Zhaoyang Zhang, Yitong Jiang, Jun Jiang, Xiaogang Wang, Ping Luo, Jinwei Gu
Abstract
Image and video enhancement such as color constancy, low light enhancement, and tone mapping on smartphones is challenging, because high-quality images should be achieved efficiently with a limited resource budget. Unlike prior works that either used very deep CNNs or large Trans-former models, we propose a structure-aware lightweight Transformer, termed STAR, for real-time image enhancement. STAR is formulated to capture long-range dependencies between image patches, which naturally and implicitly captures the structural relationships of different regions in an image. STAR is a general architecture that can be easily adapted to different image enhancement tasks. Extensive experiments show that STAR can effectively boost the quality and efficiency of many tasks such as illumination enhancement, auto white balance, and photo retouching, which are indispensable components for image processing on smartphones. For example, STAR reduces model complexity and improves image quality compared to the recent state-of-the-art [19] on the MIT-Adobe FiveK dataset [7] (i.e., 1.8dB PSNR improvements with 25% parameters and 13% float operations.)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- URetinex-Net: Retinex-based Deep Unfolding Network for Low-light Image EnhancementWenhui Wu, Jian Weng, Pingping Zhang, Xu Wang et al.CVPR 2022 · 695 citations
- FeatEnHancer: Enhancing Hierarchical Features for Object Detection and Beyond Under Low-Light VisionKhurram Azeem Hashmi, Goutham Kallempudi, Didier Stricker, Muhammad Zeshan AfzalICCV 2023 · 76 citations
- Low-Light Image Enhancement with Illumination-Aware Gamma Correction and Complete Image Modelling NetworkYinglong Wang, Zhen Liu, Jianzhuang Liu, Songcen Xu et al.ICCV 2023 · 70 citations
- ShadowFormer: Global Context Helps Shadow RemovalLanqing Guo, Siyu Huang, Ding Liu, Hao Cheng et al.AAAI 2023 · 60 citations
- Low-Light Video Enhancement with Synthetic Event GuidanceLin Liu, Junfeng An, Jianzhuang Liu, Shanxin Yuan et al.AAAI 2023 · 51 citations
Builds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu et al.ICCV 2021 · 2,462 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
Related papers
- Transition-constant Normalization for Image EnhancementJie Huang, Man Zhou, Jinghao Zhang, Gang Yang et al.NeurIPS 2023 · 3 citations
- Structure- and Texture-Aware Learning for Low-Light Image EnhancementJinghao Zhang, Jie Huang, Mingde Yao, Man Zhou et al.ACM MM 2022 · 18 citations
- SNR-Aware Low-light Image EnhancementXiaogang Xu, Ruixing Wang, Chi-Wing Fu, Jiaya JiaCVPR 2022 · 552 citations
- RT-VENet: A Convolutional Network for Real-time Video EnhancementMohan Zhang, Qiqi Gao, Jinglu Wang, Henrik Turbell et al.ACM MM 2020 · 5 citations
- Equivalent Transformation and Dual Stream Network Construction for Mobile Image Super-ResolutionJiahao Chao, Zhou Zhou, Hongfan Gao, Jiali Gong et al.CVPR 2023
