RGB No More: Minimally-Decoded JPEG Vision Transformers
Jeongsoo Park, Justin Johnson
摘要
Most neural networks for computer vision are designed to infer using RGB images. However, these RGB images are commonly encoded in JPEG before saving to disk; decoding them imposes an unavoidable overhead for RGB networks. Instead, our work focuses on training Vision Transformers (ViT) directly from the encoded features of JPEG. This way, we can avoid most of the decoding overhead, accelerating data load. Existing works have studied this aspect but they focus on CNNs. Due to how these encoded features are structured, CNNs require heavy modification to their architecture to accept such data. Here, we show that this is not the case for ViTs. In addition, we tackle data augmentation directly on these encoded features, which to our knowledge, has not been explored in-depth for training in this setting. With these two improvements -ViT and data augmentation -we show that our ViT-Ti model achieves up to 39.2% faster training and 17.9% faster inference with no accuracy loss compared to the RGB counterpart.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Faster Vision Transformers with Adaptive PatchesRohan Choudhury, JungEun Kim, Jinhyung Park, Eunho Yang 等ICLR 2026 · 被引用 8 次
- Proxy Compression for Language ModelingLin Zheng, Li Xinyu, Qian Liu, Xiachong Feng 等ICML 2026 · 被引用 3 次
- Wavelet-Driven Masked Image Modeling: A Path to Efficient Visual RepresentationWenzhao Xiang, Chang Liu, Hongyang Yu, Xilin ChenAAAI 2025 · 被引用 3 次
- Uncover Treasures in DCT: Advancing JPEG Quality Enhancement by Exploiting Latent CorrelationsJing Yang, Qunliang Xing, Mai Xu, Minglang QiaoICCV 2025
- JPEG Processing Neural Operator for Backward-Compatible CodingWoo Kyoung Han, Yongjun Lee, Byeonghun Lee, Sanghyun Park 等ICCV 2025
它引用的顶会 Paper16
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- MobileViT: Light-weight, General-purpose, and Mobile-friendly Vision TransformerSachin Mehta, Mohammad RastegariICLR 2022 · 被引用 2,162 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- EfficientFormer: Vision Transformers at MobileNet SpeedYanyu Li, Geng Yuan, Yang Wen, Ju Hu 等NeurIPS 2022 · 被引用 742 次
相关 Paper
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards End-to-End Image Compression and Analysis with TransformersYuanchao Bai, Xu Yang, Xianming Liu, Junjun Jiang 等AAAI 2022 · 被引用 68 次
- Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNetLi Yuan, Yunpeng Chen, Tao Wang, Weihao Yu 等ICCV 2021 · 被引用 2,462 次
- Understanding Robustness of Transformers for Image ClassificationSrinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li 等ICCV 2021 · 被引用 501 次
- Incorporating Convolution Designs into Visual TransformersKun Yuan, Shaopeng Guo, Ziwei Liu, Aojun Zhou 等ICCV 2021 · 被引用 581 次
