Pre-training Vision Transformers with Very Limited Synthesized Images
Ryo Nakamura, Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega, Rio Yokota, Nakamasa Inoue
Abstract
Formula-driven supervised learning (FDSL) is a pre-training method that relies on synthetic images generated from mathematical formulae such as fractals. Prior work on FDSL has shown that pre-training vision transformers on such synthetic datasets can yield competitive accuracy on a wide range of downstream tasks. These synthetic images are categorized according to the parameters in the mathematical formula that generate them. In the present work, we hypothesize that the process for generating different instances for the same category in FDSL, can be viewed as a form of data augmentation. We validate this hypothesis by replacing the instances with data augmentation, which means we only need a single image per category. Our experiments shows that this one-instance fractal database (OFDB) performs better than the original dataset where instances were explicitly generated. We further scale up OFDB to 21,000 categories and show that it matches, or even surpasses, the model pre-trained on ImageNet-21k in ImageNet-1k fine-tuning. The number of images in OFDB is 21k, whereas ImageNet-21k has 14M. This opens new possibilities for pre-training vision transformers with much smaller datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- SegFormer: Simple and Efficient Design for Semantic Segmentation with TransformersEnze Xie, Wenhai Wang, Zhiding Yu, Anima Anandkumar et al.NeurIPS 2021 · 9,661 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
Related papers
- Replacing Labeled Real-image Datasets with Auto-generated ContoursHirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima et al.CVPR 2022 · 32 citations
- Visual Atoms: Pre-Training Vision Transformers with Sinusoidal WavesSora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka et al.CVPR 2023
- Can Vision Transformers Learn without Natural Images?Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata et al.AAAI 2022 · 42 citations
- Pre-Training Vision Models with Mandelbulb VariationsBenjamin Naoto Chiche, Yuto Horikawa, Ryo FujitaCVPR 2024
- Point Cloud Pre-training with Natural 3D StructuresRyosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae et al.CVPR 2022 · 33 citations
