Visual Atoms: Pre-Training Vision Transformers with Sinusoidal Waves
Sora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka, Rio Yokota
Abstract
Formula-driven supervised learning (FDSL) has been shown to be an effective method for pre-training vision transformers, where ExFractalDB-21k was shown to exceed the pre-training effect of ImageNet-21k. These studies also indicate that contours mattered more than textures when pre-training vision transformers. However, the lack of a systematic investigation as to why these contour-oriented synthetic datasets can achieve the same accuracy as real datasets leaves much room for skepticism. In the present work, we develop a novel methodology based on circular harmonics for systematically investigating the design space of contour-oriented synthetic datasets. This allows us to efficiently search the optimal range of FDSL parameters and maximize the variety of synthetic images in the dataset, which we found to be a critical factor. When the resulting new dataset VisualAtom-21k is used for pre-training ViT-Base, the top-1 accuracy reached 83.7% when fine-tuning on ImageNet-1k. This is only 0.5% difference from the top-1 accuracy (84.2%) achieved by the JFT-300M pre-training, even though the scale of images is 1/14. Unlike JFT-300M which is a static dataset, the quality of synthetic datasets will continue to improve, and the current work is a testament to this possibility. FDSL is also free of the common issues associated with real images, e.g. privacy/copyright issues, labeling costs/errors, and ethical biases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 890f7ea0-890f-4f06-ac5e-54ce6a4c9b42Cited by top-tier papers5
- FreeMask: Synthetic Images with Dense Annotations Make Stronger Segmentation ModelsLihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi et al.NeurIPS 2023 · 94 citations
- A Vision Check-up for Language ModelsPratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Adrián Rodríguez-Muñoz et al.CVPR 2024 · 10 citations
- Can You Learn to See Without Images? Procedural Warm-Up for Vision TransformersZachary Shinnick, Liangze Jiang, Hemanth Saratchandran, Damien Teney et al.CVPR 2026 · 9 citations
- Pre-Training Vision Models with Mandelbulb VariationsBenjamin Naoto Chiche, Yuto Horikawa, Ryo FujitaCVPR 2024
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
Builds on14
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
Related papers
- Replacing Labeled Real-image Datasets with Auto-generated ContoursHirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima et al.CVPR 2022 · 32 citations
- Pre-training Vision Transformers with Very Limited Synthesized ImagesRyo Nakamura, Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega et al.ICCV 2023 · 14 citations
- Can Vision Transformers Learn without Natural Images?Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata et al.AAAI 2022 · 42 citations
- SegRCDB: Semantic Segmentation via Formula-Driven Supervised LearningRisa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue et al.ICCV 2023 · 16 citations
- PeCo: Perceptual Codebook for BERT Pre-training of Vision TransformersXiaoyi Dong, Jianmin Bao, Ting Zhang, Dongdong Chen et al.AAAI 2023 · 281 citations
