Pre-Training Vision Models with Mandelbulb Variations
Benjamin Naoto Chiche, Yuto Horikawa, Ryo Fujita
Abstract
The use of models that have been pre-trained on natural image datasets like ImageNet may face some limitations. First, this use may be restricted due to copyright and license on the training images, and privacy laws. Second, these datasets and models may incorporate societal and ethical biases. Formula-driven supervised learning (FDSL) enables model pre-training to circumvent these issues. This consists of generating a synthetic image dataset based on mathematical formulae and pre-training the model on it. In this work, we propose novel FDSL datasets based on Mandelbulb Variations. These datasets contain RGB images that are projections of colored objects deriving from the 3D Mandelbulb fractal. Pre-training ResNet-50 on one of our proposed datasets MandelbulbVAR-1k enables an average top-1 accuracy over target classification datasets that is at least 1% higher than pre-training on existing FDSL datasets. With regard to anomaly detection on MVTec AD, pre-training the WideResNet-50 backbone on MandelbulbVAR-1k enables PatchCore to achieve 97.2% average image-level AUROC. This is only 1.9% lower than pre-training on ImageNet-1k (99.1%) and 4.5% higher than pre-training on the second-best performing FDSL dataset i.e. VisualAtom-1k (92.7%). Regarding Vision Transformer (ViT) pre-training, another dataset that we propose and coin MandelbulbVAR-Hybrid-21k enables ViT-Base to achieve 82.2% top-1 accuracy on ImageNet-1k, which is 0.4% higher than pre-training on ImageNet-21k (81.8%) and only 0.1% lower than pre-training on VisualAtom-1k (82.3%).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf et al.CVPR 2022 · 1,301 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Can Vision Transformers Learn without Natural Images?Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata et al.AAAI 2022 · 42 citations
- Point Cloud Pre-training with Natural 3D StructuresRyosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae et al.CVPR 2022 · 33 citations
Related papers
- Visual Atoms: Pre-Training Vision Transformers with Sinusoidal WavesSora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka et al.CVPR 2023
- Replacing Labeled Real-image Datasets with Auto-generated ContoursHirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima et al.CVPR 2022 · 32 citations
- Pre-training Vision Transformers with Very Limited Synthesized ImagesRyo Nakamura, Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega et al.ICCV 2023 · 14 citations
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 10 citations
- SegRCDB: Semantic Segmentation via Formula-Driven Supervised LearningRisa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue et al.ICCV 2023 · 16 citations
