Pre-Training Vision Models with Mandelbulb Variations
Benjamin Naoto Chiche, Yuto Horikawa, Ryo Fujita
摘要
The use of models that have been pre-trained on natural image datasets like ImageNet may face some limitations. First, this use may be restricted due to copyright and license on the training images, and privacy laws. Second, these datasets and models may incorporate societal and ethical biases. Formula-driven supervised learning (FDSL) enables model pre-training to circumvent these issues. This consists of generating a synthetic image dataset based on mathematical formulae and pre-training the model on it. In this work, we propose novel FDSL datasets based on Mandelbulb Variations. These datasets contain RGB images that are projections of colored objects deriving from the 3D Mandelbulb fractal. Pre-training ResNet-50 on one of our proposed datasets MandelbulbVAR-1k enables an average top-1 accuracy over target classification datasets that is at least 1% higher than pre-training on existing FDSL datasets. With regard to anomaly detection on MVTec AD, pre-training the WideResNet-50 backbone on MandelbulbVAR-1k enables PatchCore to achieve 97.2% average image-level AUROC. This is only 1.9% lower than pre-training on ImageNet-1k (99.1%) and 4.5% higher than pre-training on the second-best performing FDSL dataset i.e. VisualAtom-1k (92.7%). Regarding Vision Transformer (ViT) pre-training, another dataset that we propose and coin MandelbulbVAR-Hybrid-21k enables ViT-Base to achieve 82.2% top-1 accuracy on ImageNet-1k, which is 0.4% higher than pre-training on ImageNet-21k (81.8%) and only 0.1% lower than pre-training on VisualAtom-1k (82.3%).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Total Recall in Industrial Anomaly DetectionKarsten Roth, Latha Pemula, Joaquin Zepeda, Bernhard Schölkopf 等CVPR 2022 · 被引用 1,301 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Can Vision Transformers Learn without Natural Images?Kodai Nakashima, Hirokatsu Kataoka, Asato Matsumoto, Kenji Iwata 等AAAI 2022 · 被引用 42 次
- Point Cloud Pre-training with Natural 3D StructuresRyosuke Yamada, Hirokatsu Kataoka, Naoya Chiba, Yukiyasu Domae 等CVPR 2022 · 被引用 33 次
相关 Paper
- Visual Atoms: Pre-Training Vision Transformers with Sinusoidal WavesSora Takashima, Ryo Hayamizu, Nakamasa Inoue, Hirokatsu Kataoka 等CVPR 2023
- Replacing Labeled Real-image Datasets with Auto-generated ContoursHirokatsu Kataoka, Ryo Hayamizu, Ryosuke Yamada, Kodai Nakashima 等CVPR 2022 · 被引用 32 次
- Pre-training Vision Transformers with Very Limited Synthesized ImagesRyo Nakamura, Hirokatsu Kataoka, Sora Takashima, Edgar Josafat Martinez-Noriega 等ICCV 2023 · 被引用 14 次
- Masked Image Residual Learning for Scaling Deeper Vision TransformersGuoxi Huang, Hongtao Fu, Adrian G. BorsNeurIPS 2023 · 被引用 10 次
- SegRCDB: Semantic Segmentation via Formula-Driven Supervised LearningRisa Shinoda, Ryo Hayamizu, Kodai Nakashima, Nakamasa Inoue 等ICCV 2023 · 被引用 16 次
