ImageNet-D: Benchmarking Neural Network Robustness on Diffusion Synthetic Object
Chenshuang Zhang, Fei Pan, Junmo Kim, In So Kweon, Chengzhi Mao
摘要
We establish rigorous benchmarks for visual perception robustness. Synthetic images such as ImageNet-C, ImageNet-9, and Stylized ImageNet provide specific type of evaluation over synthetic corruptions, backgrounds, and textures, yet those robustness benchmarks are restricted in specified variations and have low synthetic quality. In this work, we introduce generative model as a data source for synthesizing hard images that benchmark deep models' robustness. Leveraging diffusion models, we are able to generate images with more diversified backgrounds, textures, and materials than any prior work, where we term this benchmark as ImageNet-D. Experimental results show that ImageNet-D results in a significant accuracy drop to a range of vision models, from the standard ResNet visual classifier to the latest foundation models like CLIP and MiniGPT-4, significantly reducing their accuracy by up to 60%. Our work suggests that diffusion models can be an effective source to test vision models. The code and dataset are available at https://github.com/chenshuang- zhang/imagenet_d.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- A Sober Look at the Robustness of CLIPs to Spurious FeaturesQizhou Wang, Yong Lin, Yongqiang Chen, Ludwig Schmidt 等NeurIPS 2024 · 被引用 46 次
- DASH: Detection and Assessment of Systematic Hallucinations of VLMsMaximilian Augustin, Yannic Neuhaus, Matthias HeinICCV 2025 · 被引用 17 次
- Improving robustness to corruptions with multiplicative weight perturbationsTrung Q. Trinh, Markus Heinonen, Luigi Acerbi, Samuel KaskiNeurIPS 2024 · 被引用 8 次
- GHOST: Hallucination-Inducing Image Generation for Multimodal LLMsAryan Yazdan Parast, Parsa Hosseini, Hesam Asadollahzadeh, Arshia Soltani Moakhar 等ICLR 2026 · 被引用 4 次
- Exploring Semantic-constrained Adversarial Example with Instruction Uncertainty ReductionJin Hu, Jiakai Wang, Linna Jing, Haolin Li 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
相关 Paper
- VLM-RobustBench: A Comprehensive Benchmark for Robustness of Vision-Language ModelsRohit Saxena, Alessandro Suglia, Pasquale MinerviniICML 2026 · 被引用 7 次
- Fake it Till You Make it: Learning Transferable Representations from Synthetic ImageNet ClonesMert Bülent Sariyildiz, Karteek Alahari, Diane Larlus, Yannis KalantidisCVPR 2023
- DIRE for Diffusion-Generated Image DetectionZhendong Wang, Jianmin Bao, Wengang Zhou, Weilun Wang 等ICCV 2023 · 被引用 479 次
- LAION-C: An Out-of-Distribution Benchmark for Web-Scale Vision ModelsFanfei Li, Thomas Klein, Wieland Brendel, Robert Geirhos 等ICML 2025
- Do Image Classifiers Generalize Across Time?Vaishaal Shankar, Achal Dave, Rebecca Roelofs, Deva Ramanan 等ICCV 2021 · 被引用 87 次
