CIFAR-10-Warehouse: Broad and More Realistic Testbeds in Model Generalization Analysis
Xiaoxiao Sun, Xingjian Leng, Zijian Wang, Yang Yang, Zi Huang, Liang Zheng
摘要
Analyzing model performance in various unseen environments is a critical research problem in the machine learning community. To study this problem, it is important to construct a testbed with out-of-distribution test sets that have broad coverage of environmental discrepancies. However, existing testbeds typically either have a small number of domains or are synthesized by image corruptions, hindering algorithm design that demonstrates real-world effectiveness. In this paper, we introduce CIFAR-10-Warehouse, consisting of 180 datasets collected by prompting image search engines and diffusion models in various ways. Generally sized between 300 and 8,000 images, the datasets contain natural images, cartoons, certain colors, or objects that do not naturally appear. With CIFAR-10-W, we aim to enhance the evaluation and deepen the understanding of two generalization tasks: domain generalization and model accuracy prediction in various out-of-distribution environments. We conduct extensive benchmarking and comparison experiments and show that CIFAR-10-W offers new and interesting insights inherent to these tasks. We also discuss other fields that would benefit from CIFAR-10-W. Data and code are available at https://sites.google.com/view/CIFAR-10-warehouse/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Energy-based Automated Model EvaluationRu Peng, Heming Zou, Haobo Wang, Yawen Zeng 等ICLR 2024 · 被引用 18 次
- Efficient Lifelong Model Evaluation in an Era of Rapid ProgressAmeya Prabhu, Vishaal Udandarao, Philip Torr, Matthias Bethge 等NeurIPS 2024 · 被引用 11 次
- A Theory-Inspired Framework for Few-Shot Cross-Modal Sketch Person Re-IdentificationYunpeng Gong, Yongjie Hou, Jiangming Shi, Kim Long Diep 等AAAI 2026 · 被引用 7 次
- Bounding Box Stability against Feature Dropout Reflects Detector Generalization across EnvironmentsYang Yang, Wenhai Wang, Zhe Chen, Jifeng Dai 等ICLR 2024 · 被引用 6 次
- Buffer layers for Test-Time AdaptationHyeongyu Kim, Geonhui Han, Dosik HwangNeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper18
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang 等ICCV 2019 · 被引用 2,239 次
- Tent: Fully Test-Time Adaptation by Entropy MinimizationDequan Wang, Evan Shelhamer, Shaoteng Liu, Bruno A. Olshausen 等ICLR 2021 · 被引用 1,731 次
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang 等ICML 2021 · 被引用 1,163 次
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville 等NeurIPS 2021 · 被引用 378 次
相关 Paper
- DiffGuard: Semantic Mismatch-Guided Out-of-Distribution Detection using Pre-trained Diffusion ModelsRuiyuan Gao, Chenchen Zhao, Lanqing Hong, Qiang XuICCV 2023 · 被引用 29 次
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 被引用 154 次
- DecAug: Out-of-Distribution Generalization via Decomposed Feature Representation and Semantic AugmentationHaoyue Bai, Rui Sun, Lanqing Hong, Fengwei Zhou 等AAAI 2021 · 被引用 88 次
- Improving Out-of-Distribution Detection with Disentangled Foreground and Background FeaturesChoubo Ding, Guansong PangACM MM 2024 · 被引用 1 次
- Rethinking the Evaluation Protocol of Domain GeneralizationHan Yu, Xingxuan Zhang, Renzhe Xu, Jiashuo Liu 等CVPR 2024
