A Decade's Battle on Dataset Bias: Are We There Yet?
Zhuang Liu, Kaiming He
Abstract
We revisit the "dataset classification" experiment suggested by Torralba & Efros (2011) a decade ago, in the new era with large-scale, diverse, and hopefully less biased datasets as well as more capable neural network architectures. Surprisingly, we observe that modern neural networks can achieve excellent accuracy in classifying which dataset an image is from: e.g., we report 84.7% accuracy on held-out validation data for the three-way classification problem consisting of the YFCC, CC, and DataComp datasets. Our further experiments show that such a dataset classifier could learn semantic features that are generalizable and transferable, which cannot be explained by memorization. We hope our discovery will inspire the community to rethink issues involving dataset bias.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 11e9fb89-aa71-4a18-b80b-13397ddbb51bCited by top-tier papers20
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- Seeing What Matters: Generalizable AI-generated Video Detection with Forensic-Oriented AugmentationRiccardo Corvi, Davide Cozzolino, Ekta Prashnani, Shalini De Mello et al.NeurIPS 2025 · 25 citations
- Understanding Bias in Large-Scale Visual DatasetsBoya Zeng, Yida Yin, Zhuang LiuNeurIPS 2024 · 23 citations
- Drag-and-Drop LLMs: Zero-Shot Prompt-to-WeightsZhiyuan Liang, Dongwen Tang, Yuhao Zhou, Xuanlei Zhao et al.NeurIPS 2025 · 22 citations
- In Pursuit of Pixel Supervision for Visual Pre-trainingLihe Yang, Shang-Wen Li, Yang Li, Xinjie Lei et al.CVPR 2026 · 13 citations
Builds on18
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
Related papers
- Can Deep Learning Recognize Subtle Human Activities?Vincent Jacquot, Zhuofan Ying, Gabriel KreimanCVPR 2020
- Object Classification From Randomized EEG TrialsHamad Ahmed, Ronnie B. Wilbur, Hari M. Bharadwaj, Jeffrey Mark SiskindCVPR 2021
- Learning to Imagine: Diversify Memory for Incremental Learning using Unlabeled DataYu-Ming Tang, Yi-Xing Peng, Wei-Shi ZhengCVPR 2022 · 33 citations
- Correlated Input-Dependent Label Noise in Large-Scale Image ClassificationMark Collier, Basil Mustafa, Efi Kokiopoulou, Rodolphe Jenatton et al.CVPR 2021
- Let's Agree to Agree: Neural Networks Share Classification Order on Real DatasetsGuy Hacohen, Leshem Choshen, Daphna WeinshallICML 2020 · 63 citations
