The Pursuit of Human Labeling: A New Perspective on Unsupervised Learning
Artyom Gadetsky, Maria Brbic
摘要
We present HUME, a simple model-agnostic framework for inferring human labeling of a given dataset without any external supervision. The key insight behind our approach is that classes defined by many human labelings are linearly separable regardless of the representation space used to represent a dataset. HUME utilizes this insight to guide the search over all possible labelings of a dataset to discover an underlying human labeling. We show that the proposed optimization objective is strikingly well-correlated with the ground truth labeling of the dataset. In effect, we only train linear classifiers on top of pretrained representations that remain fixed during training, making our framework compatible with any large pretrained and self-supervised model. Despite its simplicity, HUME outperforms a supervised linear classifier on top of self-supervised representations on the STL-10 dataset by a large margin and achieves comparable performance on the CIFAR-10 dataset. Compared to the existing unsupervised baselines, HUME achieves state-of-the-art performance on four benchmark image classification datasets including the large-scale ImageNet-1000 dataset. Altogether, our work provides a fundamentally new view to tackle unsupervised learning by searching for consistent labelings between different representation spaces.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Let Go of Your Labels with Unsupervised TransferArtyom Gadetsky, Yulun Jiang, Maria BrbicICML 2024 · 被引用 16 次
- Cross-domain Open-world DiscoveryShuo Wen, Maria BrbicICML 2024 · 被引用 8 次
- Fine-grained Classes and How to Find ThemMatej Grcic, Artyom Gadetsky, Maria BrbicICML 2024 · 被引用 5 次
- Large (Vision) Language Models are Unsupervised In-Context LearnersArtyom Gadetsky, Andrei Atanov, Yulun Jiang, Zhitong Gao 等ICLR 2025
它引用的顶会 Paper25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec 等NeurIPS 2020 · 被引用 9,171 次
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
相关 Paper
- HUME: Measuring the Human-Model Performance Gap in Text Embedding TasksAdnan El Assadi, Isaac Chung, Roman Solomatin, Niklas Muennighoff 等ICLR 2026 · 被引用 8 次
- DreamTeacher: Pretraining Image Backbones with Deep Generative ModelsDaiqing Li, Huan Ling, Amlan Kar, David Acuna 等ICCV 2023 · 被引用 37 次
- FreeSOLO: Learning to Segment Objects without AnnotationsXinlong Wang, Zhiding Yu, Shalini De Mello, Jan Kautz 等CVPR 2022 · 被引用 100 次
- One-Shot Exemplars for Class Grounding in Self-Supervised LearningHaowen Cui, Shuo Chen, Jun Li, Jian YangICLR 2026
- Unsupervised Visual Representation Learning via Mutual Information Regularized AssignmentDong Hoon Lee, Sungik Choi, Hyunwoo J. Kim, Sae-Young ChungNeurIPS 2022 · 被引用 9 次
