Classification Done Right for Vision-Language Pre-Training
Zilong Huang, Qinghao Ye, Bingyi Kang, Jiashi Feng, Haoqi Fan
Abstract
We introduce SuperClass, a super simple classification method for vision-language pre-training on image-text data. Unlike its contrastive counterpart CLIP who contrast with a text encoder, SuperClass directly utilizes tokenized raw text as supervised classification labels, without the need for additional text filtering or selection. Due to the absence of the text encoding as contrastive target, SuperClass does not require a text encoder and does not need to maintain a large batch size as CLIP does. SuperClass demonstrated superior performance on various downstream tasks, including classic computer vision benchmarks and vision language downstream tasks. We further explored the scaling behavior of SuperClass on model size, training length, or data size, and reported encouraging results and comparisons to CLIP. https://github.com/x-cls/superclass
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c2591dbc-0ecf-4c52-98e9-cc81b61b2fb8Cited by top-tier papers7
- EXP-Bench: Can AI Conduct AI Research Experiments?Patrick Tser Jern Kon, Qiuyi Ding, Jiachen Liu, Xinyi Zhu et al.ICLR 2026 · 35 citations
- SuperCLIP: CLIP with Simple Classification SupervisionWeiheng Zhao, Zilong Huang, Jiashi Feng, Xinggang WangNeurIPS 2025 · 6 citations
- The Scalability of Simplicity: Empirical Analysis of Vision-Language Learning with a Single TransformerWeixian Lei, Jiacong Wang, Haochen Wang, Xiangtai Li et al.ICCV 2025 · 3 citations
- Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training MethodsXuanru Zhou, Yiwen Shao, Wei-Cheng Tseng, Dong YuCVPR 2026 · 1 citation
- HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP ModelsZhixiang Wei, Guangting Wang, Xiaoxiao Ma, Ke Mei et al.ICCV 2025 · 1 citation
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- Is a Caption Worth a Thousand Images? A Study on Representation LearningShibani Santurkar, Yann Dubois, Rohan Taori, Percy Liang et al.ICLR 2023 · 9 citations
- RankCLIP: Ranking-Consistent Language-Image PretrainingYiming Zhang, Zhuokai Zhao, Zhaorun Chen, Zhili Feng et al.ICCV 2025 · 1 citation
- How Much Can CLIP Benefit Vision-and-Language Tasks?Sheng Shen, Liunian Harold Li, Hao Tan, Mohit Bansal et al.ICLR 2022 · 503 citations
- Non-Contrastive Learning Meets Language-Image Pre-TrainingJinghao Zhou, Li Dong, Zhe Gan, Lijuan Wang et al.CVPR 2023
- Scaling Language-Image Pre-Training via MaskingYanghao Li, Haoqi Fan, Ronghang Hu, Christoph Feichtenhofer et al.CVPR 2023
