Towards Free Data Selection with General-Purpose Models
Yichen Xie, Mingyu Ding, Masayoshi Tomizuka, Wei Zhan
Abstract
A desirable data selection algorithm can efficiently choose the most informative samples to maximize the utility of limited annotation budgets. However, current approaches, represented by active learning methods, typically follow a cumbersome pipeline that iterates the time-consuming model training and batch data selection repeatedly. In this paper, we challenge this status quo by designing a distinct data selection pipeline that utilizes existing general-purpose models to select data from various datasets with a single-pass inference without the need for additional training or supervision. A novel free data selection (FreeSel) method is proposed following this new pipeline. Specifically, we define semantic patterns extracted from inter-mediate features of the general-purpose model to capture subtle local information in each image. We then enable the selection of all data samples in a single pass through distance-based sampling at the fine-grained semantic pattern level. FreeSel bypasses the heavy batch selection process, achieving a significant improvement in efficiency and being 530x faster than existing active learning methods. Extensive experiments verify the effectiveness of FreeSel on various computer vision tasks. Our code is available at https://github.com/yichen928/FreeSel.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d597b2ed-6190-439e-a8ab-0a34d9e53a84Cited by top-tier papers10
- Class Balance Matters to Active Class-Incremental LearningZitong Huang, Ze Chen, Yuanze Li, Bowen Dong et al.ACM MM 2024 · 8 citations
- Boundary Matters: A Bi-Level Active Finetuning MethodHan Lu, Yichen Xie, Xiaokang Yang, Junchi YanNeurIPS 2024 · 7 citations
- ActiveDC: Distribution Calibration for Active FinetuningWenshuai Xu, Zhenghui Hu, Yu Lu, Jinzhou Meng et al.CVPR 2024 · 5 citations
- Debiased Active Learning with Variational Gradient RectifierWeiguo Chen, Changjian Wang, Shijun Li, Kele Xu et al.AAAI 2025 · 2 citations
- MDS-VQA: Model-Informed Data Selection for Video Quality AssessmentJian Zou, Xiaoyu Xu, Zhihua Wang, Yilin Wang et al.CVPR 2026 · 1 citation
Builds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- Active Finetuning: Exploiting Annotation Budget in the Pretraining-Finetuning ParadigmYichen Xie, Han Lu, Junchi Yan, Xiaokang Yang et al.CVPR 2023
- Enhancing Semi-Supervised Learning via Representative and Diverse Sample SelectionQian Shao, Jiangrui Kang, Qiyuan Chen, Zepeng Li et al.NeurIPS 2024 · 3 citations
- State-Relabeling Adversarial Active LearningBeichen Zhang, Liang Li, Shijie Yang, Shuhui Wang et al.CVPR 2020
- Similarity Search for Efficient Active Learning and Search of Rare ConceptsCody Coleman, Edward Chou, Julian Katz-Samuels, Sean Culatana et al.AAAI 2022 · 47 citations
- Influence Selection for Active LearningZhuoming Liu, Hao Ding, Huaping Zhong, Weijia Li et al.ICCV 2021 · 125 citations
