Info-Coevolution: An Efficient Framework for Data Model Coevolution
Ziheng Qin, Hailun Xu, Wei Chee Yew, Qi Jia, Yang Luo, Kanchan Sarkar, Danhui Guan, Kai Wang, Yang You
摘要
Machine learning relies heavily on data, yet the continuous growth of real-world data poses challenges for efficient dataset construction and training. A fundamental yet unsolved question is: given our current model and data, does a new data (sample/batch) need annotation/learning? Conventional approaches retain all available data, leading to non-optimal data and training efficiency. Active learning aims to reduce data redundancy by selecting a subset of samples to annotate, while it increases pipeline complexity and introduces bias. In this work, we propose Info-Coevolution, a novel framework that efficiently enables models and data to coevolve through online selective annotation with no bias. Leveraging task-specific models (and open-source models), it selectively annotates and integrates online and web data to improve datasets efficiently. For real-world datasets like ImageNet-1K, Info-Coevolution reduces annotation and training costs by 32% without performance loss. It is able to automatically give the saving ratio without tuning the ratio. It can further reduce the annotation ratio to 50% with semi-supervised learning. We also explore retrieval-based dataset enhancement using unlabeled open-source data. Code is available at https://github.com/NUS-HPC-AI-Lab/Info-Coevolution/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
相关 Paper
- When Dynamic Data Selection Meets Data Augmentation: Achieving Enhanced Training AccelerationSuorong Yang, Peng Ye, Furao Shen, Dongzhan ZhouICML 2025
- Active Finetuning: Exploiting Annotation Budget in the Pretraining-Finetuning ParadigmYichen Xie, Han Lu, Junchi Yan, Xiaokang Yang 等CVPR 2023
- Towards Free Data Selection with General-Purpose ModelsYichen Xie, Mingyu Ding, Masayoshi Tomizuka, Wei ZhanNeurIPS 2023 · 被引用 18 次
- Active Learning with Foundation Model Priors: Efficient Learning under Class ImbalanceJiancheng Zhang, Meiqing Li, Qi Zhang, Yinglun ZhuICML 2026
- Efficiency for Free: Ideal Data Are Transportable RepresentationsPeng Sun, Yi Jiang, Tao LinNeurIPS 2024 · 被引用 2 次
