OASIS: Online Sample Selection for Continual Instruction Tuning
Minjae Lee, Minhyuk Seo, Tingyu Qu, Tinne Tuytelaars, Jonghyun Choi
Abstract
In continual instruction tuning (CIT) scenarios, where new instruction tuning data continuously arrive in an online streaming manner, training delays from large-scale data significantly hinder real-time adaptation. Data selection can mitigate this overhead, but existing strategies often rely on pre-trained reference models, which are impractical in CIT setups since future data are unknown. Recent reference model-free online sample selection methods address this, but typically select a fixed number of samples per batch (e.g., top-k), making them vulnerable to distribution shifts where informativeness varies across batches. To address these limitations, we propose OASIS, an adaptive online sample selection approach for CIT that (1) selects informative samples by estimating each sample's informativeness relative to all previously seen data, beyond batch-level constraints, and (2) minimizes informative redundancy of selected samples through iterative selection score updates. Experiments on large foundation models show that OASIS, using only 25% of the data, achieves comparable performance to full-data training and outperforms the state-of-the-art sampling methods. ‡
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5bba0901-6fc2-4867-ad0a-96b1cfdc7829Builds on33
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- Overfitting in adversarially robust deep learningLeslie Rice, Eric Wong, J. Zico KolterICML 2020 · 935 citations
- Deep Learning on a Data Diet: Finding Important Examples Early in TrainingMansheej Paul, Surya Ganguli, Gintare Karolina DziugaiteNeurIPS 2021 · 806 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
Related papers
- OnEDIT: Online Editing with Decoupled Implicit Task for Large Language ModelsChae-Won Lee, Jae-Hong Lee, Ji-Hun Kang, Joon-Hyuk ChangAAAI 2026
- Adapt-∞: Scalable Continual Multimodal Instruction Tuning via Dynamic Data SelectionAdyasha Maharana, Jaehong Yoon, Tianlong Chen, Mohit BansalICLR 2025
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma et al.ICLR 2026 · 4 citations
- Don't Half-listen: Capturing Key-part Information in Continual Instruction TuningYongquan He, Wenyuan Zhang, Xuancheng Huang, Peng Zhang et al.ACL 2025
- TarGATE: Target-Aware Data Selection via Token-Attenuation GatesXiandi Luo, Shiwei Li, Haozhao Wang, Yihao Ouyang et al.ICML 2026
