A Bayesian Approach to Data Point Selection
Xinnuo Xu, Minyoung Kim, Royson Lee, Brais Martínez, Timothy M. Hospedales
Abstract
Data point selection (DPS) is becoming a critical topic in deep learning due to the ease of acquiring uncurated training data compared to the difficulty of obtaining curated or processed data. Existing approaches to DPS are predominantly based on a bi-level optimisation (BLO) formulation, which is demanding in terms of memory and computation, and exhibits some theoretical defects regarding minibatches. Thus, we propose a novel Bayesian approach to DPS. We view the DPS problem as posterior inference in a novel Bayesian model where the posterior distributions of the instance-wise weights and the main neural network parameters are inferred under a reasonable prior and likelihood model. We employ stochastic gradient Langevin MCMC sampling to learn the main network and instance-wise weights jointly, ensuring convergence even with minibatches. Our update equation is comparable to the widely used SGD and much more efficient than existing BLO-based methods. Through controlled experiments in both the vision and language domains, we present the proof-of-concept. Additionally, we demonstrate that our method scales effectively to large language models and facilitates automated per-task optimization for instruction fine-tuning datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 811410a0-29c2-43e9-bbd2-3b9bed779f76Cited by top-tier papers2
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei et al.NeurIPS 2025 · 16 citations
- Fair Bayesian Data Selection via Generalized Discrepancy MeasuresYixuan Zhang, Jiabin Luo, Zhenggang Wang, Feng Zhou et al.AAAI 2026
Builds on25
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Multitask Prompted Training Enables Zero-Shot Task GeneralizationVictor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach et al.ICLR 2022 · 1,976 citations
Related papers
- A Stochastic Approach to Bi-Level Optimization for Hyperparameter Optimization and Meta LearningMinyoung Kim, Timothy M. HospedalesAAAI 2025 · 3 citations
- BayesTune: Bayesian Sparse Deep Model Fine-tuningMinyoung Kim, Timothy M. HospedalesNeurIPS 2023 · 10 citations
- A Contour Stochastic Gradient Langevin Dynamics Algorithm for Simulations of Multi-modal DistributionsWei Deng, Guang Lin, Faming LiangNeurIPS 2020 · 37 citations
- Microcanonical Langevin Ensembles: Advancing the Sampling of Bayesian Neural NetworksEmanuel Sommer, Jakob Robnik, Giorgi Nozadze, Uros Seljak et al.ICLR 2025
- Amortising Inference and Meta-Learning Priors in Neural NetworksTommy Rochussen, Vincent FortuinICLR 2026
