Adaptive Data Collection for Robust Learning Across Multiple Distributions
Chengbo Zang, Mehmet Kerem Türkcan, Gil Zussman, Zoran Kostic, Javad Ghaderi
Abstract
We propose a framework for adaptive data collection aimed at robust learning in multi-distribution scenarios under a fixed data collection budget. In each round, the algorithm selects a distribution source to sample from for data collection and updates the model parameters accordingly. The objective is to find the model parameters that minimize the expected loss across all the data sources. Our approach integrates upper-confidence-bound (UCB) sampling with online gradient descent (OGD) to dynamically collect and annotate data from multiple sources. By bridging online optimization and multi-armed bandits, we provide theoretical guarantees for our UCB-OGD approach, demonstrating that it achieves a minimax regret of O(T 1 2 (K ln T ) 1 2 ) over K data sources after T rounds. We further provide a lower bound showing that the result is optimal up to a ln T factor. Extensive evaluations on standard datasets and a real-world testbed for object detection in smartcity intersections validate the consistent performance improvements of our method compared to baselines such as random sampling and various active learning methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4f1d235c-e841-48b5-92f4-060a05ca8f93Cited by top-tier papers1
Ask how each one uses itBuilds on16
- Deep Batch Active Learning by Diverse, Uncertain Gradient Lower BoundsJordan T. Ash, Chicheng Zhang, Akshay Krishnamurthy, John Langford et al.ICLR 2020 · 974 citations
- Variational Adversarial Active LearningSamarth Sinha, Sayna Ebrahimi, Trevor DarrellICCV 2019 · 662 citations
- What matters when building vision-language models?Hugo Laurençon, Léo Tronchon, Matthieu Cord, Victor SanhNeurIPS 2024 · 401 citations
- Invariant Risk Minimization GamesKartik Ahuja, Karthikeyan Shanmugam, Kush R. Varshney, Amit DhurandharICML 2020 · 289 citations
- Active Learning for Deep Detection Neural NetworksHamed H. Aghdam, Abel Gonzalez-Garcia, Antonio M. López, Joost van de WeijerICCV 2019 · 155 citations
Related papers
- An active learning framework for multi-group mean estimationAbdellah Aznag, Rachel Cummings, Adam N. ElmachtoubNeurIPS 2023 · 8 citations
- Distributionally Robust Classification for Multi-source Unsupervised Domain AdaptationSeonghwi Kim, Sungho Jo, Wooseok Ha, Minwoo ChaeICLR 2026 · 4 citations
- Sample-Adaptivity Tradeoff in On-Demand SamplingNika Haghtalab, Omar Montasser, Mingda QiaoNeurIPS 2025 · 1 citation
- Leveraging (Biased) Information: Multi-armed Bandits with Offline DataWang Chi Cheung, Lixing LyuICML 2024 · 3 citations
- Budgeted Multi-Armed Bandits with Asymmetric Confidence IntervalsMarco Heyden, Vadim Arzamasov, Edouard Fouché, Klemens BöhmKDD 2024 · 1 citation
