Low-Bandwidth Self-Improving Transmission of Rare Training Data
Shilpa Anna George, Haithem Turki, Ziqiang Feng, Deva Ramanan, Padmanabhan Pillai, Mahadev Satyanarayanan
Abstract
A severe bandwidth mismatch between incoming sensor data rate and wireless backhaul bandwidth often exists on unmanned probes when collecting new training data for machine learning (ML). To overcome this mismatch, we describe a self-improving ML-based transmission system called Hawk. Starting from a weak model that is trained on just a few examples, it seamlessly pipelines semi-supervised learning, active learning, and transfer learning, with asynchronous bandwidth-sensitive data transmission to a distant human for labeling. When a significant number of true positives (TPs) have been labeled, Hawk trains an improved model to replace the old model. This iterative workflow, called Live Learning, continues until a sufficient number of TPs have been collected. For very rare events on challenging datasets, and bandwidths as low as 12 kbps, a team of 7 probes using Hawk discovers up to 87% of the TPs that could have been discovered via full preview, transmission and labeling of all mission data. Hawk also uses diversity sampling and few-shot learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- Curriculum Labeling: Revisiting Pseudo-Labeling for Semi-Supervised LearningPaola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, Vicente OrdonezAAAI 2021 · 362 citations
- Distribution Aligning Refinery of Pseudo-label for Imbalanced Semi-supervised LearningJaehyung Kim, Youngbum Hur, Sejun Park, Eunho Yang et al.NeurIPS 2020 · 209 citations
- Laplacian Regularized Few-Shot LearningImtiaz Masud Ziko, Jose Dolz, Eric Granger, Ismail Ben AyedICML 2020 · 205 citations
- Learning Rare Category Classifiers on a Tight Labeling BudgetRavi Teja Mullapudi, Fait Poms, William R. Mark, Deva Ramanan et al.ICCV 2021 · 17 citations
- Background Splitting: Finding Rare Classes in a Sea of BackgroundRavi Teja Mullapudi, Fait Poms, William R. Mark, Deva Ramanan et al.CVPR 2021
Related papers
- Self-Training for Sample-Efficient Active Learning for Text Classification with Pre-Trained Language ModelsChristopher Schröder, Gerhard HeyerEMNLP 2024 · 2 citations
- ATSO: Asynchronous Teacher-Student Optimization for Semi-Supervised Image SegmentationXinyue Huo, Lingxi Xie, Jianzhong He, Zijie Yang et al.CVPR 2021
- Extending the WILDS Benchmark for Unsupervised AdaptationShiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao et al.ICLR 2022 · 116 citations
- Self-supervised Label Augmentation via Input TransformationsHankook Lee, Sung Ju Hwang, Jinwoo ShinICML 2020 · 218 citations
- Active Learning for Multiple Target ModelsYing-Peng Tang, Sheng-Jun HuangNeurIPS 2022 · 7 citations
