Camel: Managing Data for Efficient Stream Learning
Yiming Li, Yanyan Shen, Lei Chen
摘要
Many real-world applications rely on predictive models that are incrementally learned online. Specifically, models are updated with a single pass over continuously arriving data batches in a typical stream learning framework. However, this framework has three shortcomings: high training cost, low data effectiveness, and catastrophic forgetting. We describe Camel, a system that addresses the above issues. Camel includes two independent data management components: coreset selection and buffer update. To accelerate model training, Camel selects a coreset from each streaming data batch for model update. Selecting a coreset with worst-case guarantees is NP-hard. To solve this problem, we reformulate coreset selection as a submodular maximization problem by deriving an upper bound on the objective function. To mitigate catastrophic forgetting, Camel maintains a buffer of past representative samples as new data arrive. Moreover, Camel quantizes numerical data in buffer via a quantile sketch to reduce the memory footprint. Finally, extensive experiments validate the effectiveness and efficiency of Camel. In particular, our coreset selection algorithm can achieve a linear speedup with a marginal accuracy loss on redundant datasets. Furthermore, our buffer update algorithms can outperform the state-of-the-art methods for anti-forgetting on various data distributions.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper10
- QCore: Data-Efficient, On-Device Continual Calibration for Quantized ModelsDavid Campos, Bin Yang, Tung Kieu, Miao Zhang 等VLDB 2024 · 被引用 11 次
- BRIEF: Bi-level Coreset Selection for Efficient Instruction Tuning in LLMsChaoyuan Shen, Chi Zhang, Chengliang Chai, Jiacheng Wang 等VLDB 2026 · 被引用 2 次
- FreewayML: An Adaptive and Stable Streaming Learning Framework for Dynamic Data StreamsZheng Qin, Zheheng Liang, Lijie Xu, Wentao Wu 等ICDE 2025 · 被引用 2 次
- PECJ: Stream Window Join on Disorder Data Streams with Proactive Error CompensationXianzhi Zeng, Shuhao Zhang, Hongbin Zhong, Hao Zhang 等SIGMOD 2024 · 被引用 1 次
- C2TC: A Training-Free Framework for Efficient Tabular Data CondensationSijia Xu, Fan Li, Xiaoyang Wang, Zhengyi Yang 等ICDE 2026 · 被引用 1 次
相关 Paper
- StreamFP: Fingerprint-guided Data Selection for Efficient Stream LearningChangwu Li, Tongjun Shi, Shuhao Zhang, Binbin Chen 等WWW 2026
- GCR: Gradient Coreset based Replay Buffer Selection for Continual LearningRishabh Tiwari, KrishnaTeja Killamsetty, Rishabh K. Iyer, Pradeep ShenoyCVPR 2022 · 被引用 102 次
- Salient Frequency-aware Exemplar Compression for Resource-constrained Online Continual LearningJunsu Kim, Suhyun KimAAAI 2025 · 被引用 1 次
- Coreset Selection via Reducible Loss in Continual LearningRuilin Tong, Yuhang Liu, Javen Qinfeng Shi, Dong GongICLR 2025
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 被引用 181 次
