StreamFP: Fingerprint-guided Data Selection for Efficient Stream Learning
Changwu Li, Tongjun Shi, Shuhao Zhang, Binbin Chen, Bingsheng He, Xiaofei Liao, Hai Jin
Abstract
Modern web applications—ranging from personalized recommendation to real-time fraud detection—rely on AI models to deliver timely and personalized services, yet the underlying user interaction data arrives as massive and evolving streams. Stream Learning (SL) offers a natural paradigm for building adaptive models, but it struggles with challenges such as redundant training data and catastrophic forgetting, which can undermine long-term predictive performance. To address these issues, recent studies have explored data selection strategies like coreset selection and buffer update, typically implemented through rule-based or model-based methods. However, fixed selection rules hinder the adaptability of rule-based approaches to changing data distributions, while model-based methods often depend on costly per-sample gradients, leading to throttled updates and reduced coverage of informative samples. In this paper, we propose StreamFP, a lightweight SL framework that introduces fingerprints-a set of compact, learnable parameter vectors that summarize the model state. Fingerprints compute similarity scores to jointly guide coreset selection and buffer update, prioritizing informative incoming samples while retaining representative historical ones. A lightweight fingerprint attunement plugin further calibrates fingerprints using pre-trained ViT attention with negligible overhead, thereby improving accuracy while mitigating forgetting. Extensive experiments demonstrate that StreamFP consistently achieves superior accuracy and efficiency compared with state-of-the-art methods across diverse real-world datasets and varying data arrival rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on20
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- An Empirical Study of Training Self-Supervised Vision TransformersXinlei Chen, Saining Xie, Kaiming HeICCV 2021 · 2,340 citations
Related papers
- Camel: Managing Data for Efficient Stream LearningYiming Li, Yanyan Shen, Lei ChenSIGMOD 2022 · 19 citations
- GPT4Rec: Graph Prompt Tuning for Streaming RecommendationPeiyan Zhang, Yuchen Yan, Xi Zhang, Liying Kang et al.SIGIR 2024 · 15 citations
- Salient Frequency-aware Exemplar Compression for Resource-constrained Online Continual LearningJunsu Kim, Suhyun KimAAAI 2025 · 1 citation
- Stable Learned Bloom Filters for Data StreamsQiyu Liu, Libin Zheng, Yanyan Shen, Lei ChenVLDB 2020 · 45 citations
- FreewayML: An Adaptive and Stable Streaming Learning Framework for Dynamic Data StreamsZheng Qin, Zheheng Liang, Lijie Xu, Wentao Wu et al.ICDE 2025 · 2 citations
