A Two-Stage Data Selection Framework for Data-Efficient Model Training on Edge Devices
Chen Gong, Rui Xing, Zhenzhe Zheng, Fan Wu
Abstract
The demand for machine learning (ML) model training on edge devices is escalating due to data privacy and personalized service needs. However, we observe that current on-device model training is hampered by the under-utilization of on-device data, due to low training throughput, limited storage and diverse data importance. To improve data resource utilization, we propose a two-stage data selection framework Titan to select the most important data batch from streaming data for model training with guaranteed efficiency and effectiveness. Specifically, in the first stage, Titan filters out a candidate dataset with potentially high importance in a coarsegrained manner. In the second stage of fine-grained selection, we propose a theoretically optimal data selection strategy to identify the data batch with the highest model performance improvement to current training round. To further enhance time-and-resource efficiency, Titan leverages a pipeline to co-execute data selection and model training, and avoids resource conflicts by exploiting idle computing resources. We evaluate Titan on real-world edge devices and three representative edge computing tasks with diverse models and data modalities. Empirical results demonstrate that Titan achieves up to 43% reduction in training time and 6.2% increase in final accuracy with minor system overhead, such as data processing delay, memory footprint and energy consumption. CCS Concepts • Human-centered computing → Mobile computing; • Computing methodologies → Machine learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 86ff4bab-97bf-4df1-941a-97cc9113ba5cBuilds on18
- Coresets for Data-efficient Training of Machine Learning ModelsBaharan Mirzasoleiman, Jeff A. Bilmes, Jure LeskovecICML 2020 · 494 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- PyramidFL: a fine-grained client selection framework for efficient federated learningChenning Li, Xiao Zeng, Mi Zhang, Zhichao CaoMobiCom 2022 · 190 citations
- Online Coreset Selection for Rehearsal-based Continual LearningJaehong Yoon, Divyam Madaan, Eunho Yang, Sung Ju HwangICLR 2022 · 181 citations
- PMC: A Privacy-preserving Deep Learning Model Customization Framework for Edge ComputingBingyan Liu, Yuanchun Li, Yunxin Liu, Yao Guo et al.UbiComp 2021 · 96 citations
Related papers
- To Store or Not? Online Data Selection for Federated Learning with Limited StorageChen Gong, Zhenzhe Zheng, Fan Wu, Yunfeng Shao et al.WWW 2023 · 28 citations
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeYoung D. Kwon, Rui Li, Stylianos I. Venieris, Jagmohan Chauhan et al.ICML 2024 · 25 citations
- Context-Aware Compilation of DNN Training Pipelines across Edge and CloudDixi Yao, Liyao Xiang, Zifan Wang, Jiayu Xu et al.UbiComp 2022 · 25 citations
- EdgeMove: Pipelining Device-Edge Model Training for Mobile IntelligenceZeqian Dong, Qiang He, Feifei Chen, Hai Jin et al.WWW 2023 · 12 citations
- FedEL: Federated Elastic Learning for Heterogeneous DevicesLetian Zhang, Bo Chen, Jieming Bian, Lei Wang et al.NeurIPS 2025 · 7 citations
