TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine Learning
Yihong Li, Tianyu Zeng, Xiaoxi Zhang, Jingpu Duan, Chuan Wu
Abstract
Machine learning (ML) tasks are one of the major workloads in today's edge computing networks. Existing edge-cloud schedulers allocate the requested amounts of resources to each task, falling short of best utilizing the limited edge resources flexibly for ML task performance optimization. This paper proposes TapFinger, a distributed scheduler that minimizes the total completion time of ML tasks in a multi-cluster edge network, through co-optimizing task placement and fine-grained multi-resource allocation. To learn the tasks' uncertain resource sensitivity and enable distributed online scheduling, we adopt multi-agent reinforcement learning (MARL), and propose several techniques to make it efficient for our ML-task resource allocation. First, TapFinger uses a heterogeneous graph attention network as the MARL backbone to abstract inter-related state features into more learnable environmental patterns. Second, the actor network is augmented through a tailored task selection phase, which decomposes the actions and encodes the optimization constraints. Third, to mitigate decision conflicts among agents, we novelly combine Bayes' theorem and masking schemes to facilitate our MARL model training. Extensive experiments using synthetic and test-bed ML task traces show that TapFinger can achieve up to 28.6% reduction in the average task completion time and improve resource efficiency as compared to state-of-the- art resource schedulers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62edec38-27f9-48c9-a970-87b1bee4b726Cited by top-tier papers2
- A Two Time-Scale Joint Optimization Approach for UAV-assisted MECZemin Sun, Geng Sun, Long He, Fang Mei et al.INFOCOM 2024 · 18 citations
- FaaSConf: QoS-aware Hybrid Resources Configuration for Serverless WorkflowsYilun Wang, Pengfei Chen, Hui Dou, Yiwen Zhang et al.ASE 2024 · 3 citations
Builds on8
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang et al.OSDI 2020 · 260 citations
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger et al.OSDI 2021 · 258 citations
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen et al.SC 2021 · 136 citations
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin et al.NSDI 2021 · 111 citations
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang et al.INFOCOM 2021 · 93 citations
Related papers
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan et al.INFOCOM 2022 · 27 citations
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang et al.NSDI 2025 · 23 citations
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 9 citations
- Multi-Task Reinforcement Learning for Collaborative Network Optimization in Data CentersTing Wang, Kai Cheng, Xiao DuINFOCOM 2025 · 2 citations
- NeuRO: Inference-time Profiling and Orchestration of ML Applications at the EdgeArshad Javeed, György Dán, Viktoria FodorINFOCOM 2026 · 1 citation
