TapFinger: Task Placement and Fine-Grained Resource Allocation for Edge Machine Learning
Yihong Li, Tianyu Zeng, Xiaoxi Zhang, Jingpu Duan, Chuan Wu
摘要
Machine learning (ML) tasks are one of the major workloads in today's edge computing networks. Existing edge-cloud schedulers allocate the requested amounts of resources to each task, falling short of best utilizing the limited edge resources flexibly for ML task performance optimization. This paper proposes TapFinger, a distributed scheduler that minimizes the total completion time of ML tasks in a multi-cluster edge network, through co-optimizing task placement and fine-grained multi-resource allocation. To learn the tasks' uncertain resource sensitivity and enable distributed online scheduling, we adopt multi-agent reinforcement learning (MARL), and propose several techniques to make it efficient for our ML-task resource allocation. First, TapFinger uses a heterogeneous graph attention network as the MARL backbone to abstract inter-related state features into more learnable environmental patterns. Second, the actor network is augmented through a tailored task selection phase, which decomposes the actions and encodes the optimization constraints. Third, to mitigate decision conflicts among agents, we novelly combine Bayes' theorem and masking schemes to facilitate our MARL model training. Extensive experiments using synthetic and test-bed ML task traces show that TapFinger can achieve up to 28.6% reduction in the average task completion time and improve resource efficiency as compared to state-of-the- art resource schedulers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- A Two Time-Scale Joint Optimization Approach for UAV-assisted MECZemin Sun, Geng Sun, Long He, Fang Mei 等INFOCOM 2024 · 被引用 18 次
- FaaSConf: QoS-aware Hybrid Resources Configuration for Serverless WorkflowsYilun Wang, Pengfei Chen, Hui Dou, Yiwen Zhang 等ASE 2024 · 被引用 3 次
它引用的顶会 Paper8
- AntMan: Dynamic Scaling on GPU Clusters for Deep LearningWencong Xiao, Shiru Ren, Yong Li, Yang Zhang 等OSDI 2020 · 被引用 260 次
- Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep LearningAurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger 等OSDI 2021 · 被引用 258 次
- Characterization and prediction of deep learning workloads in large-scale GPU datacentersQinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen 等SC 2021 · 被引用 136 次
- Elastic Resource Sharing for Distributed Deep LearningChangho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin 等NSDI 2021 · 被引用 111 次
- Tailored Learning-Based Scheduling for Kubernetes-Oriented Edge-Cloud SystemYiwen Han, Shihao Shen, Xiaofei Wang, Shiqiang Wang 等INFOCOM 2021 · 被引用 93 次
相关 Paper
- EdgeTuner: Fast Scheduling Algorithm Tuning for Dynamic Edge-Cloud Workloads and ResourcesRui Han, Shilin Wen, Chi Harold Liu, Ye Yuan 等INFOCOM 2022 · 被引用 27 次
- GREEN: Carbon-efficient Resource Scheduling for Machine Learning ClustersKaiqiang Xu, Decang Sun, Han Tian, Junxue Zhang 等NSDI 2025 · 被引用 23 次
- RESPECT: Reinforcement Learning based Edge Scheduling on Pipelined Coral Edge TPUsJiaqi Yin, Yingjie Li, Daniel Robinson, Cunxi YuDAC 2023 · 被引用 9 次
- Multi-Task Reinforcement Learning for Collaborative Network Optimization in Data CentersTing Wang, Kai Cheng, Xiao DuINFOCOM 2025 · 被引用 2 次
- NeuRO: Inference-time Profiling and Orchestration of ML Applications at the EdgeArshad Javeed, György Dán, Viktoria FodorINFOCOM 2026 · 被引用 1 次
