Reinforcement learning for one-shot DAG scheduling with comparability identification and dense reward
Xumai Qi, Dongdong Zhang, Taotao Liu, Hongcheng Wang
摘要
In recent years, many studies proposed to generate solutions for Directed Acyclic Graph (DAG) scheduling problem in one shot by combining reinforcement learning and list scheduling heuristic. However, these existing methods suffer from biased estimation of sampling probabilities and inefficient guidance in training, due to redundant comparisons among node priorities and the sparse reward challenge. To address these issues, we analyze of the limitations of these existing methods, and propose a novel one-shot DAG scheduling method with comparability identification and dense reward signal, based on the policy gradient framework. In our method, a comparable antichain identification mechanism is proposed to eliminate the problem of redundant nodewise priority comparison. We also propose a dense reward signal for node level decision-making optimization in training, effectively addressing the sparse reward challenge. The experimental results show that the proposed method can yield superior results of scheduling objectives compared to other learning-based DAG scheduling methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper10
- Do Transformers Really Perform Badly for Graph Representation?Chengxuan Ying, Tianle Cai, Shengjie Luo, Shuxin Zheng 等NeurIPS 2021 · 被引用 1,632 次
- POMO: Policy Optimization with Multiple Optima for Reinforcement LearningYeong-Dae Kwon, Jinho Choo, Byoungjip Kim, Iljoo Yoon 等NeurIPS 2020 · 被引用 731 次
- Learning to Dispatch for Job Shop Scheduling via Deep Reinforcement LearningCong Zhang, Wen Song, Zhiguang Cao, Jie Zhang 等NeurIPS 2020 · 被引用 497 次
- Directed Acyclic Graph Neural NetworksVeronika Thost, Jie ChenICLR 2021 · 被引用 134 次
- Simulation-guided Beam Search for Neural Combinatorial OptimizationJinho Choo, Yeong-Dae Kwon, Jihoon Kim, Jeongwoo Jae 等NeurIPS 2022 · 被引用 123 次
相关 Paper
- Neural DAG Scheduling via One-Shot Priority SamplingWonseok Jeon, Mukul Gagrani, Burak Bartan, Weiliang Will Zeng 等ICLR 2023
- Reinforcement Causal Structure Learning on Order GraphDezhi Yang, Guoxian Yu, Jun Wang, Zhengtian Wu 等AAAI 2023 · 被引用 20 次
- MARLIN: Multi-Agent Reinforcement Learning for Incremental DAG DiscoveryDong Li, Zhengzhang Chen, Xujiang Zhao, Linlin Yu 等AAAI 2026
- Priority-Aware Attention Meets Generative Flow Networks for Global Fixed-Priority AssignmentShiwu Li, Jianjun Li, Quan Zhou, Yuan FuRTSS 2025
- Deep Reinforcement Learning Guided Improvement Heuristic for Job Shop SchedulingCong Zhang, Zhiguang Cao, Wen Song, Yaoxin Wu 等ICLR 2024 · 被引用 28 次
