MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Baheri, Betis; Tronge, Jacob; Fang, Bo; Li, Ang; Chaudhary, Vipin; Guan, Qiang

doi:10.1109/IPCCC55026.2022.9894315

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2005.01584 (cs)

[Submitted on 4 May 2020 (v1), last revised 23 Dec 2022 (this version, v3)]

Title:MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Authors:Betis Baheri, Jacob Tronge, Bo Fang, Ang Li, Vipin Chaudhary, Qiang Guan

View PDF

Abstract:In this paper, we introduce MARS, a new scheduling system for HPC-cloud infrastructures based on a cost-aware, flexible reinforcement learning approach, which serves as an intermediate layer for next generation HPC-cloud resource manager. MARS ensembles the pre-trained models from heuristic workloads and decides on the most cost-effective strategy for optimization. A whole workflow application would be split into several optimizable dependent sub-tasks, then based on the pre-defined resource management plan, a reward will be generated after executing a scheduled task. Lastly, MARS updates the Deep Neural Network (DNN) model based on the reward. MARS is designed to optimize the existing models through reinforcement mechanisms. MARS adapts to the dynamics of workflow applications, selects the most cost-effective scheduling solution among pre-built scheduling strategies (backfilling, SJF, etc.) and self-learning deep neural network model at run-time. We evaluate MARS with different real-world workflow traces. MARS can achieve 5%-60% increased performance compared to the state-of-the-art approaches.

Comments:	10 pages, HPC, Cloud System, Scheduling, Workflow Management, Reinforcement Learning, Deep Learning
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as:	arXiv:2005.01584 [cs.DC]
	(or arXiv:2005.01584v3 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2005.01584
Journal reference:	2022 IEEE International Performance Computing and Communications Conference (IPCCC) 217-226
Related DOI:	https://doi.org/10.1109/IPCCC55026.2022.9894315

Submission history

From: Betis Baheri [view email]
[v1] Mon, 4 May 2020 15:51:41 UTC (2,582 KB)
[v2] Mon, 22 Aug 2022 18:08:06 UTC (2,576 KB)
[v3] Fri, 23 Dec 2022 07:14:29 UTC (2,576 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:MARS: Malleable Actor-Critic Reinforcement Learning Scheduler

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators