SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Mei, Zhiyu; Fu, Wei; Gao, Jiaxuan; Wang, Guangju; Zhang, Huanchen; Wu, Yi

Computer Science > Distributed, Parallel, and Cluster Computing

arXiv:2306.16688 (cs)

[Submitted on 29 Jun 2023 (v1), last revised 21 Jun 2024 (this version, v3)]

Title:SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Authors:Zhiyu Mei, Wei Fu, Jiaxuan Gao, Guangju Wang, Huanchen Zhang, Yi Wu

View PDF HTML (experimental)

Abstract:The ever-growing complexity of reinforcement learning (RL) tasks demands a distributed system to efficiently generate and process a massive amount of data. However, existing open-source libraries suffer from various limitations, which impede their practical use in challenging scenarios where large-scale training is necessary. In this paper, we present a novel abstraction on the dataflows of RL training, which unifies diverse RL training applications into a general framework. Following this abstraction, we develop a scalable, efficient, and extensible distributed RL system called ReaLlyScalableRL, which allows efficient and massively parallelized training and easy development of customized algorithms. Our evaluation shows that SRL outperforms existing academic libraries, reaching at most 21x higher training throughput in a distributed setting. On learning performance, beyond performing and scaling well on common RL benchmarks with different RL algorithms, SRL can reproduce the same solution in the challenging hide-and-seek environment as reported by OpenAI with up to 5x speedup in wall-clock time. Notably, SRL is the first in the academic community to perform RL experiments at a large scale with over 15k CPU cores. SRL source code is available at: this https URL .

Comments:	Published at ICLR 2024. 10 pages (24 pages with references and appendix), 7 figures
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as:	arXiv:2306.16688 [cs.DC]
	(or arXiv:2306.16688v3 [cs.DC] for this version)
	https://doi.org/10.48550/arXiv.2306.16688

Submission history

From: Zhiyu Mei [view email]
[v1] Thu, 29 Jun 2023 05:16:25 UTC (1,180 KB)
[v2] Wed, 5 Jul 2023 08:16:29 UTC (1,180 KB)
[v3] Fri, 21 Jun 2024 08:02:57 UTC (3,464 KB)

Computer Science > Distributed, Parallel, and Cluster Computing

Title:SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Distributed, Parallel, and Cluster Computing

Title:SRL: Scaling Distributed Reinforcement Learning to Over Ten Thousand Cores

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators