Decentralised Q-Learning for Multi-Agent Markov Decision Processes with a Satisfiability Criterion

Keval, Keshav P.; Borkar, Vivek S.

Electrical Engineering and Systems Science > Systems and Control

arXiv:2311.12613 (eess)

[Submitted on 21 Nov 2023]

Title:Decentralised Q-Learning for Multi-Agent Markov Decision Processes with a Satisfiability Criterion

Authors:Keshav P. Keval, Vivek S. Borkar

View PDF

Abstract:In this paper, we propose a reinforcement learning algorithm to solve a multi-agent Markov decision process (MMDP). The goal, inspired by Blackwell's Approachability Theorem, is to lower the time average cost of each agent to below a pre-specified agent-specific bound. For the MMDP, we assume the state dynamics to be controlled by the joint actions of agents, but the per-stage costs to only depend on the individual agent's actions. We combine the Q-learning algorithm for a weighted combination of the costs of each agent, obtained by a gossip algorithm with the Metropolis-Hastings or Multiplicative Weights formalisms to modulate the averaging matrix of the gossip. We use multiple timescales in our algorithm and prove that under mild conditions, it approximately achieves the desired bounds for each of the agents. We also demonstrate the empirical performance of this algorithm in the more general setting of MMDPs having jointly controlled per-stage costs.

Subjects:	Systems and Control (eess.SY); Machine Learning (cs.LG); Multiagent Systems (cs.MA)
Cite as:	arXiv:2311.12613 [eess.SY]
	(or arXiv:2311.12613v1 [eess.SY] for this version)
	https://doi.org/10.48550/arXiv.2311.12613

Submission history

From: Keshav Patel Keval [view email]
[v1] Tue, 21 Nov 2023 13:56:44 UTC (1,391 KB)

Electrical Engineering and Systems Science > Systems and Control

Title:Decentralised Q-Learning for Multi-Agent Markov Decision Processes with a Satisfiability Criterion

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Systems and Control

Title:Decentralised Q-Learning for Multi-Agent Markov Decision Processes with a Satisfiability Criterion

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators