-
Reinforcement Learning in Non-Markovian Environments
Abstract: Motivated by the novel paradigm developed by Van Roy and coauthors for reinforcement learning in arbitrary non-Markovian environments, we propose a related formulation and explicitly pin down the error caused by non-Markovianity of observations when the Q-learning algorithm is applied on this formulation. Based on this observation, we propose that the criterion for agent design should be to seek g… ▽ More
Submitted 13 February, 2024; v1 submitted 3 November, 2022; originally announced November 2022.
Comments: 19 pages, accepted for publication at Systems and Control Letters
-
Concentration of Contractive Stochastic Approximation and Reinforcement Learning
Abstract: Using a martingale concentration inequality, concentration bounds `from time $n_0$ on' are derived for stochastic approximation algorithms with contractive maps and both martingale difference and Markov noises. These are applied to reinforcement learning algorithms, in particular to asynchronous Q-learning and TD(0).
Submitted 11 June, 2022; v1 submitted 27 June, 2021; originally announced June 2021.
Comments: 20 pages, Accepted for publication in Stochastic Systems