Showing 1–2 of 2 results for author: Bas-Serrano, J

Search v0.5.6 released 2020-02-24

arXiv:2010.11151 [pdf, other]

cs.LG cs.AI stat.ML

Logistic Q-Learning

Authors: Joan Bas-Serrano, Sebastian Curi, Andreas Krause, Gergely Neu

Abstract: We propose a new reinforcement learning algorithm derived from a regularized linear-programming formulation of optimal control in MDPs. The method is closely related to the classic Relative Entropy Policy Search (REPS) algorithm of Peters et al. (2010), with the key difference that our method introduces a Q-function that enables efficient exact model-free implementation. The main feature of our al… ▽ More We propose a new reinforcement learning algorithm derived from a regularized linear-programming formulation of optimal control in MDPs. The method is closely related to the classic Relative Entropy Policy Search (REPS) algorithm of Peters et al. (2010), with the key difference that our method introduces a Q-function that enables efficient exact model-free implementation. The main feature of our algorithm (called QREPS) is a convex loss function for policy evaluation that serves as a theoretically sound alternative to the widely used squared Bellman error. We provide a practical saddle-point optimization method for minimizing this loss function and provide an error-propagation analysis that relates the quality of the individual updates to the performance of the output policy. Finally, we demonstrate the effectiveness of our method on a range of benchmark problems. △ Less

Submitted 25 February, 2021; v1 submitted 21 October, 2020; originally announced October 2020.
arXiv:1909.10904 [pdf, ps, other]

math.OC cs.LG stat.ML

Faster saddle-point optimization for solving large-scale Markov decision processes

Authors: Joan Bas-Serrano, Gergely Neu

Abstract: We consider the problem of computing optimal policies in average-reward Markov decision processes. This classical problem can be formulated as a linear program directly amenable to saddle-point optimization methods, albeit with a number of variables that is linear in the number of states. To address this issue, recent work has considered a linearly relaxed version of the resulting saddle-point pro… ▽ More We consider the problem of computing optimal policies in average-reward Markov decision processes. This classical problem can be formulated as a linear program directly amenable to saddle-point optimization methods, albeit with a number of variables that is linear in the number of states. To address this issue, recent work has considered a linearly relaxed version of the resulting saddle-point problem. Our work aims at achieving a better understanding of this relaxed optimization problem by characterizing the conditions necessary for convergence to the optimal policy, and designing an optimization algorithm enjoying fast convergence rates that are independent of the size of the state space. Notably, our characterization points out some potential issues with previous work. △ Less

Submitted 10 January, 2020; v1 submitted 22 September, 2019; originally announced September 2019.

Search v0.5.6 released 2020-02-24