Skip to main content

Showing 1–2 of 2 results for author: Bas-Serrano, J

.
  1. arXiv:2010.11151  [pdf, other

    cs.LG cs.AI stat.ML

    Logistic Q-Learning

    Authors: Joan Bas-Serrano, Sebastian Curi, Andreas Krause, Gergely Neu

    Abstract: We propose a new reinforcement learning algorithm derived from a regularized linear-programming formulation of optimal control in MDPs. The method is closely related to the classic Relative Entropy Policy Search (REPS) algorithm of Peters et al. (2010), with the key difference that our method introduces a Q-function that enables efficient exact model-free implementation. The main feature of our al… ▽ More

    Submitted 25 February, 2021; v1 submitted 21 October, 2020; originally announced October 2020.

  2. arXiv:1909.10904  [pdf, ps, other

    math.OC cs.LG stat.ML

    Faster saddle-point optimization for solving large-scale Markov decision processes

    Authors: Joan Bas-Serrano, Gergely Neu

    Abstract: We consider the problem of computing optimal policies in average-reward Markov decision processes. This classical problem can be formulated as a linear program directly amenable to saddle-point optimization methods, albeit with a number of variables that is linear in the number of states. To address this issue, recent work has considered a linearly relaxed version of the resulting saddle-point pro… ▽ More

    Submitted 10 January, 2020; v1 submitted 22 September, 2019; originally announced September 2019.