Offline Estimation of Controlled Markov Chains: Minimax Nonparametric Estimators and Sample Efficiency

Banerjee, Imon; Honnappa, Harsha; Rao, Vinayak

Statistics > Machine Learning

arXiv:2211.07092v1 (stat)

[Submitted on 14 Nov 2022 (this version), latest version 26 Jan 2024 (v4)]

Title:Offline Estimation of Controlled Markov Chains: Minimax Nonparametric Estimators and Sample Efficiency

Authors:Imon Banerjee, Harsha Honnappa, Vinayak Rao

View PDF

Abstract:Controlled Markov chains (CMCs) have wide applications in engineering and machine learning, forming a key component in many reinforcement learning problems. In this work, we consider the estimation of the transition probabilities of a finite-state finite-control CMC, and develop a minimax sample complexity bounds for nonparametric estimation of these transition probability matrices. Unlike most studies that have been done in the online setup, we consider offline MDPs. Our results are quite general, since we do not assume anything specific about the logging policy. Instead, the dependence of our statistical bounds on the logging policy comes in the form of a natural mixing coefficient. We demonstrate an interesting trade-off between stronger assumptions on mixing versus requiring more samples to achieve a particular PAC-bound. We demonstrate the validity of our results under various examples, like ergodic Markov chains, weakly ergodic inhomogenous Markov chains, and controlled Markov chains with non-stationary Markov, episodic, and greedy controls. Lastly, we use the properties of the estimated transition matrix to perform estimate the value function when the controls are stationary and Markov.

Comments:	67 pages, 23 main-47 appendix
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as:	arXiv:2211.07092 [stat.ML]
	(or arXiv:2211.07092v1 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2211.07092

Submission history

From: Imon Banerjee [view email]
[v1] Mon, 14 Nov 2022 03:39:59 UTC (83 KB)
[v2] Tue, 15 Nov 2022 16:26:30 UTC (83 KB)
[v3] Wed, 1 Feb 2023 13:31:47 UTC (81 KB)
[v4] Fri, 26 Jan 2024 20:23:18 UTC (812 KB)

Statistics > Machine Learning

Title:Offline Estimation of Controlled Markov Chains: Minimax Nonparametric Estimators and Sample Efficiency

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:Offline Estimation of Controlled Markov Chains: Minimax Nonparametric Estimators and Sample Efficiency

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators