From naive trees to Random Forests: A general approach for proving consistency of tree-based methods

Föge, Nico; Pauly, Markus; Schmid, Lena; Ditzhaus, Marc

Mathematics > Statistics Theory

arXiv:2404.06850 (math)

This paper has been withdrawn by Nico Föge

[Submitted on 10 Apr 2024 (v1), last revised 22 Apr 2024 (this version, v3)]

Title:From naive trees to Random Forests: A general approach for proving consistency of tree-based methods

Authors:Nico Föge, Markus Pauly, Lena Schmid, Marc Ditzhaus

No PDF available, click to view other formats

Abstract:Tree-based methods such as Random Forests are learning algorithms that have become an integral part of the statistical toolbox. The last decade has shed some light on theoretical properties such as their consistency for regression tasks. However, the usual proofs assume normal error terms as well as an additive regression function and are rather technical. We overcome these issues by introducing a simple and catchy technique for proving consistency under quite general assumptions. To this end, we introduce a new class of naive trees, which do the subspacing completely at random and independent of the data. We then give a direct proof of their consistency. Using them to bound the error of more complex tree-based approaches such as univariate and multivariate CARTs, Extra Randomized Trees, or Random Forests, we deduce the consistency of all of them. Since naive trees appear to be too simple for actual application, we further analyze their finite sample properties in a simulation and small benchmark study. We find a slow convergence speed and a rather poor predictive performance. Based on these results, we finally discuss to what extent consistency proofs help to justify the application of complex learning algorithms.

Comments:	Incorrect Proof
Subjects:	Statistics Theory (math.ST)
MSC classes:	Primary 62G05, secondary 62G20
Cite as:	arXiv:2404.06850 [math.ST]
	(or arXiv:2404.06850v3 [math.ST] for this version)
	https://doi.org/10.48550/arXiv.2404.06850

Submission history

From: Nico Föge [view email]
[v1] Wed, 10 Apr 2024 09:20:06 UTC (724 KB)
[v2] Thu, 11 Apr 2024 10:30:59 UTC (724 KB)
[v3] Mon, 22 Apr 2024 08:46:02 UTC (1 KB) (withdrawn)

Mathematics > Statistics Theory

Title:From naive trees to Random Forests: A general approach for proving consistency of tree-based methods

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Mathematics > Statistics Theory

Title:From naive trees to Random Forests: A general approach for proving consistency of tree-based methods

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators