FACT: High-Dimensional Random Forests Inference

Chi, Chien-Ming; Fan, Yingying; Lv, **chi

Statistics > Machine Learning

arXiv:2207.01678 (stat)

[Submitted on 4 Jul 2022 (v1), last revised 13 Nov 2023 (this version, v2)]

Title:FACT: High-Dimensional Random Forests Inference

Authors:Chien-Ming Chi, Yingying Fan, **chi Lv

View PDF

Abstract:Quantifying the usefulness of individual features in random forests learning can greatly enhance its interpretability. Existing studies have shown that some popularly used feature importance measures for random forests suffer from the bias issue. In addition, there lack comprehensive size and power analyses for most of these existing methods. In this paper, we approach the problem via hypothesis testing, and suggest a framework of the self-normalized feature-residual correlation test (FACT) for evaluating the significance of a given feature in the random forests model with bias-resistance property, where our null hypothesis concerns whether the feature is conditionally independent of the response given all other features. Such an endeavor on random forests inference is empowered by some recent developments on high-dimensional random forests consistency. Under a fairly general high-dimensional nonparametric model setting with dependent features, we formally establish that FACT can provide theoretically justified feature importance test with controlled type I error and enjoy appealing power property. The theoretical results and finite-sample advantages of the newly suggested method are illustrated with several simulation examples and an economic forecasting application.

Comments:	42 pages, 3 figures
Subjects:	Machine Learning (stat.ML); Machine Learning (cs.LG); Statistics Theory (math.ST)
Cite as:	arXiv:2207.01678 [stat.ML]
	(or arXiv:2207.01678v2 [stat.ML] for this version)
	https://doi.org/10.48550/arXiv.2207.01678

Submission history

From: **chi Lv [view email]
[v1] Mon, 4 Jul 2022 19:05:08 UTC (404 KB)
[v2] Mon, 13 Nov 2023 04:08:57 UTC (2,057 KB)

Statistics > Machine Learning

Title:FACT: High-Dimensional Random Forests Inference

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Statistics > Machine Learning

Title:FACT: High-Dimensional Random Forests Inference

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators