-
Gaussian processes Correlated Bayesian Additive Regression Trees
Authors:
Xuetao Lu a,
Robert E. McCulloch
Abstract:
In recent years, Bayesian Additive Regression Trees (BART) has garnered increased attention, leading to the development of various extensions for diverse applications. However, there has been limited exploration of its utility in analyzing correlated data. This paper introduces a novel extension of BART, named Correlated BART (CBART). Unlike the original BART with independent errors, CBART is spec…
▽ More
In recent years, Bayesian Additive Regression Trees (BART) has garnered increased attention, leading to the development of various extensions for diverse applications. However, there has been limited exploration of its utility in analyzing correlated data. This paper introduces a novel extension of BART, named Correlated BART (CBART). Unlike the original BART with independent errors, CBART is specifically designed to handle correlated (dependent) errors. Additionally, we propose the integration of CBART with Gaussian processes (GP) to create a new model termed GP-CBART. This innovative model combines the strengths of the Gaussian processes and CBART, making it particularly well-suited for analyzing time series or spatial data. In the GP-CBART framework, CBART captures the nonlinearity in the mean regression (covariates) function, while the Gaussian processes adeptly models the correlation structure within the response. Additionally, given the high flexibility of both CBART and GP models, their combination may lead to identification issues. We provide methods to address these challenges. To demonstrate the effectiveness of CBART and GP-CBART, we present corresponding simulated and real-world examples.
△ Less
Submitted 30 November, 2023;
originally announced November 2023.
-
Influential Observations in Bayesian Regression Tree Models
Authors:
Matthew T. Pratola,
Edward I. George,
Robert E. McCulloch
Abstract:
BCART (Bayesian Classification and Regression Trees) and BART (Bayesian Additive Regression Trees) are popular Bayesian regression models widely applicable in modern regression problems. Their popularity is intimately tied to the ability to flexibly model complex responses depending on high-dimensional inputs while simultaneously being able to quantify uncertainties. This ability to quantify uncer…
▽ More
BCART (Bayesian Classification and Regression Trees) and BART (Bayesian Additive Regression Trees) are popular Bayesian regression models widely applicable in modern regression problems. Their popularity is intimately tied to the ability to flexibly model complex responses depending on high-dimensional inputs while simultaneously being able to quantify uncertainties. This ability to quantify uncertainties is key, as it allows researchers to perform appropriate inferential analyses in settings that have generally been too difficult to handle using the Bayesian approach. However, surprisingly little work has been done to evaluate the sensitivity of these modern regression models to violations of modeling assumptions. In particular, we will consider influential observations, which one reasonably would imagine to be common -- or at least a concern -- in the big-data setting. In this paper, we consider both the problem of detecting influential observations and adjusting predictions to not be unduly affected by such potentially problematic data. We consider three detection diagnostics for Bayesian tree models, one an analogue of Cook's distance and the others taking the form of a divergence measure and a conditional predictive density metric, and then propose an importance sampling algorithm to re-weight previously sampled posterior draws so as to remove the effects of influential data in a computationally efficient manner. Finally, our methods are demonstrated on real-world data where blind application of the models can lead to poor predictions and inference.
△ Less
Submitted 17 May, 2023; v1 submitted 26 March, 2022;
originally announced March 2022.
-
Causal Inference with the Instrumental Variable Approach and Bayesian Nonparametric Machine Learning
Authors:
Robert E. McCulloch,
Rodney A. Sparapani,
Brent R. Logan,
Purushottam W. Laud
Abstract:
We provide a new flexible framework for inference with the instrumental variable model. Rather than using linear specifications, functions characterizing the effects of instruments and other explanatory variables are estimated using machine learning via Bayesian Additive Regression Trees (BART). Error terms and their distribution are inferred using Dirichlet Process mixtures. Simulated and real ex…
▽ More
We provide a new flexible framework for inference with the instrumental variable model. Rather than using linear specifications, functions characterizing the effects of instruments and other explanatory variables are estimated using machine learning via Bayesian Additive Regression Trees (BART). Error terms and their distribution are inferred using Dirichlet Process mixtures. Simulated and real examples show that when the true functions are linear, little is lost. But when nonlinearities are present, dramatic improvements are obtained with virtually no manual tuning.
△ Less
Submitted 1 February, 2021;
originally announced February 2021.
-
Nonparametric competing risks analysis using Bayesian Additive Regression Trees (BART)
Authors:
Rodney Sparapani,
Brent R. Logan,
Robert E. McCulloch,
Purushottam W. Laud
Abstract:
Many time-to-event studies are complicated by the presence of competing risks. Such data are often analyzed using Cox models for the cause specific hazard function or Fine-Gray models for the subdistribution hazard. In practice regression relationships in competing risks data with either strategy are often complex and may include nonlinear functions of covariates, interactions, high-dimensional pa…
▽ More
Many time-to-event studies are complicated by the presence of competing risks. Such data are often analyzed using Cox models for the cause specific hazard function or Fine-Gray models for the subdistribution hazard. In practice regression relationships in competing risks data with either strategy are often complex and may include nonlinear functions of covariates, interactions, high-dimensional parameter spaces and nonproportional cause specific or subdistribution hazards. Model misspecification can lead to poor predictive performance. To address these issues, we propose a novel approach to flexible prediction modeling of competing risks data using Bayesian Additive Regression Trees (BART). We study the simulation performance in two-sample scenarios as well as a complex regression setting, and benchmark its performance against standard regression techniques as well as random survival forests. We illustrate the use of the proposed method on a recently published study of patients undergoing hematopoietic stem cell transplantation.
△ Less
Submitted 28 June, 2018;
originally announced June 2018.
-
Decision making and uncertainty quantification for individualized treatments
Authors:
Brent R. Logan,
Rodney Sparapani,
Robert E. McCulloch,
Purushottam W. Laud
Abstract:
Individualized treatment rules (ITR) can improve health outcomes by recognizing that patients may respond differently to treatment and assigning therapy with the most desirable predicted outcome for each individual. Flexible and efficient prediction models are desired as a basis for such ITRs to handle potentially complex interactions between patient factors and treatment. Modern Bayesian semipara…
▽ More
Individualized treatment rules (ITR) can improve health outcomes by recognizing that patients may respond differently to treatment and assigning therapy with the most desirable predicted outcome for each individual. Flexible and efficient prediction models are desired as a basis for such ITRs to handle potentially complex interactions between patient factors and treatment. Modern Bayesian semiparametric and nonparametric regression models provide an attractive avenue in this regard as these allow natural posterior uncertainty quantification of patient specific treatment decisions as well as the population wide value of the prediction-based ITR. In addition, via the use of such models, inference is also available for the value of the Optimal ITR. We propose such an approach and implement it using Bayesian Additive Regression Trees (BART) as this model has been shown to perform well in fitting nonparametric regression functions to continuous and binary responses, even with many covariates. It is also computationally efficient for use in practice. With BART we investigate a treatment strategy which utilizes individualized predictions of patient outcomes from BART models. Posterior distributions of patient outcomes under each treatment are used to assign the treatment that maximizes the expected posterior utility. We also describe how to approximate such a treatment policy with a clinically interpretable ITR, and quantify its expected outcome. The proposed method performs very well in extensive simulation studies in comparison with several existing methods. We illustrate the usage of the proposed method to identify an individualized choice of conditioning regimen for patients undergoing hematopoietic cell transplantation and quantify the value of this method of choice in relation to the Optimal ITR as well as non-individualized treatment strategies.
△ Less
Submitted 21 September, 2017;
originally announced September 2017.
-
mBART: Multidimensional Monotone BART
Authors:
Hugh A. Chipman,
Edward I. George,
Robert E. McCulloch,
Thomas S. Shively
Abstract:
For the discovery of regression relationships between Y and a large set of p potential predictors x 1 , . . . , x p , the flexible nonparametric nature of BART (Bayesian Additive Regression Trees) allows for a much richer set of possibilities than restrictive parametric approaches. However, subject matter considerations sometimes warrant a minimal assumption of monotonicity in at least some of the…
▽ More
For the discovery of regression relationships between Y and a large set of p potential predictors x 1 , . . . , x p , the flexible nonparametric nature of BART (Bayesian Additive Regression Trees) allows for a much richer set of possibilities than restrictive parametric approaches. However, subject matter considerations sometimes warrant a minimal assumption of monotonicity in at least some of the predictors. For such contexts, we introduce mBART, a constrained version of BART that can flexibly incorporate monotonicity in any predesignated subset of predictors using a multivariate basis of monotone trees, while avoiding the further confines of a full parametric form. For such monotone relationships, mBART provides (i) function estimates that are smoother and more interpretable, (ii) better out-of-sample predictive performance, and (iii) less post-data uncertainty. While many key aspects of the unconstrained BART model carry over directly to mBART, the introduction of monotonicity constraints necessitates a fundamental rethinking of how the model is implemented. In particular, the original BART Markov Chain Monte Carlo algorithm relied on a conditional conjugacy that is no longer available in a monotonically constrained space. Various simulated and real examples demonstrate the wide ranging potential of mBART.
△ Less
Submitted 8 October, 2021; v1 submitted 5 December, 2016;
originally announced December 2016.
-
BART: Bayesian additive regression trees
Authors:
Hugh A. Chipman,
Edward I. George,
Robert E. McCulloch
Abstract:
We develop a Bayesian "sum-of-trees" model where each tree is constrained by a regularization prior to be a weak learner, and fitting and inference are accomplished via an iterative Bayesian backfitting MCMC algorithm that generates samples from a posterior. Effectively, BART is a nonparametric Bayesian regression approach which uses dimensionally adaptive random basis elements. Motivated by ensem…
▽ More
We develop a Bayesian "sum-of-trees" model where each tree is constrained by a regularization prior to be a weak learner, and fitting and inference are accomplished via an iterative Bayesian backfitting MCMC algorithm that generates samples from a posterior. Effectively, BART is a nonparametric Bayesian regression approach which uses dimensionally adaptive random basis elements. Motivated by ensemble methods in general, and boosting algorithms in particular, BART is defined by a statistical model: a prior and a likelihood. This approach enables full posterior inference including point and interval estimates of the unknown regression function as well as the marginal effects of potential predictors. By kee** track of predictor inclusion frequencies, BART can also be used for model-free variable selection. BART's many features are illustrated with a bake-off against competing methods on 42 different data sets, with a simulation experiment and on a drug discovery classification problem.
△ Less
Submitted 7 October, 2010; v1 submitted 19 June, 2008;
originally announced June 2008.