-
A Bayesian joint longitudinal-survival model with a latent stochastic process for intensive longitudinal data
Authors:
Madeline R. Abbott,
Walter H. Dempsey,
Inbal Nahum-Shani,
Lindsey N. Potter,
David W. Wetter,
Cho Y. Lam,
Jeremy M. G. Taylor
Abstract:
The availability of mobile health (mHealth) technology has enabled increased collection of intensive longitudinal data (ILD). ILD have potential to capture rapid fluctuations in outcomes that may be associated with changes in the risk of an event. However, existing methods for jointly modeling longitudinal and event-time outcomes are not well-equipped to handle ILD due to the high computational co…
▽ More
The availability of mobile health (mHealth) technology has enabled increased collection of intensive longitudinal data (ILD). ILD have potential to capture rapid fluctuations in outcomes that may be associated with changes in the risk of an event. However, existing methods for jointly modeling longitudinal and event-time outcomes are not well-equipped to handle ILD due to the high computational cost. We propose a joint longitudinal and time-to-event model suitable for analyzing ILD. In this model, we summarize a multivariate longitudinal outcome as a smaller number of time-varying latent factors. These latent factors, which are modeled using an Ornstein-Uhlenbeck stochastic process, capture the risk of a time-to-event outcome in a parametric hazard model. We take a Bayesian approach to fit our joint model and conduct simulations to assess its performance. We use it to analyze data from an mHealth study of smoking cessation. We summarize the longitudinal self-reported intensity of nine emotions as the psychological states of positive and negative affect. These time-varying latent states capture the risk of the first smoking lapse after attempted quit. Understanding factors associated with smoking lapse is of keen interest to smoking cessation researchers.
△ Less
Submitted 30 April, 2024;
originally announced May 2024.
-
Optimizing Dynamic Predictions from Joint Models using Super Learning
Authors:
Dimitris Rizopoulos,
Jeremy M. G. Taylor
Abstract:
Joint models for longitudinal and time-to-event data are often employed to calculate dynamic individualized predictions used in numerous applications of precision medicine. Two components of joint models that influence the accuracy of these predictions are the shape of the longitudinal trajectories and the functional form linking the longitudinal outcome history to the hazard of the event. Finding…
▽ More
Joint models for longitudinal and time-to-event data are often employed to calculate dynamic individualized predictions used in numerous applications of precision medicine. Two components of joint models that influence the accuracy of these predictions are the shape of the longitudinal trajectories and the functional form linking the longitudinal outcome history to the hazard of the event. Finding a single well-specified model that produces accurate predictions for all subjects and follow-up times can be challenging, especially when considering multiple longitudinal outcomes. In this work, we use the concept of super learning and avoid selecting a single model. In particular, we specify a weighted combination of the dynamic predictions calculated from a library of joint models with different specifications. The weights are selected to optimize a predictive accuracy metric using V-fold cross-validation. We use as predictive accuracy measures the expected quadratic prediction error and the expected predictive cross-entropy. In a simulation study, we found that the super learning approach produces results very similar to the Oracle model, which was the model with the best performance in the test datasets. All proposed methodology is implemented in the freely available R package JMbayes2.
△ Less
Submitted 1 December, 2023; v1 submitted 20 September, 2023;
originally announced September 2023.
-
Using Joint Models for Longitudinal and Time-to-Event Data to Investigate the Causal Effect of Salvage Therapy after Prostatectomy
Authors:
Dimitris Rizopoulos,
Jeremy M. G. Taylor,
Grigorios Papageorgiou,
Todd M. Morgan
Abstract:
Prostate cancer patients who undergo prostatectomy are closely monitored for recurrence and metastasis using routine prostate-specific antigen (PSA) measurements. When PSA levels rise, salvage therapies are recommended to decrease the risk of metastasis. However, due to the side effects of these therapies and to avoid over-treatment, it is important to understand which patients and when to initiat…
▽ More
Prostate cancer patients who undergo prostatectomy are closely monitored for recurrence and metastasis using routine prostate-specific antigen (PSA) measurements. When PSA levels rise, salvage therapies are recommended to decrease the risk of metastasis. However, due to the side effects of these therapies and to avoid over-treatment, it is important to understand which patients and when to initiate these salvage therapies. In this work, we use the University of Michigan Prostatectomy registry Data to tackle this question. Due to the observational nature of this data, we face the challenge that PSA is simultaneously a time-varying confounder and an intermediate variable for salvage therapy. We define different causal salvage therapy effects defined conditionally on different specifications of the longitudinal PSA history. We then illustrate how these effects can be estimated using the framework of joint models for longitudinal and time-to-event data. All proposed methodology is implemented in the freely-available R package JMbayes2.
△ Less
Submitted 5 September, 2023;
originally announced September 2023.
-
A Continuous-Time Dynamic Factor Model for Intensive Longitudinal Data Arising from Mobile Health Studies
Authors:
Madeline R. Abbott,
Walter H. Dempsey,
Inbal Nahum-Shani,
Cho Y. Lam,
David W. Wetter,
Jeremy M. G. Taylor
Abstract:
Intensive longitudinal data (ILD) collected in mobile health (mHealth) studies contain rich information on multiple outcomes measured frequently over time that have the potential to capture short-term and long-term dynamics. Motivated by an mHealth study of smoking cessation in which participants self-report the intensity of many emotions multiple times per day, we describe a dynamic factor model…
▽ More
Intensive longitudinal data (ILD) collected in mobile health (mHealth) studies contain rich information on multiple outcomes measured frequently over time that have the potential to capture short-term and long-term dynamics. Motivated by an mHealth study of smoking cessation in which participants self-report the intensity of many emotions multiple times per day, we describe a dynamic factor model that summarizes the ILD as a low-dimensional, interpretable latent process. This model consists of two submodels: (i) a measurement submodel--a factor model--that summarizes the multivariate longitudinal outcome as lower-dimensional latent variables and (ii) a structural submodel--an Ornstein-Uhlenbeck (OU) stochastic process--that captures the temporal dynamics of the multivariate latent process in continuous time. We derive a closed-form likelihood for the marginal distribution of the outcome and the computationally-simpler sparse precision matrix for the OU process. We propose a block coordinate descent algorithm for estimation. Finally, we apply our method to the mHealth data to summarize the dynamics of 18 different emotions as two latent processes. These latent processes are interpreted by behavioral scientists as the psychological constructs of positive and negative affect and are key in understanding vulnerability to lapsing back to tobacco use among smokers attempting to quit.
△ Less
Submitted 20 February, 2024; v1 submitted 28 July, 2023;
originally announced July 2023.
-
Surrogacy Validation for Time-to-Event Outcomes with Illness-Death Frailty Models
Authors:
Emily K. Roberts,
Michael R. Elliott,
Jeremy M. G. Taylor
Abstract:
A common practice in clinical trials is to evaluate a treatment effect on an intermediate endpoint when the true outcome of interest would be difficult or costly to measure. We consider how to validate intermediate endpoints in a causally-valid way when the trial outcomes are time-to-event. Using counterfactual outcomes, those that would be observed if the counterfactual treatment had been given,…
▽ More
A common practice in clinical trials is to evaluate a treatment effect on an intermediate endpoint when the true outcome of interest would be difficult or costly to measure. We consider how to validate intermediate endpoints in a causally-valid way when the trial outcomes are time-to-event. Using counterfactual outcomes, those that would be observed if the counterfactual treatment had been given, the causal association paradigm assesses the relationship of the treatment effect on the surrogate $S$ with the treatment effect on the true endpoint $T$. In particular, we propose illness death models to accommodate the censored and semi-competing risk structure of survival data. The proposed causal version of these models involves estimable and counterfactual frailty terms. Via these multi-state models, we characterize what a valid surrogate would look like using a causal effect predictiveness plot. We evaluate the estimation properties of a Bayesian method using Markov Chain Monte Carlo and assess the sensitivity of our model assumptions. Our motivating data source is a localized prostate cancer clinical trial where the two survival endpoints are time to distant metastasis and time to death.
△ Less
Submitted 28 November, 2022;
originally announced November 2022.
-
Understanding the dynamic impact of COVID-19 through competing risk modeling with bivariate varying coefficients
Authors:
Wenbo Wu,
John D. Kalbfleisch,
Jeremy M. G. Taylor,
Jian Kang,
Kevin He
Abstract:
The coronavirus disease 2019 (COVID-19) pandemic has exerted a profound impact on patients with end-stage renal disease relying on kidney dialysis to sustain their lives. Motivated by a request by the U.S. Centers for Medicare & Medicaid Services, our analysis of their postdischarge hospital readmissions and deaths in 2020 revealed that the COVID-19 effect has varied significantly with postdischar…
▽ More
The coronavirus disease 2019 (COVID-19) pandemic has exerted a profound impact on patients with end-stage renal disease relying on kidney dialysis to sustain their lives. Motivated by a request by the U.S. Centers for Medicare & Medicaid Services, our analysis of their postdischarge hospital readmissions and deaths in 2020 revealed that the COVID-19 effect has varied significantly with postdischarge time and time since the onset of the pandemic. However, the complex dynamics of the COVID-19 effect trajectories cannot be characterized by existing varying coefficient models. To address this issue, we propose a bivariate varying coefficient model for competing risks within a cause-specific hazard framework, where tensor-product B-splines are used to estimate the surface of the COVID-19 effect. An efficient proximal Newton algorithm is developed to facilitate the fitting of the new model to the massive Medicare data for dialysis patients. Difference-based anisotropic penalization is introduced to mitigate model overfitting and the wiggliness of the estimated trajectories; various cross-validation methods are considered in the determination of optimal tuning parameters. Hypothesis testing procedures are designed to examine whether the COVID-19 effect varies significantly with postdischarge time and the time since pandemic onset, either jointly or separately. Simulation experiments are conducted to evaluate the estimation accuracy, type I error rate, statistical power, and model selection procedures. Applications to Medicare dialysis patients demonstrate the real-world performance of the proposed methods.
△ Less
Submitted 31 August, 2022;
originally announced September 2022.
-
A synthetic data integration framework to leverage external summary-level information from heterogeneous populations
Authors:
Tian Gu,
Jeremy M. G. Taylor,
Bhramar Mukherjee
Abstract:
There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in multiple forms, through regression coefficient estimates or predicted values of the outcome variable. Different external models may use different sets of predictors an…
▽ More
There is a growing need for flexible general frameworks that integrate individual-level data with external summary information for improved statistical inference. External information relevant for a risk prediction model may come in multiple forms, through regression coefficient estimates or predicted values of the outcome variable. Different external models may use different sets of predictors and the algorithm they used to predict the outcome Y given these predictors may or may not be known. The underlying populations corresponding to each external model may be different from each other and from the internal study population. Motivated by a prostate cancer risk prediction problem where novel biomarkers are measured only in the internal study, this paper proposes an imputation-based methodology where the goal is to fit a target regression model with all available predictors in the internal study while utilizing summary information from external models that may have used only a subset of the predictors. The method allows for heterogeneity of covariate effects across the external populations. The proposed approach generates synthetic outcome data in each external population, uses stacked multiple imputation technique to create a long dataset with complete covariate information. The final analysis of the stacked imputed data is conducted by weighted regression. This flexible and unified approach can improve statistical efficiency of the estimated coefficients in the internal study, improve predictions by utilizing even partial information available from models that use a subset of the full set of covariates used in the internal study, and provide statistical inference for the external population with potentially different covariate effects from the internal population.
△ Less
Submitted 1 June, 2022; v1 submitted 12 June, 2021;
originally announced June 2021.
-
Incorporating baseline covariates to validate surrogate endpoints with a constant biomarker under control arm
Authors:
Emily Roberts,
Michael Elliott,
Jeremy M. G. Taylor
Abstract:
A surrogate endpoint S in a clinical trial is an outcome that may be measured earlier or more easily than the true outcome of interest T. In this work, we extend causal inference approaches to validate such a surrogate using potential outcomes. The causal association paradigm assesses the relationship of the treatment effect on the surrogate with the treatment effect on the true endpoint. Using th…
▽ More
A surrogate endpoint S in a clinical trial is an outcome that may be measured earlier or more easily than the true outcome of interest T. In this work, we extend causal inference approaches to validate such a surrogate using potential outcomes. The causal association paradigm assesses the relationship of the treatment effect on the surrogate with the treatment effect on the true endpoint. Using the principal surrogacy criteria, we utilize the joint conditional distribution of the potential outcomes T, given the potential outcomes S. In particular, our setting of interest allows us to assume the surrogate under the placebo, S(0), is zero-valued, and we incorporate baseline covariates in the setting of normally-distributed endpoints. We develop Bayesian methods to incorporate conditional independence and other modeling assumptions and explore their impact on the assessment of surrogacy. We demonstrate our approach via simulation and data that mimics an ongoing study of a muscular dystrophy gene therapy.
△ Less
Submitted 2 February, 2022; v1 submitted 26 April, 2021;
originally announced April 2021.
-
Multiple imputation with missing data indicators
Authors:
Lauren J Beesley,
Irina Bondarenko,
Michael R Elliott,
Allison W Kurian,
Steven J Katz,
Jeremy M G Taylor
Abstract:
Multiple imputation is a well-established general technique for analyzing data with missing values. A convenient way to implement multiple imputation is sequential regression multiple imputation (SRMI), also called chained equations multiple imputation. In this approach, we impute missing values using regression models for each variable, conditional on the other variables in the data. This approac…
▽ More
Multiple imputation is a well-established general technique for analyzing data with missing values. A convenient way to implement multiple imputation is sequential regression multiple imputation (SRMI), also called chained equations multiple imputation. In this approach, we impute missing values using regression models for each variable, conditional on the other variables in the data. This approach, however, assumes that the missingness mechanism is missing at random, and it is not well-justified under not-at-random missingness without additional modification. In this paper, we describe how we can generalize the SRMI imputation procedure to handle not-at-random missingness (MNAR) in the setting where missingness may depend on other variables that are also missing. We provide algebraic justification for several generalizations of standard SRMI using Taylor series and other approximations of the target imputation distribution under MNAR. Resulting regression model approximations include indicators for missingness, interactions, or other functions of the MNAR missingness model and observed data. In a simulation study, we demonstrate that the proposed SRMI modifications result in reduced bias in the final analysis compared to standard SRMI, with an approximation strategy involving inclusion of an offset in the imputation model performing the best overall. The method is illustrated in a breast cancer study, where the goal is to estimate the prevalence of a specific genetic pathogenic variant.
△ Less
Submitted 2 March, 2021;
originally announced March 2021.
-
Accounting for not-at-random missingness through imputation stacking
Authors:
Lauren J Beesley,
Jeremy M G Taylor
Abstract:
Not-at-random missingness presents a challenge in addressing missing data in many health research applications. In this paper, we propose a new approach to account for not-at-random missingness after multiple imputation through weighted analysis of stacked multiple imputations. The weights are easily calculated as a function of the imputed data and assumptions about the not-at-random missingness.…
▽ More
Not-at-random missingness presents a challenge in addressing missing data in many health research applications. In this paper, we propose a new approach to account for not-at-random missingness after multiple imputation through weighted analysis of stacked multiple imputations. The weights are easily calculated as a function of the imputed data and assumptions about the not-at-random missingness. We demonstrate through simulation that the proposed method has excellent performance when the missingness model is correctly specified. In practice, the missingness mechanism will not be known. We show how we can use our approach in a sensitivity analysis framework to evaluate the robustness of model inference to different assumptions about the missingness mechanism, and we provide R package StackImpute to facilitate implementation as part of routine sensitivity analyses. We apply the proposed method to account for not-at-random missingness in human papillomavirus test results in a study of survival for patients diagnosed with oropharyngeal cancer.
△ Less
Submitted 19 January, 2021;
originally announced January 2021.
-
Kullback-Leibler-Based Discrete Failure Time Models for Integration of Published Prediction Models with New Time-To-Event Dataset
Authors:
Di Wang,
Wen Ye,
Randall Sung,
Hui Jiang,
Jeremy M. G. Taylor,
Lisa Ly,
Kevin He
Abstract:
Prediction of time-to-event data often suffers from rare event rates, small sample sizes, high dimensionality and low signal-to-noise ratios. Incorporating published prediction models from large-scale studies is expected to improve the performance of prognosis prediction on internal individual-level time-to-event data. However, existing integration approaches typically assume that underlying distr…
▽ More
Prediction of time-to-event data often suffers from rare event rates, small sample sizes, high dimensionality and low signal-to-noise ratios. Incorporating published prediction models from large-scale studies is expected to improve the performance of prognosis prediction on internal individual-level time-to-event data. However, existing integration approaches typically assume that underlying distributions from the external and internal data sources are similar, which is often invalid. To account for challenges including heterogeneity, data sharing, and privacy constraints, we propose a discrete failure time modeling procedure, which utilizes a discrete hazard-based Kullback-Leibler discriminatory information measuring the discrepancy between the published models and the internal dataset. Simulations show the advantage of the proposed method compared with those solely based on the internal data or published models. We apply the proposed method to improve prediction performance on a kidney transplant dataset from a local hospital by integrating this small-scale dataset with published survival models obtained from the national transplant registry.
△ Less
Submitted 28 July, 2022; v1 submitted 6 January, 2021;
originally announced January 2021.
-
A meta-inference framework to integrate multiple external models into a current study
Authors:
Tian Gu,
Jeremy M. G. Taylor,
Bhramar Mukherjee
Abstract:
It is becoming increasingly common for researchers to consider incorporating external information from large studies to improve the accuracy of statistical inference instead of relying on a modestly sized dataset collected internally. With some new predictors only available internally, we aim to build improved regression models based on individual-level data from an "internal" study while incorpor…
▽ More
It is becoming increasingly common for researchers to consider incorporating external information from large studies to improve the accuracy of statistical inference instead of relying on a modestly sized dataset collected internally. With some new predictors only available internally, we aim to build improved regression models based on individual-level data from an "internal" study while incorporating summary-level information from "external" models. We propose a meta-analysis framework along with two weighted estimators as the composite of empirical Bayes estimators, which combines the estimates from the different external models. The proposed framework is flexible and robust in the ways that (i) it is capable of incorporating external models that use a slightly different set of covariates; (ii) it can identify the most relevant external information and diminish the influence of information that is less compatible with the internal data; and (iii) it nicely balances the bias-variance trade-off while preserving the most efficiency gain. The proposed estimators are more efficient than the naive analysis of the internal data and other naive combinations of external estimators.
△ Less
Submitted 9 April, 2021; v1 submitted 19 October, 2020;
originally announced October 2020.
-
Using Multiple Imputation to Classify Potential Outcomes Subgroups
Authors:
Yun Li,
Irina Bondarenko,
Michael R. Elliott,
Timothy P. Hofer,
Jeremy M. G. Taylor
Abstract:
With medical tests becoming increasingly available, concerns about over-testing and over-treatment dramatically increase. Hence, it is important to understand the influence of testing on treatment selection in general practice. Most statistical methods focus on average effects of testing on treatment decisions. However, this may be ill-advised, particularly for patient subgroups that tend not to b…
▽ More
With medical tests becoming increasingly available, concerns about over-testing and over-treatment dramatically increase. Hence, it is important to understand the influence of testing on treatment selection in general practice. Most statistical methods focus on average effects of testing on treatment decisions. However, this may be ill-advised, particularly for patient subgroups that tend not to benefit from such tests. Furthermore, missing data are common, representing large and often unaddressed threats to the validity of statistical methods. Finally, it is desirable to conduct analyses that can be interpreted causally. We propose to classify patients into four potential outcomes subgroups, defined by whether or not a patient's treatment selection is changed by the test result and by the direction of how the test result changes treatment selection. This subgroup classification naturally captures the differential influence of medical testing on treatment selections for different patients, which can suggest targets to improve the utilization of medical tests. We can then examine patient characteristics associated with patient potential outcomes subgroup memberships. We used multiple imputation methods to simultaneously impute the missing potential outcomes as well as regular missing values. This approach can also provide estimates of many traditional causal quantities. We find that explicitly incorporating causal inference assumptions into the multiple imputation process can improve the precision for some causal estimates of interest. We also find that bias can occur when the potential outcomes conditional independence assumption is violated; sensitivity analyses are proposed to assess the impact of this violation. We applied the proposed methods to examine the influence of 21-gene assay, the most commonly used genomic test, on chemotherapy selection among breast cancer patients.
△ Less
Submitted 10 August, 2020;
originally announced August 2020.
-
A stacked approach for chained equations multiple imputation incorporating the substantive model
Authors:
Lauren Beesley,
Jeremy M G Taylor
Abstract:
Multiple imputation by chained equations (MICE) has emerged as a popular approach for handling missing data. A central challenge for applying MICE is determining how to incorporate outcome information into covariate imputation models, particularly for complicated outcomes. Often, we have a particular analysis model in mind, and we would like to ensure congeniality between the imputation and analys…
▽ More
Multiple imputation by chained equations (MICE) has emerged as a popular approach for handling missing data. A central challenge for applying MICE is determining how to incorporate outcome information into covariate imputation models, particularly for complicated outcomes. Often, we have a particular analysis model in mind, and we would like to ensure congeniality between the imputation and analysis models.
We propose a novel strategy for directly incorporating the analysis model into the handling of missing data. In our proposed approach, multiple imputations of missing covariates are obtained without using outcome information. We then utilize the strategy of imputation stacking, where multiple imputations are stacked on top of each other to create a large dataset. The analysis model is then incorporated through weights. Instead of applying multiple imputation combining rules, we obtain parameter estimates by fitting a weighted version of the analysis model on the stacked dataset. We propose a novel estimator for obtaining standard errors for this stacked and weighted analysis. Our estimator is based on the observed data information principle in Louis (1982) and can be applied for analyzing stacked multiple imputations more generally. Our approach for analyzing stacked multiple imputations is the first well-motivated method that can be easily applied for a wide variety of standard analysis models and missing data settings.
In simulations, the proposed strategy produced unbiased parameter estimates when the analysis model was correctly specified. We developed an R package, StackImpute, allowing this imputation approach to be easily implemented for many standard analysis models.
△ Less
Submitted 10 October, 2019;
originally announced October 2019.
-
Vertical modeling: analysis of competing risks data with a cure proportion
Authors:
M. A. Nicolaie,
J. M. G. Taylor,
C. Legrand
Abstract:
In this paper, we extend the vertical modeling approach for the analysis of survival data with competing risks to incorporate a cured fraction in the population, that is, a proportion of the population for which none of the competing events can occur. The proposed method has three components: the proportion of cure, the risk of failure, irrespective of the cause, and the relative risk of a certain…
▽ More
In this paper, we extend the vertical modeling approach for the analysis of survival data with competing risks to incorporate a cured fraction in the population, that is, a proportion of the population for which none of the competing events can occur. The proposed method has three components: the proportion of cure, the risk of failure, irrespective of the cause, and the relative risk of a certain cause of failure, given a failure occurred. Covariates may affect each of these components. An appealing aspect of the method is that it is a natural extension to competing risks of the semi-parametric mixture cure model in ordinary survival analysis; thus, causes of failure are assigned only if a failure occurs. This contrasts with the existing mixture cure model for competing risks of Larson and Dinse, which conditions at the onset on the future status presumably attained. Regression parameter estimates are obtained using an EM-algorithm. The performance of the estimators is evaluated in a simulation study. The method is illustrated using a melanoma cancer data set.
△ Less
Submitted 16 August, 2015;
originally announced August 2015.
-
Personalized Screening Intervals for Biomarkers using Joint Models for Longitudinal and Survival Data
Authors:
Dimitris Rizopoulos,
Jeremy M. G. Taylor,
Joost van Rosmalen,
Ewout W. Steyerberg,
Johanna J. M. Takkenberg
Abstract:
Screening and surveillance are routinely used in medicine for early detection of disease and close monitoring of progression. Biomarkers are one of the primarily tools used for these tasks, but their successful translation to clinical practice is closely linked to their ability to accurately predict clinical endpoints during follow-up. Motivated by a study of patients who received a human tissue v…
▽ More
Screening and surveillance are routinely used in medicine for early detection of disease and close monitoring of progression. Biomarkers are one of the primarily tools used for these tasks, but their successful translation to clinical practice is closely linked to their ability to accurately predict clinical endpoints during follow-up. Motivated by a study of patients who received a human tissue valve in the aortic position, in this work we are interested in optimizing and personalizing screening intervals for longitudinal biomarker measurements. Our aim in this paper is twofold: First, to appropriately select the model to use at time t, the time point the patient was still event-free, and second, based on this model to select the optimal time point u > t to plan the next measurement. To achieve these two goals we develop measures based on information theory quantities that assess the information we gain for the conditional survival process given the history of the subject that includes both baseline information and his/her accumulated longitudinal measurements.
△ Less
Submitted 22 March, 2015;
originally announced March 2015.
-
Bayesian shrinkage methods for partially observed data with many predictors
Authors:
Philip S. Boonstra,
Bhramar Mukherjee,
Jeremy M. G. Taylor
Abstract:
Motivated by the increasing use of and rapid changes in array technologies, we consider the prediction problem of fitting a linear regression relating a continuous outcome $Y$ to a large number of covariates $\mathbf {X}$, for example, measurements from current, state-of-the-art technology. For most of the samples, only the outcome $Y$ and surrogate covariates, $\mathbf {W}$, are available. These…
▽ More
Motivated by the increasing use of and rapid changes in array technologies, we consider the prediction problem of fitting a linear regression relating a continuous outcome $Y$ to a large number of covariates $\mathbf {X}$, for example, measurements from current, state-of-the-art technology. For most of the samples, only the outcome $Y$ and surrogate covariates, $\mathbf {W}$, are available. These surrogates may be data from prior studies using older technologies. Owing to the dimension of the problem and the large fraction of missing information, a critical issue is appropriate shrinkage of model parameters for an optimal bias-variance trade-off. We discuss a variety of fully Bayesian and Empirical Bayes algorithms which account for uncertainty in the missing data and adaptively shrink parameter estimates for superior prediction. These methods are evaluated via a comprehensive simulation study. In addition, we apply our methods to a lung cancer data set, predicting survival time ($Y$) using qRT-PCR ($\mathbf {X}$) and microarray ($\mathbf {W}$) measurements.
△ Less
Submitted 10 January, 2014;
originally announced January 2014.
-
Signal extraction and breakpoint identification for array CGH data using robust state space model
Authors:
Bin Zhu,
Jeremy M. G. Taylor,
Peter X. -K. Song
Abstract:
Array comparative genomic hybridization(CGH) is a high resolution technique to assess DNA copy number variation. Identifying breakpoints where copy number changes will enhance the understanding of the pathogenesis of human diseases, such as cancers. However, the biological variation and experimental errors contained in array CGH data may lead to false positive identification of breakpoints. We pro…
▽ More
Array comparative genomic hybridization(CGH) is a high resolution technique to assess DNA copy number variation. Identifying breakpoints where copy number changes will enhance the understanding of the pathogenesis of human diseases, such as cancers. However, the biological variation and experimental errors contained in array CGH data may lead to false positive identification of breakpoints. We propose a robust state space model for array CGH data analysis. The model consists of two equations: an observation equation and a state equation, in which both the measurement error and evolution error are specified to follow t-distributions with small degrees of freedom. The completely unspecified CGH profiles are estimated by a Markov Chain Monte Carlo(MCMC) algorithm. Breakpoints and outliers are identified by a novel backward selection procedure based on posterior draws of the CGH profiles. Compared to three other popular methods, our method demonstrates several desired features, including false positive rate control, robustness against outliers, and superior power of breakpoint detection. All these properties are illustrated using simulated and real datasets.
△ Less
Submitted 24 January, 2012;
originally announced January 2012.