-
High-dimensional multiple imputation (HDMI) for partially observed confounders including natural language processing-derived auxiliary covariates
Authors:
Janick Weberpals,
Pamela A. Shaw,
Kueiyu Joshua Lin,
Richard Wyss,
Joseph M Plasek,
Li Zhou,
Kerry Ngan,
Thomas DeRamus,
Sudha R. Raman,
Bradley G. Hammill,
Hana Lee,
Sengwee Toh,
John G. Connolly,
Kimberly J. Dandreo,
Fang Tian,
Wei Liu,
Jie Li,
José J. Hernández-Muñoz,
Sebastian Schneeweiss,
Rishi J. Desai
Abstract:
Multiple imputation (MI) models can be improved by including auxiliary covariates (AC), but their performance in high-dimensional data is not well understood. We aimed to develop and compare high-dimensional MI (HDMI) approaches using structured and natural language processing (NLP)-derived AC in studies with partially observed confounders. We conducted a plasmode simulation study using data from…
▽ More
Multiple imputation (MI) models can be improved by including auxiliary covariates (AC), but their performance in high-dimensional data is not well understood. We aimed to develop and compare high-dimensional MI (HDMI) approaches using structured and natural language processing (NLP)-derived AC in studies with partially observed confounders. We conducted a plasmode simulation study using data from opioid vs. non-steroidal anti-inflammatory drug (NSAID) initiators (X) with observed serum creatinine labs (Z2) and time-to-acute kidney injury as outcome. We simulated 100 cohorts with a null treatment effect, including X, Z2, atrial fibrillation (U), and 13 other investigator-derived confounders (Z1) in the outcome generation. We then imposed missingness (MZ2) on 50% of Z2 measurements as a function of Z2 and U and created different HDMI candidate AC using structured and NLP-derived features. We mimicked scenarios where U was unobserved by omitting it from all AC candidate sets. Using LASSO, we data-adaptively selected HDMI covariates associated with Z2 and MZ2 for MI, and with U to include in propensity score models. The treatment effect was estimated following propensity score matching in MI datasets and we benchmarked HDMI approaches against a baseline imputation and complete case analysis with Z1 only. HDMI using claims data showed the lowest bias (0.072). Combining claims and sentence embeddings led to an improvement in the efficiency displaying the lowest root-mean-squared-error (0.173) and coverage (94%). NLP-derived AC alone did not perform better than baseline MI. HDMI approaches may decrease bias in studies with partially observed confounders where missingness depends on unobserved factors.
△ Less
Submitted 17 May, 2024;
originally announced May 2024.
-
Analysis of the 24-Hour Activity Cycle: An illustration examining the association with cognitive function in the Adult Changes in Thought (ACT) Study
Authors:
Yinxiang Wu,
Dori E. Rosenberg,
Mikael Anne Greenwood-Hickman,
Susan M. McCurry,
Cecile Proust-Lima,
Jennifer C. Nelson,
Paul K. Crane,
Andrea Z. LaCroix,
Eric B. Larson,
Pamela A. Shaw
Abstract:
The 24-hour activity cycle (24HAC) is a new paradigm for studying activity behaviors in relation to health outcomes. This approach captures the interrelatedness of the daily time spent in physical activity (PA), sedentary behavior (SB), and sleep. We illustrate and compare the use of three popular approaches, namely isotemporal substitution model (ISM), compositional data analysis (CoDA), and late…
▽ More
The 24-hour activity cycle (24HAC) is a new paradigm for studying activity behaviors in relation to health outcomes. This approach captures the interrelatedness of the daily time spent in physical activity (PA), sedentary behavior (SB), and sleep. We illustrate and compare the use of three popular approaches, namely isotemporal substitution model (ISM), compositional data analysis (CoDA), and latent profile analysis (LPA) for modeling outcome associations with the 24HAC. We apply these approaches to assess an association with a cognitive outcome, measured by CASI item response theory (IRT) score, in a cohort of 1034 older adults (mean [range] age = 77 [65-100]; 55.8% female; 90% White) who were part of the Adult Changes in Thought (ACT) Activity Monitoring (ACT-AM) sub-study. PA and SB were assessed with thigh-worn activPAL accelerometers for 7 days. We highlight differences in assumptions between the three approaches, discuss statistical challenges, and provide guidance on interpretation and selecting an appropriate approach. ISM is easiest to apply and interpret; however, the typical ISM model assumes a linear association. CoDA specifies a non-linear association through isometric logratio transformations that are more challenging to apply and interpret. LPA can classify individuals into groups with similar time-use patterns. Inference on associations of latent profiles with health outcomes need to account for the uncertainty of the LPA classifications which is often ignored. The selection of the most appropriate method should be guided by the scientific questions of interest and the applicability of each model's assumptions. The analytic results did not suggest that less time spent on SB and more in PA was associated with better cognitive function. Further research is needed into the health implications of the distinct 24HAC patterns identified in this cohort.
△ Less
Submitted 19 January, 2023;
originally announced January 2023.
-
Issues in Implementing Regression Calibration Analyses
Authors:
Lillian Boe,
Pamela A. Shaw,
Douglas Midthune,
Paul Gustafson,
Victor Kipnis,
Eunyoung Park,
Daniela Sotres-Alvarez,
Laurence Freedman
Abstract:
Regression calibration is a popular approach for correcting biases in estimated regression parameters when exposure variables are measured with error. This approach involves building a calibration equation to estimate the value of the unknown true exposure given the error-prone measurement and other confounding covariates. The estimated, or calibrated, exposure is then substituted for the true exp…
▽ More
Regression calibration is a popular approach for correcting biases in estimated regression parameters when exposure variables are measured with error. This approach involves building a calibration equation to estimate the value of the unknown true exposure given the error-prone measurement and other confounding covariates. The estimated, or calibrated, exposure is then substituted for the true exposure in the health outcome regression model. When used properly, regression calibration can greatly reduce the bias induced by exposure measurement error. Here, we first provide an overview of the statistical framework for regression calibration, specifically discussing how a special type of error, called Berkson error, arises in the estimated exposure. We then present practical issues to consider when applying regression calibration, including: (1) how to develop the calibration equation and which covariates to include; (2) valid ways to calculate standard errors (SE) of estimated regression coefficients; and (3) problems arising if one of the covariates in the calibration model is a mediator of the relationship between the exposure and outcome. Throughout the paper, we provide illustrative examples using data from the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) and simulations. We conclude with recommendations for how to perform regression calibration.
△ Less
Submitted 25 September, 2022;
originally announced September 2022.
-
Practical considerations for sandwich variance estimation in two-stage regression settings
Authors:
Lillian A. Boe,
Thomas Lumley,
Pamela A. Shaw
Abstract:
We present a practical approach for computing the sandwich variance estimator in two-stage regression model settings. As a motivating example for two-stage regression, we consider regression calibration, a popular approach for addressing covariate measurement error. The sandwich variance approach has been rarely applied in regression calibration, despite that it requires less computation time than…
▽ More
We present a practical approach for computing the sandwich variance estimator in two-stage regression model settings. As a motivating example for two-stage regression, we consider regression calibration, a popular approach for addressing covariate measurement error. The sandwich variance approach has been rarely applied in regression calibration, despite that it requires less computation time than popular resampling approaches for variance estimation, specifically the bootstrap. This is likely due to requiring specialized statistical coding. In practice, a simple bootstrap approach with Wald confidence intervals is often applied, but this approach can yield confidence intervals that do not achieve the nominal coverage level. We first outline the steps needed to compute the sandwich variance estimator. We then develop a convenient method of computation in R for sandwich variance estimation, which leverages standard regression model outputs and existing R functions and can be applied in the case of a simple random sample or complex survey design. We use a simulation study to compare the performance of the sandwich to a resampling variance approach for both data settings. Finally, we further compare these two variance estimation approaches for data examples from the Women's Health Initiative (WHI) and Hispanic Community Health Study/Study of Latinos (HCHS/SOL).
△ Less
Submitted 20 September, 2022;
originally announced September 2022.
-
Three-phase generalized raking and multiple imputation estimators to address error-prone data
Authors:
Gustavo Amorim,
Ran Tao,
Sarah Lotspeich,
Pamela A. Shaw,
Thomas Lumley,
Rena C. Patel,
Bryan E. Shepherd
Abstract:
Validation studies are often used to obtain more reliable information in settings with error-prone data. Validated data on a subsample of subjects can be used together with error-prone data on all subjects to improve estimation. In practice, more than one round of data validation may be required, and direct application of standard approaches for combining validation data into analyses may lead to…
▽ More
Validation studies are often used to obtain more reliable information in settings with error-prone data. Validated data on a subsample of subjects can be used together with error-prone data on all subjects to improve estimation. In practice, more than one round of data validation may be required, and direct application of standard approaches for combining validation data into analyses may lead to inefficient estimators since the information available from intermediate validation steps is only partially considered or even completely ignored. In this paper, we present two novel extensions of multiple imputation and generalized raking estimators that make full use of all available data. We show through simulations that incorporating information from intermediate steps can lead to substantial gains in efficiency. This work is motivated by and illustrated in a study of contraceptive effectiveness among 82,957 women living with HIV whose data were originally extracted from electronic medical records, of whom 4855 had their charts reviewed, and a subsequent 1203 also had a telephone interview to validate key study variables.
△ Less
Submitted 3 May, 2022;
originally announced May 2022.
-
A Model-Free and Finite-Population-Exact Framework for Randomized Experiments Subject to Outcome Misclassification via Integer Programming
Authors:
Siyu Heng,
Pamela A. Shaw
Abstract:
Results from randomized experiments (trials) can be severely distorted by outcome misclassification, such as from measurement error or reporting bias in binary outcomes. All existing approaches to outcome misclassification rely on some data-generating (super-population) model and therefore may not be applicable to randomized experiments without additional assumptions. We propose a model-free and f…
▽ More
Results from randomized experiments (trials) can be severely distorted by outcome misclassification, such as from measurement error or reporting bias in binary outcomes. All existing approaches to outcome misclassification rely on some data-generating (super-population) model and therefore may not be applicable to randomized experiments without additional assumptions. We propose a model-free and finite-population-exact framework for randomized experiments subject to outcome misclassification. A central quantity in our framework is "warning accuracy," defined as the threshold such that the causal conclusion drawn from the measured outcomes may differ from that based on the true outcomes if the outcome measurement accuracy did not surpass that threshold. We show how learning the warning accuracy and related concepts can benefit a randomized experiment subject to outcome misclassification. We show that the warning accuracy can be computed efficiently (even for large datasets) by adaptively reformulating an integer program with respect to the randomization design. Our framework covers both Fisher's sharp null and Neyman's weak null, works for a wide range of randomization designs, and can also be applied to observational studies adopting randomization-based inference. We apply our framework to a large randomized clinical trial for the prevention of prostate cancer.
△ Less
Submitted 27 April, 2022; v1 submitted 9 January, 2022;
originally announced January 2022.
-
Nutritional blood concentration biomarkers in the Hispanic Community Health Study/Study of Latinos: Measurement characteristics and power
Authors:
Lillian A. Boe,
Yasmin Mossavar-Rahmani,
Daniela Sotres-Alvarez,
Martha L. Daviglus,
Ramon A. Durazo-Arvizu,
Bharat Thyagarajan,
Robert C. Kaplan,
Pamela A. Shaw
Abstract:
Measurement error is a major issue in self-reported diet that can distort diet-disease relationships. Use of blood concentration biomarkers has the potential to mitigate the subjective bias inherent in self-report. As part of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) baseline visit (2008-2011), self-reported diet was collected on all participants (N=16,415). Blood concentrati…
▽ More
Measurement error is a major issue in self-reported diet that can distort diet-disease relationships. Use of blood concentration biomarkers has the potential to mitigate the subjective bias inherent in self-report. As part of the Hispanic Community Health Study/Study of Latinos (HCHS/SOL) baseline visit (2008-2011), self-reported diet was collected on all participants (N=16,415). Blood concentration biomarkers for carotenoids, tocopherols, retinol, vitamin B12 and folate were collected on a subset (N=476), as part of the Study of Latinos: Nutrition and Physical Activity Assessment Study (SOLNAS). We examine the relationship between biomarker levels, self-reported intake, Hispanic/Latino background, and other participant characteristics in this diverse cohort. We build regression calibration-based prediction equations for ten nutritional biomarkers and use a simulation to study the power of detecting a diet-disease association in a multivariable Cox model using a predicted concentration level. Good power was observed for some nutrients with high prediction model R2 values, but further research is needed to understand how best to realize the potential of these dietary biomarkers. This study provides a comprehensive examination of several nutritional biomarkers within the HCHS/SOL, characterizing their associations with subject characteristics and the influence of the measurement characteristics on the power to detect associations with health outcomes.
△ Less
Submitted 20 September, 2022; v1 submitted 22 December, 2021;
originally announced December 2021.
-
An Augmented Likelihood Approach for the Discrete Proportional Hazards Model Using Auxiliary and Validated Outcome Data -- with Application to the HCHS/SOL Study
Authors:
Lillian A. Boe,
Pamela A. Shaw
Abstract:
In large epidemiologic studies, it is typical for an inexpensive, non-invasive procedure to be used to record disease status during regular follow-up visits, with less frequent assessment by a gold standard test. Inexpensive outcome measures like self-reported disease status are practical to obtain, but can be error-prone. Association analysis reliant on error-prone outcomes may lead to biased res…
▽ More
In large epidemiologic studies, it is typical for an inexpensive, non-invasive procedure to be used to record disease status during regular follow-up visits, with less frequent assessment by a gold standard test. Inexpensive outcome measures like self-reported disease status are practical to obtain, but can be error-prone. Association analysis reliant on error-prone outcomes may lead to biased results; however, restricting analyses to only data from the less frequently observed error-free outcome could be inefficient. We have developed an augmented likelihood that incorporates data from both error-prone outcomes and a gold standard assessment. We conduct a numerical study to show how we can improve statistical efficiency by using the proposed method over standard approaches for interval-censored survival data that do not leverage auxiliary data. We extend this method for the complex survey design setting so that it can be applied in our motivating data example. Our method is applied to data from the Hispanic Community Health Study/Study of Latinos to assess the association between energy and protein intake and the risk of incident diabetes. In our application, we demonstrate how our method can be used in combination with regression calibration to additionally address the covariate measurement error in self-reported diet.
△ Less
Submitted 20 September, 2022; v1 submitted 24 November, 2021;
originally announced November 2021.
-
Analysis of Error-prone Electronic Health Records with Multi-wave Validation Sampling: Association of Maternal Weight Gain during Pregnancy with Childhood Outcomes
Authors:
Bryan E. Shepherd,
Kyunghee Han,
Tong Chen,
Aihua Bian,
Shannon Pugh,
Stephany N. Duda,
Thomas Lumley,
William J. Heerman,
Pamela A. Shaw
Abstract:
Electronic health record (EHR) data are increasingly used for biomedical research, but these data have recognized data quality challenges. Data validation is necessary to use EHR data with confidence, but limited resources typically make complete data validation impossible. Using EHR data, we illustrate prospective, multi-wave, two-phase validation sampling to estimate the association between mate…
▽ More
Electronic health record (EHR) data are increasingly used for biomedical research, but these data have recognized data quality challenges. Data validation is necessary to use EHR data with confidence, but limited resources typically make complete data validation impossible. Using EHR data, we illustrate prospective, multi-wave, two-phase validation sampling to estimate the association between maternal weight gain during pregnancy and the risks of her child develo** obesity or asthma. The optimal validation sampling design depends on the unknown efficient influence functions of regression coefficients of interest. In the first wave of our multi-wave validation design, we estimate the influence function using the unvalidated (phase 1) data to determine our validation sample; then in subsequent waves, we re-estimate the influence function using validated (phase 2) data and update our sampling. For efficiency, estimation combines obesity and asthma sampling frames while calibrating sampling weights using generalized raking. We validated 996 of 10,335 mother-child EHR dyads in 6 sampling waves. Estimated associations between childhood obesity/asthma and maternal weight gain, as well as other covariates, are compared to naive estimates that only use unvalidated data. In some cases, estimates markedly differ, underscoring the importance of efficient validation sampling to obtain accurate estimates incorporating validated data.
△ Less
Submitted 28 September, 2021;
originally announced September 2021.
-
Optimal Multi-Wave Validation of Secondary Use Data with Outcome and Exposure Misclassification
Authors:
Sarah C. Lotspeich,
Gustavo G. C. Amorim,
Pamela A. Shaw,
Ran Tao,
Bryan E. Shepherd
Abstract:
The growing availability of observational databases like electronic health records (EHR) provides unprecedented opportunities for secondary use of such data in biomedical research. However, these data can be error-prone and need to be validated before use. It is usually unrealistic to validate the whole database due to resource constraints. A cost-effective alternative is to implement a two-phase…
▽ More
The growing availability of observational databases like electronic health records (EHR) provides unprecedented opportunities for secondary use of such data in biomedical research. However, these data can be error-prone and need to be validated before use. It is usually unrealistic to validate the whole database due to resource constraints. A cost-effective alternative is to implement a two-phase design that validates a subset of patient records that are enriched for information about the research question of interest. Herein, we consider odds ratio estimation under differential outcome and exposure misclassification. We propose optimal designs that minimize the variance of the maximum likelihood odds ratio estimator. We develop a novel adaptive grid search algorithm that can locate the optimal design in a computationally feasible and numerically accurate manner. Because the optimal design requires specification of unknown parameters at the outset and thus is unattainable without prior information, we introduce a multi-wave sampling strategy to approximate it in practice. We demonstrate the efficiency gains of the proposed designs over existing ones through extensive simulations and two large observational studies. We provide an R package and Shiny app to facilitate the use of the optimal designs.
△ Less
Submitted 12 September, 2022; v1 submitted 30 August, 2021;
originally announced August 2021.
-
Optimum Allocation for Adaptive Multi-Wave Sampling in R: The R Package optimall
Authors:
Jasper B. Yang,
Bryan E. Shepherd,
Thomas Lumley,
Pamela A. Shaw
Abstract:
The R package optimall offers a collection of functions that efficiently streamline the design process of sampling in surveys ranging from simple to complex. The package's main functions allow users to interactively define and adjust strata cut points based on values or quantiles of auxiliary covariates, adaptively calculate the optimum number of samples to allocate to each stratum using Neyman or…
▽ More
The R package optimall offers a collection of functions that efficiently streamline the design process of sampling in surveys ranging from simple to complex. The package's main functions allow users to interactively define and adjust strata cut points based on values or quantiles of auxiliary covariates, adaptively calculate the optimum number of samples to allocate to each stratum using Neyman or Wright allocation, and select specific IDs to sample based on a stratified sampling design. Using real-life epidemiological study examples, we demonstrate how optimall facilitates an efficient workflow for the design and implementation of surveys in R. Although tailored towards multi-wave sampling under two- or three-phase designs, the R package optimall may be useful for any sampling survey.
△ Less
Submitted 17 June, 2021;
originally announced June 2021.
-
Improved Generalized Raking Estimators to Address Dependent Covariate and Failure-Time Outcome Error
Authors:
Eric J. Oh,
Bryan E. Shepherd,
Thomas Lumley,
Pamela A. Shaw
Abstract:
Biomedical studies that use electronic health records (EHR) data for inference are often subject to bias due to measurement error. The measurement error present in EHR data is typically complex, consisting of errors of unknown functional form in covariates and the outcome, which can be dependent. To address the bias resulting from such errors, generalized raking has recently been proposed as a rob…
▽ More
Biomedical studies that use electronic health records (EHR) data for inference are often subject to bias due to measurement error. The measurement error present in EHR data is typically complex, consisting of errors of unknown functional form in covariates and the outcome, which can be dependent. To address the bias resulting from such errors, generalized raking has recently been proposed as a robust method that yields consistent estimates without the need to model the error structure. We provide rationale for why these previously proposed raking estimators can be expected to be inefficient in failure-time outcome settings involving misclassification of the event indicator. We propose raking estimators that utilize multiple imputation, to impute either the target variables or auxiliary variables, to improve the efficiency. We also consider outcome-dependent sampling designs and investigate their impact on the efficiency of the raking estimators, either with or without multiple imputation. We present an extensive numerical study to examine the performance of the proposed estimators across various measurement error settings. We then apply the proposed methods to our motivating setting, in which we seek to analyze HIV outcomes in an observational cohort with electronic health records data from the Vanderbilt Comprehensive Care Clinic.
△ Less
Submitted 12 June, 2020;
originally announced June 2020.
-
Two-phase analysis and study design for survival models with error-prone exposures
Authors:
Kyunghee Han,
Thomas Lumley,
Bryan E. Shepherd,
Pamela A. Shaw
Abstract:
Increasingly, medical research is dependent on data collected for non-research purposes, such as electronic health records data (EHR). EHR data and other large databases can be prone to measurement error in key exposures, and unadjusted analyses of error-prone data can bias study results. Validating a subset of records is a cost-effective way of gaining information on the error structure, which in…
▽ More
Increasingly, medical research is dependent on data collected for non-research purposes, such as electronic health records data (EHR). EHR data and other large databases can be prone to measurement error in key exposures, and unadjusted analyses of error-prone data can bias study results. Validating a subset of records is a cost-effective way of gaining information on the error structure, which in turn can be used to adjust analyses for this error and improve inference. We extend the mean score method for the two-phase analysis of discrete-time survival models, which uses the unvalidated covariates as auxiliary variables that act as surrogates for the unobserved true exposures. This method relies on a two-phase sampling design and an estimation approach that preserves the consistency of complete case regression parameter estimates in the validated subset, with increased precision leveraged from the auxiliary data. Furthermore, we develop optimal sampling strategies which minimize the variance of the mean score estimator for a target exposure under a fixed cost constraint. We consider the setting where an internal pilot is necessary for the optimal design so that the phase two sample is split into a pilot and an adaptive optimal sample. Through simulations and data example, we evaluate efficiency gains of the mean score estimator using the derived optimal validation design compared to balanced and simple random sampling for the phase two sample. We also empirically explore efficiency gains that the proposed discrete optimal design can provide for the Cox proportional hazards model in the setting of a continuous-time survival outcome.
△ Less
Submitted 11 May, 2020;
originally announced May 2020.
-
An Approximate Quasi-Likelihood Approach for Error-Prone Failure Time Outcomes and Exposures
Authors:
Lillian A. Boe,
Lesley F. Tinker,
Pamela A. Shaw
Abstract:
Measurement error arises commonly in clinical research settings that rely on data from electronic health records or large observational cohorts. In particular, self-reported outcomes are typical in cohort studies for chronic diseases such as diabetes in order to avoid the burden of expensive diagnostic tests. Dietary intake, which is also commonly collected by self-report and subject to measuremen…
▽ More
Measurement error arises commonly in clinical research settings that rely on data from electronic health records or large observational cohorts. In particular, self-reported outcomes are typical in cohort studies for chronic diseases such as diabetes in order to avoid the burden of expensive diagnostic tests. Dietary intake, which is also commonly collected by self-report and subject to measurement error, is a major factor linked to diabetes and other chronic diseases. These errors can bias exposure-disease associations that ultimately can mislead clinical decision-making. We have extended an existing semiparametric likelihood-based method for handling error-prone, discrete failure time outcomes to also address covariate error. We conduct an extensive numerical study to compare the proposed method to the naive approach that ignores measurement error in terms of bias and efficiency in the estimation of the regression parameter of interest. In all settings considered, the proposed method showed minimal bias and maintained coverage probability, thus outperforming the naive analysis which showed extreme bias and low coverage. This method is applied to data from the Women's Health Initiative to assess the association between energy and protein intake and the risk of incident diabetes mellitus. Our results show that correcting for errors in both the self-reported outcome and dietary exposures leads to considerably different hazard ratio estimates than those from analyses that ignore measurement error, which demonstrates the importance of correcting for both outcome and covariate error. Computational details and R code for implementing the proposed method are presented in Section S1 of the Supplementary Materials.
△ Less
Submitted 4 February, 2021; v1 submitted 2 April, 2020;
originally announced April 2020.
-
Combining multiple imputation with raking of weights: An efficient and robust approach in the setting of nearly-true models
Authors:
Kyunghee Han,
Pamela A. Shaw,
Thomas Lumley
Abstract:
Multiple imputation provides us with efficient estimators in model-based methods for handling missing data under the true model. It is also well-understood that design-based estimators are robust methods that do not require accurately modeling the missing data; however, they can be inefficient. In any applied setting, it is difficult to know whether a missing data model may be good enough to win t…
▽ More
Multiple imputation provides us with efficient estimators in model-based methods for handling missing data under the true model. It is also well-understood that design-based estimators are robust methods that do not require accurately modeling the missing data; however, they can be inefficient. In any applied setting, it is difficult to know whether a missing data model may be good enough to win the bias-efficiency trade-off. Raking of weights is one approach that relies on constructing an auxiliary variable from data observed on the full cohort, which is then used to adjust the weights for the usual Horvitz-Thompson estimator. Computing the optimally efficient raking estimator requires evaluating the expectation of the efficient score given the full cohort data, which is generally infeasible. We demonstrate multiple imputation (MI) as a practical method to compute a raking estimator that will be optimal. We compare this estimator to common parametric and semi-parametric estimators, including standard multiple imputation. We show that while estimators, such as the semi-parametric maximum likelihood and MI estimator, obtain optimal performance under the true model, the proposed raking estimator utilizing MI maintains a better robustness-efficiency trade-off even under mild model misspecification. We also show that the standard raking estimator, without MI, is often competitive with the optimal raking estimator. We demonstrate these properties through several numerical examples and provide a theoretical discussion of conditions for asymptotically superior relative efficiency of the proposed raking estimator.
△ Less
Submitted 9 June, 2020; v1 submitted 2 October, 2019;
originally announced October 2019.
-
Regression to the Mean's Impact on the Synthetic Control Method: Bias and Sensitivity Analysis
Authors:
Nicholas Illenberger,
Dylan S. Small,
Pamela A. Shaw
Abstract:
To make informed policy recommendations from observational data, we must be able to discern true treatment effects from random noise and effects due to confounding. Difference-in-Difference techniques which match treated units to control units based on pre-treatment outcomes, such as the synthetic control approach have been presented as principled methods to account for confounding. However, we sh…
▽ More
To make informed policy recommendations from observational data, we must be able to discern true treatment effects from random noise and effects due to confounding. Difference-in-Difference techniques which match treated units to control units based on pre-treatment outcomes, such as the synthetic control approach have been presented as principled methods to account for confounding. However, we show that use of synthetic controls or other matching procedures can introduce regression to the mean (RTM) bias into estimates of the average treatment effect on the treated. Through simulations, we show RTM bias can lead to inflated type I error rates as well as decreased power in typical policy evaluation settings. Further, we provide a novel correction for RTM bias which can reduce bias and attain appropriate type I error rates. This correction can be used to perform a sensitivity analysis which determines how results may be affected by RTM. We use our proposed correction and sensitivity analysis to reanalyze data concerning the effects of California's Proposition 99, a large-scale tobacco control program, on statewide smoking rates.
△ Less
Submitted 10 September, 2019;
originally announced September 2019.
-
Raking and Regression Calibration: Methods to Address Bias from Correlated Covariate and Time-to-Event Error
Authors:
Eric J. Oh,
Bryan E. Shepherd,
Thomas Lumley,
Pamela A. Shaw
Abstract:
Medical studies that depend on electronic health records (EHR) data are often subject to measurement error, as the data are not collected to support research questions under study. These data errors, if not accounted for in study analyses, can obscure or cause spurious associations between patient exposures and disease risk. Methodology to address covariate measurement error has been well develope…
▽ More
Medical studies that depend on electronic health records (EHR) data are often subject to measurement error, as the data are not collected to support research questions under study. These data errors, if not accounted for in study analyses, can obscure or cause spurious associations between patient exposures and disease risk. Methodology to address covariate measurement error has been well developed; however, time-to-event error has also been shown to cause significant bias but methods to address it are relatively underdeveloped. More generally, it is possible to observe errors in both the covariate and the time-to-event outcome that are correlated. We propose regression calibration (RC) estimators to simultaneously address correlated error in the covariates and the censored event time. Although RC can perform well in many settings with covariate measurement error, it is biased for nonlinear regression models, such as the Cox model. Thus, we additionally propose raking estimators which are consistent estimators of the parameter defined by the population estimating equation. Raking can improve upon RC in certain settings with failure-time data, require no explicit modeling of the error structure, and can be utilized under outcome-dependent sampling designs. We discuss features of the underlying estimation problem that affect the degree of improvement the raking estimator has over the RC approach. Detailed simulation studies are presented to examine the performance of the proposed estimators under varying levels of signal, error, and censoring. The methodology is illustrated on observational EHR data on HIV outcomes from the Vanderbilt Comprehensive Care Clinic.
△ Less
Submitted 9 March, 2020; v1 submitted 20 May, 2019;
originally announced May 2019.
-
Epidemiologic analyses with error-prone exposures: Review of current practice and recommendations
Authors:
Pamela A. Shaw,
Veronika Deffner,
Ruth H. Keogh,
Janet A. Tooze,
Kevin W. Dodd,
Helmut Küchenhoff,
Victor Kipnis,
Laurence S. Freedman
Abstract:
Background: Variables in epidemiological observational studies are commonly subject to measurement error and misclassification, but the impact of such errors is frequently not appreciated or ignored. As part of the STRengthening Analytical Thinking for Observational Studies (STRATOS) Initiative, a Task Group on measurement error and misclassification (TG4) seeks to describe the scope of this probl…
▽ More
Background: Variables in epidemiological observational studies are commonly subject to measurement error and misclassification, but the impact of such errors is frequently not appreciated or ignored. As part of the STRengthening Analytical Thinking for Observational Studies (STRATOS) Initiative, a Task Group on measurement error and misclassification (TG4) seeks to describe the scope of this problem and the analysis methods currently in use to address measurement error. Methods: TG4 conducted a literature survey of four types of research studies that are typically impacted by exposure measurement error: 1) dietary intake cohort studies, 2) dietary intake population surveys, 3) physical activity cohort studies, and 4) air pollution cohort studies. The survey was conducted to understand current practice for acknowledging and addressing measurement error. Results: The survey revealed that while researchers were generally aware that measurement error affected their studies, very few adjusted their analysis for the error. Most articles provided incomplete discussion of the potential effects of measurement error on their results. Regression calibration was the most widely used method of adjustment. Conclusions: Even in areas of epidemiology where measurement error is a known problem, the dominant current practice is to ignore errors in analyses. Methods to correct for measurement error are available but require additional data to inform the error structure. There is a great need to incorporate such data collection within study designs and improve the analytical approach. Increased efforts by investigators, editors and reviewers are also needed to improve presentation of research when data are subject to error.
△ Less
Submitted 28 February, 2018;
originally announced February 2018.