-
Running in circles: is practical application feasible for data fission and data thinning in post-clustering differential analysis?
Authors:
Benjamin Hivert,
Denis Agniel,
Rodolphe ThiƩbaut,
Boris P. Hejblum
Abstract:
The standard pipeline to analyse single-cell RNA sequencing (scRNA-seq) often involves two steps : clustering and Differential Expression Analysis (DEA) to annotate cell populations based on gene expression. However, using clustering results for data-driven hypothesis formulation compromises statistical properties, especially Type I error control. Data fission was introduced to split the informati…
▽ More
The standard pipeline to analyse single-cell RNA sequencing (scRNA-seq) often involves two steps : clustering and Differential Expression Analysis (DEA) to annotate cell populations based on gene expression. However, using clustering results for data-driven hypothesis formulation compromises statistical properties, especially Type I error control. Data fission was introduced to split the information contained in each observation into two independent parts that can be used for clustering and testing. However, data fission was originally designed for non-mixture distributions, and adapting it for mixtures requires knowledge of the unknown clustering structure to estimate component-specific scale parameters. As components are typically unavailable in practice, scale parameter estimators often exhibit bias. We explicitly quantify how this bias affects subsequent post-clustering differential analysis Type I error rate despite employing data fission. In response, we propose a novel approach that involves modeling each observation as a realization of its distribution, with scale parameters estimated non-parametrically. Simulations study showcase the efficacy of our method when component are clearly separated. However, the level of separability required to reach good performance presents complexities in its application to real scRNA-seq data.
△ Less
Submitted 22 May, 2024;
originally announced May 2024.
-
Robust Evaluation of Longitudinal Surrogate Markers with Censored Data
Authors:
Denis Agniel,
Layla Parast
Abstract:
The development of statistical methods to evaluate surrogate markers is an active area of research. In many clinical settings, the surrogate marker is not simply a single measurement but is instead a longitudinal trajectory of measurements over time, e.g., fasting plasma glucose measured every 6 months for 3 years. In general, available methods developed for the single-surrogate setting cannot acc…
▽ More
The development of statistical methods to evaluate surrogate markers is an active area of research. In many clinical settings, the surrogate marker is not simply a single measurement but is instead a longitudinal trajectory of measurements over time, e.g., fasting plasma glucose measured every 6 months for 3 years. In general, available methods developed for the single-surrogate setting cannot accommodate a longitudinal surrogate marker. Furthermore, many of the methods have not been developed for use with primary outcomes that are time-to-event outcomes and/or subject to censoring. In this paper, we propose robust methods to evaluate a longitudinal surrogate marker in a censored time-to-event outcome setting. Specifically, we propose a method to define and estimate the proportion of the treatment effect on a censored primary outcome that is explained by the treatment effect on a longitudinal surrogate marker measured up to time $t_0$. We accommodate both potential censoring of the primary outcome and of the surrogate marker. A simulation study demonstrates good finite-sample performance of our proposed methods. We illustrate our procedures by examining repeated measures of fasting plasma glucose, a surrogate marker for diabetes diagnosis, using data from the Diabetes Prevention Program (DPP).
△ Less
Submitted 26 February, 2024;
originally announced February 2024.
-
De-Biasing the Bias: Methods for Improving Disparity Assessments with Noisy Group Measurements
Authors:
Solvejg Wastvedt,
Joshua Snoke,
Denis Agniel,
Julie Lai,
Marc N. Elliott,
Steven C. Martino
Abstract:
Health care decisions are increasingly informed by clinical decision support algorithms, but these algorithms may perpetuate or increase racial and ethnic disparities in access to and quality of health care. Further complicating the problem, clinical data often have missing or poor quality racial and ethnic information, which can lead to misleading assessments of algorithmic bias. We present novel…
▽ More
Health care decisions are increasingly informed by clinical decision support algorithms, but these algorithms may perpetuate or increase racial and ethnic disparities in access to and quality of health care. Further complicating the problem, clinical data often have missing or poor quality racial and ethnic information, which can lead to misleading assessments of algorithmic bias. We present novel statistical methods that allow for the use of probabilities of racial/ethnic group membership in assessments of algorithm performance and quantify the statistical bias that results from error in these imputed group probabilities. We propose a sensitivity analysis approach to estimating the statistical bias that allows practitioners to assess disparities in algorithm performance under a range of assumed levels of group probability error. We also prove theoretical bounds on the statistical bias for a set of commonly used fairness metrics and describe real-world scenarios where our theoretical results are likely to apply. We present a case study using imputed race and ethnicity from the Bayesian Improved Surname Geocoding (BISG) algorithm for estimation of disparities in a clinical decision support algorithm used to inform osteoporosis treatment. Our novel methods allow policy makers to understand the range of potential disparities under a given algorithm even when race and ethnicity information is missing and to make informed decisions regarding the implementation of machine learning for clinical decision support.
△ Less
Submitted 26 February, 2024; v1 submitted 20 February, 2024;
originally announced February 2024.
-
Generalized difference-in-differences
Authors:
Denis Agniel,
Max Rubinstein,
Jessie Coe,
Maria DeYoreo
Abstract:
We propose a new method for estimating causal effects in longitudinal/panel data settings that we call generalized difference-in-differences. Our approach unifies two alternative approaches in these settings: ignorability estimators (e.g., synthetic controls) and difference-in-differences (DiD) estimators. We propose a new identifying assumption -- a stable bias assumption -- which generalizes the…
▽ More
We propose a new method for estimating causal effects in longitudinal/panel data settings that we call generalized difference-in-differences. Our approach unifies two alternative approaches in these settings: ignorability estimators (e.g., synthetic controls) and difference-in-differences (DiD) estimators. We propose a new identifying assumption -- a stable bias assumption -- which generalizes the conditional parallel trends assumption in DiD, leading to the proposed generalized DiD framework. This change gives generalized DiD estimators the flexibility of ignorability estimators while maintaining the robustness to unobserved confounding of DiD. We also show how ignorability and DiD estimators are special cases of generalized DiD. We then propose influence-function based estimators of the observed data functional, allowing the use of double/debiased machine learning for estimation. We also show how generalized DiD easily extends to include clustered treatment assignment and staggered adoption settings, and we discuss how the framework can facilitate estimation of other treatment effects beyond the average treatment effect on the treated. Finally, we provide simulations which show that generalized DiD outperforms ignorability and DiD estimators when their identifying assumptions are not met, while being competitive with these special cases when their identifying assumptions are met.
△ Less
Submitted 8 December, 2023;
originally announced December 2023.
-
Post-clustering difference testing: valid inference and practical considerations
Authors:
Benjamin Hivert,
Denis Agniel,
Rodolphe ThiƩbaut,
Boris P Hejblum
Abstract:
Clustering is part of unsupervised analysis methods that consist in grou** samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to infer the variables that significantly separate the estimated clusters from each other. However, data-driven hypotheses are considered for the inference process,…
▽ More
Clustering is part of unsupervised analysis methods that consist in grou** samples into homogeneous and separate subgroups of observations also called clusters. To interpret the clusters, statistical hypothesis testing is often used to infer the variables that significantly separate the estimated clusters from each other. However, data-driven hypotheses are considered for the inference process, since the hypotheses are derived from the clustering results. This double use of the data leads traditional hypothesis test to fail to control the Type I error rate particularly because of uncertainty in the clustering process and the potential artificial differences it could create. We propose three novel statistical hypothesis tests which account for the clustering process. Our tests efficiently control the Type I error rate by identifying only variables that contain a true signal separating groups of observations.
△ Less
Submitted 24 October, 2022;
originally announced October 2022.
-
Doubly-robust evaluation of high-dimensional surrogate markers
Authors:
Denis Agniel,
Layla Parast,
Boris Hejblum
Abstract:
When evaluating the effectiveness of a treatment, policy, or intervention, the desired measure of effectiveness may be expensive to collect, not routinely available, or may take a long time to occur. In these cases, it is sometimes possible to identify a surrogate outcome that can more easily/quickly/cheaply capture the effect of interest. Theory and methods for evaluating the strength of surrogat…
▽ More
When evaluating the effectiveness of a treatment, policy, or intervention, the desired measure of effectiveness may be expensive to collect, not routinely available, or may take a long time to occur. In these cases, it is sometimes possible to identify a surrogate outcome that can more easily/quickly/cheaply capture the effect of interest. Theory and methods for evaluating the strength of surrogate markers have been well studied in the context of a single surrogate marker measured in the course of a randomized clinical study. However, methods are lacking for quantifying the utility of surrogate markers when the dimension of the surrogate grows and/or when study data are observational. We propose an efficient nonparametric method for evaluating high-dimensional surrogate markers in studies where the treatment need not be randomized. Our approach draws on a connection between quantifying the utility of a surrogate marker and the most fundamental tools of causal inference -- namely, methods for estimating the average treatment effect. We show that recently developed methods for incorporating machine learning methods into the estimation of average treatment effects can be used for evaluating surrogate markers. This allows us to derive limiting asymptotic distributions for key quantities, and we demonstrate their good performance in simulation.
△ Less
Submitted 2 December, 2020; v1 submitted 2 December, 2020;
originally announced December 2020.
-
Synthetic estimation for the complier average causal effect
Authors:
Denis Agniel,
Bing Han,
Matthew Cefalu
Abstract:
We propose an improved estimator of the complier average causal effect (CACE). Researchers typically choose a presumably-unbiased estimator for the CACE in studies with noncompliance, when many other lower-variance estimators may be available. We propose a synthetic estimator that combines information across all available estimators, leveraging the efficiency in lower-variance estimators while mai…
▽ More
We propose an improved estimator of the complier average causal effect (CACE). Researchers typically choose a presumably-unbiased estimator for the CACE in studies with noncompliance, when many other lower-variance estimators may be available. We propose a synthetic estimator that combines information across all available estimators, leveraging the efficiency in lower-variance estimators while maintaining low bias. Our approach minimizes an estimate of the mean squared error of all convex combinations of the candidate estimators. We derive the asymptotic distribution of the synthetic estimator and demonstrate its good performance in simulation, displaying a robustness to inclusion of even high-bias estimators.
△ Less
Submitted 12 September, 2019;
originally announced September 2019.
-
Functional principal variance component testing for a genetic association study of HIV progression
Authors:
Denis Agniel,
Wen Xie,
Myron Essex,
Tianxi Cai
Abstract:
HIV-1C is the most prevalent subtype of HIV-1 and accounts for over half of HIV-1 infections worldwide. Host genetic influence of HIV infection has been previously studied in HIV-1B, but little attention has been paid to the more prevalent subtype C. To understand the role of host genetics in HIV-1C disease progression, we perform a study to assess the association between longitudinally collected…
▽ More
HIV-1C is the most prevalent subtype of HIV-1 and accounts for over half of HIV-1 infections worldwide. Host genetic influence of HIV infection has been previously studied in HIV-1B, but little attention has been paid to the more prevalent subtype C. To understand the role of host genetics in HIV-1C disease progression, we perform a study to assess the association between longitudinally collected measures of disease and more than 100,000 genetic markers located on chromosome 6. The most common approach to analyzing longitudinal data in this context is linear mixed effects models, which may be overly simplistic in this case. On the other hand, existing non-parametric methods may suffer from low power due to high degrees of freedom (DF) and may be computationally infeasible at the large scale. We propose a functional principal variance component (FPVC) testing framework which captures the nonlinearity in the CD4 and viral load with potentially low DF and is fast enough to carry out thousands or millions of times. The FPVC testing unfolds in two stages. In the first stage, we summarize the markers of disease progression according to their major patterns of variation via functional principal components analysis (FPCA). In the second stage, we employ a simple working model and variance component testing to examine the association between the summaries of disease progression and a set of single nucleotide polymorphisms. We supplement this analysis with simulation results which indicate that FPVC testing can offer large power gains over the standard linear mixed effects model.
△ Less
Submitted 9 June, 2017;
originally announced June 2017.
-
Doubly robust matching estimators for high dimensional confounding adjustment
Authors:
Joseph Antonelli,
Matthew Cefalu,
Nathan Palmer,
Denis Agniel
Abstract:
Valid estimation of treatment effects from observational data requires proper control of confounding. If the number of covariates is large relative to the number of observations, then controlling for all available covariates is infeasible. In cases where a sparsity condition holds, variable selection or penalization can reduce the dimension of the covariate space in a manner that allows for valid…
▽ More
Valid estimation of treatment effects from observational data requires proper control of confounding. If the number of covariates is large relative to the number of observations, then controlling for all available covariates is infeasible. In cases where a sparsity condition holds, variable selection or penalization can reduce the dimension of the covariate space in a manner that allows for valid estimation of treatment effects. In this article, we propose matching on both the estimated propensity score and the estimated prognostic scores when the number of covariates is large relative to the number of observations. We derive asymptotic results for the matching estimator and show that it is doubly robust, in the sense that only one of the two score models need be correct to obtain a consistent estimator. We show via simulation its effectiveness in controlling for confounding and highlight its potential to address nonlinear confounding. Finally, we apply the proposed procedure to analyze the effect of gender on prescription opioid use using insurance claims data.
△ Less
Submitted 10 January, 2018; v1 submitted 1 December, 2016;
originally announced December 2016.
-
Variance component score test for time-course gene set analysis of longitudinal RNA-seq data
Authors:
Denis Agniel,
Boris P Hejblum
Abstract:
As gene expression measurement technology is shifting from microarrays to sequencing, the statistical tools available for their analysis must be adapted since RNA-seq data are measured as counts. Recently, it has been proposed to tackle the count nature of these data by modeling log-count reads per million as continuous variables, using nonparametric regression to account for their inherent hetero…
▽ More
As gene expression measurement technology is shifting from microarrays to sequencing, the statistical tools available for their analysis must be adapted since RNA-seq data are measured as counts. Recently, it has been proposed to tackle the count nature of these data by modeling log-count reads per million as continuous variables, using nonparametric regression to account for their inherent heteroscedasticity. Adopting such a framework, we propose tcgsaseq, a principled, model-free and efficient top-down method for detecting longitudinal changes in RNA-seq gene sets. Considering gene sets defined a priori, tcgsaseq identifies those whose expression vary over time, based on an original variance component score test accounting for both covariates and heteroscedasticity without assuming any specific parametric distribution for the transformed counts. We demonstrate that despite the presence of a nonparametric component, our test statistic has a simple form and limiting distribution, and both may be computed quickly. A permutation version of the test is additionally proposed for very small sample sizes. Applied to both simulated data and two real datasets, the proposed method is shown to exhibit very good statistical properties, with an increase in stability and power when compared to state of the art methods ROAST, edgeR and DESeq2, which can fail to control the type I error under certain realistic settings. We have made the method available for the community in the R package tcgsaseq.
△ Less
Submitted 6 January, 2017; v1 submitted 8 May, 2016;
originally announced May 2016.
-
Estimation and testing for multiple regulation of multivariate mixed outcomes
Authors:
Denis Agniel,
Katherine P. Liao,
Tianxi Cai
Abstract:
Considerable interest has recently been focused on studying multiple phenotypes simultaneously in both epidemiological and genomic studies, either to capture the multidimensionality of complex disorders or to understand shared etiology of related disorders. We seek to identify {\em multiple regulators} or predictors that are associated with multiple outcomes when these outcomes may be measured on…
▽ More
Considerable interest has recently been focused on studying multiple phenotypes simultaneously in both epidemiological and genomic studies, either to capture the multidimensionality of complex disorders or to understand shared etiology of related disorders. We seek to identify {\em multiple regulators} or predictors that are associated with multiple outcomes when these outcomes may be measured on very different scales or composed of a mixture of continuous, binary, and not-fully-observed elements. We first propose an estimation technique to put all effects on similar scales, and we induce sparsity on the estimated effects. We provide standard asymptotic results for this estimator and show that resampling can be used to quantify uncertainty in finite samples. We finally provide a multiple testing procedure which can be geared specifically to the types of multiple regulators of interest, and we establish that, under standard regularity conditions, the familywise error rate will approach 0 as sample size diverges. Simulation results indicate that our approach can improve over unregularized methods both in reducing bias in estimation and improving power for testing.
△ Less
Submitted 25 November, 2015;
originally announced November 2015.