-
Map**-to-Parameter Nonlinear Functional Regression with Novel B-spline Free Knot Placement Algorithm
Authors:
Chengdong Shi,
Ching-Hsun Tseng,
Wei Zhao,
Xiao-Jun Zeng
Abstract:
We propose a novel approach to nonlinear functional regression, called the Map**-to-Parameter function model, which addresses complex and nonlinear functional regression problems in parameter space by employing any supervised learning technique. Central to this model is the map** of function data from an infinite-dimensional function space to a finite-dimensional parameter space. This is accom…
▽ More
We propose a novel approach to nonlinear functional regression, called the Map**-to-Parameter function model, which addresses complex and nonlinear functional regression problems in parameter space by employing any supervised learning technique. Central to this model is the map** of function data from an infinite-dimensional function space to a finite-dimensional parameter space. This is accomplished by concurrently approximating multiple functions with a common set of B-spline basis functions by any chosen order, with their knot distribution determined by the Iterative Local Placement Algorithm, a newly proposed free knot placement algorithm. In contrast to the conventional equidistant knot placement strategy that uniformly distributes knot locations based on a predefined number of knots, our proposed algorithms determine knot location according to the local complexity of the input or output functions. The performance of our knot placement algorithms is shown to be robust in both single-function approximation and multiple-function approximation contexts. Furthermore, the effectiveness and advantage of the proposed prediction model in handling both function-on-scalar regression and function-on-function regression problems are demonstrated through several real data applications, in comparison with four groups of state-of-the-art methods.
△ Less
Submitted 26 January, 2024;
originally announced January 2024.
-
Guided Semi-Supervised Non-negative Matrix Factorization on Legal Documents
Authors:
Pengyu Li,
Christine Tseng,
Yaxuan Zheng,
Joyce A. Chew,
Longxiu Huang,
Benjamin Jarman,
Deanna Needell
Abstract:
Classification and topic modeling are popular techniques in machine learning that extract information from large-scale datasets. By incorporating a priori information such as labels or important features, methods have been developed to perform classification and topic modeling tasks; however, most methods that can perform both do not allow for guidance of the topics or features. In this paper, we…
▽ More
Classification and topic modeling are popular techniques in machine learning that extract information from large-scale datasets. By incorporating a priori information such as labels or important features, methods have been developed to perform classification and topic modeling tasks; however, most methods that can perform both do not allow for guidance of the topics or features. In this paper, we propose a method, namely Guided Semi-Supervised Non-negative Matrix Factorization (GSSNMF), that performs both classification and topic modeling by incorporating supervision from both pre-assigned document class labels and user-designed seed words. We test the performance of this method through its application to legal documents provided by the California Innocence Project, a nonprofit that works to free innocent convicted persons and reform the justice system. The results show that our proposed method improves both classification accuracy and topic coherence in comparison to past methods like Semi-Supervised Non-negative Matrix Factorization (SSNMF) and Guided Non-negative Matrix Factorization (Guided NMF).
△ Less
Submitted 31 January, 2022;
originally announced January 2022.
-
Association study between gene expression and multiple phenotypes in omics applications of complex diseases
Authors:
Yujia Li,
Yusi Fang,
Peng Liu,
George C. Tseng
Abstract:
Studying phenotype-gene association can uncover mechanism of diseases and develop efficient treatments. In complex disease where multiple phenotypes are available and correlated, analyzing and interpreting associated genes for each phenotype respectively may decrease statistical power and lose intepretation due to not considering the correlation between phenotypes. The typical approaches are many…
▽ More
Studying phenotype-gene association can uncover mechanism of diseases and develop efficient treatments. In complex disease where multiple phenotypes are available and correlated, analyzing and interpreting associated genes for each phenotype respectively may decrease statistical power and lose intepretation due to not considering the correlation between phenotypes. The typical approaches are many global testing methods, such as multivariate analysis of variance (MANOVA), which tests the overall association between phenotypes and each gene, without considersing the heterogeneity among phenotypes. In this paper, we extend and evaluate two p-value combination methods, adaptive weighted Fisher's method (AFp) and adaptive Fisher's method (AFz), to tackle this problem, where AFp stands out as our final proposed method, based on extensive simulations and a real application. Our proposed AFp method has three advantages over traditional global testing methods. Firstly, it can consider the heterogeneity of phenotypes and determines which specific phenotypes a gene is associated with, using phenotype specific 0-1 weights. Secondly, AFp takes the p-values from the test of association of each phenotype as input, thus can accommodate different types of phenotypes (continuous, binary and count). Thirdly, we also apply bootstrap** to construct a variability index for the weight estimator of AFp and generate a co-membership matrix to categorize (cluster) genes based on their association-patterns for intuitive biological investigations. Through extensive simulations, AFp shows superior performance over global testing methods in terms of type I error control and statistical power, as well as higher accuracy of 0-1 weights estimation over AFz. A real omics application with transcriptomic and clinical data of complex lung diseases demonstrates insightful biological findings of AFp.
△ Less
Submitted 10 December, 2021;
originally announced December 2021.
-
Sample size planning for pilot studies
Authors:
Chi-Hong Tseng,
Danielle Sim
Abstract:
Pilot studies are often the first step of experimental research. It is usually on a smaller scale and the results can inform intervention development, study feasibility and how the study implementation will play out, if such a larger main study is undertaken. This paper illustrates the relationship between pilot study sample size and the performance study design of main studies. We present two sim…
▽ More
Pilot studies are often the first step of experimental research. It is usually on a smaller scale and the results can inform intervention development, study feasibility and how the study implementation will play out, if such a larger main study is undertaken. This paper illustrates the relationship between pilot study sample size and the performance study design of main studies. We present two simple sample size calculation methods to ensure adequate study planning for main studies. We use numerical examples and simulations to demonstrate the use and performance of proposed methods. Practical heuristic guidelines are provided based on the results.
△ Less
Submitted 12 May, 2021;
originally announced May 2021.
-
Heavy-tailed distribution for combining dependent $p$-values with asymptotic robustness
Authors:
Yusi Fang,
George C. Tseng,
Chung Chang
Abstract:
The issue of combining individual $p$-values to aggregate multiple small effects is prevalent in many scientific investigations and is a long-standing statistical topic. Many classical methods are designed for combining independent and frequent signals in a traditional meta-analysis sense using the sum of transformed $p$-values with the transformation of light-tailed distributions, in which Fisher…
▽ More
The issue of combining individual $p$-values to aggregate multiple small effects is prevalent in many scientific investigations and is a long-standing statistical topic. Many classical methods are designed for combining independent and frequent signals in a traditional meta-analysis sense using the sum of transformed $p$-values with the transformation of light-tailed distributions, in which Fisher's method and Stouffer's method are the most well-known. Since the early 2000, advances in big data promoted methods to aggregate independent, sparse and weak signals, such as the renowned higher criticism and Berk-Jones tests. Recently, Liu and Xie(2020) and Wilson(2019) independently proposed Cauchy and harmonic mean combination tests to robustly combine $p$-values under "arbitrary" dependency structure, where a notable application is to combine $p$-values from a set of often correlated SNPs in genome-wide association studies. The proposed tests are the transformation of heavy-tailed distributions for improved power with the sparse signal. It calls for a natural question to investigate heavy-tailed distribution transformation, to understand the connection among existing methods, and to explore the conditions for a method to possess robustness to dependency. In this paper, we investigate the regularly varying distribution, which is a rich family of heavy-tailed distribution and includes Pareto distribution as a special case. We show that only an equivalent class of Cauchy and harmonic mean tests have sufficient robustness to dependency in a practical sense. We also show an issue caused by large negative penalty in the Cauchy method and propose a simple, yet practical modification. Finally, we present simulations and apply to a neuroticism GWAS application to verify the discovered theoretical insights and provide practical guidance.
△ Less
Submitted 7 September, 2021; v1 submitted 23 March, 2021;
originally announced March 2021.
-
Outcome-Guided Disease Subty** for High-Dimensional Omics Data
Authors:
Peng Liu,
Yusi Fang,
Zhao Ren,
Lu Tang,
George C. Tseng
Abstract:
High-throughput microarray and sequencing technology have been used to identify disease subtypes that could not be observed otherwise by using clinical variables alone. The classical unsupervised clustering strategy concerns primarily the identification of subpopulations that have similar patterns in gene features. However, as the features corresponding to irrelevant confounders (e.g. gender or ag…
▽ More
High-throughput microarray and sequencing technology have been used to identify disease subtypes that could not be observed otherwise by using clinical variables alone. The classical unsupervised clustering strategy concerns primarily the identification of subpopulations that have similar patterns in gene features. However, as the features corresponding to irrelevant confounders (e.g. gender or age) may dominate the clustering process, the resulting clusters may or may not capture clinically meaningful disease subtypes. This gives rise to a fundamental problem: can we find a subty** procedure guided by a pre-specified disease outcome? Existing methods, such as supervised clustering, apply a two-stage approach and depend on an arbitrary number of selected features associated with outcome. In this paper, we propose a unified latent generative model to perform outcome-guided disease subty** constructed from omics data, which improves the resulting subtypes concerning the disease of interest. Feature selection is embedded in a regularization regression. A modified EM algorithm is applied for numerical computation and parameter estimation. The proposed method performs feature selection, latent subtype characterization and outcome prediction simultaneously. To account for possible outliers or violation of mixture Gaussian assumption, we incorporate robust estimation using adaptive Huber or median-truncated loss function. Extensive simulations and an application to complex lung diseases with transcriptomic and clinical data demonstrate the ability of the proposed method to identify clinically relevant disease subtypes and signature genes suitable to explore toward precision medicine.
△ Less
Submitted 21 July, 2020;
originally announced July 2020.
-
Variable screening with multiple studies
Authors:
Tianzhou Ma,
Zhao Ren,
George C. Tseng
Abstract:
Advancement in technology has generated abundant high-dimensional data that allows integration of multiple relevant studies. Due to their huge computational advantage, variable screening methods based on marginal correlation have become promising alternatives to the popular regularization methods for variable selection. However, all these screening methods are limited to single study so far. In th…
▽ More
Advancement in technology has generated abundant high-dimensional data that allows integration of multiple relevant studies. Due to their huge computational advantage, variable screening methods based on marginal correlation have become promising alternatives to the popular regularization methods for variable selection. However, all these screening methods are limited to single study so far. In this paper, we consider a general framework for variable screening with multiple related studies, and further propose a novel two-step screening procedure using a self-normalized estimator for high-dimensional regression analysis in this framework. Compared to the one-step procedure and rank-based sure independence screening (SIS) procedure, our procedure greatly reduces false negative errors while kee** a low false positive rate. Theoretically, we show that our procedure possesses the sure screening property with weaker assumptions on signal strengths and allows the number of features to grow at an exponential rate of the sample size. In addition, we relax the commonly used normality assumption and allow sub-Gaussian distributions. Simulations and a real transcriptomic application illustrate the advantage of our method as compared to the rank-based SIS method.
△ Less
Submitted 10 October, 2017;
originally announced October 2017.
-
Imputation of truncated p-values for meta-analysis methods and its genomic application
Authors:
Shaowu Tang,
Ying Ding,
Etienne Sibille,
Jeffrey S. Mogil,
William R. Lariviere,
George C. Tseng
Abstract:
Microarray analysis to monitor expression activities in thousands of genes simultaneously has become routine in biomedical research during the past decade. A tremendous amount of expression profiles are generated and stored in the public domain and information integration by meta-analysis to detect differentially expressed (DE) genes has become popular to obtain increased statistical power and val…
▽ More
Microarray analysis to monitor expression activities in thousands of genes simultaneously has become routine in biomedical research during the past decade. A tremendous amount of expression profiles are generated and stored in the public domain and information integration by meta-analysis to detect differentially expressed (DE) genes has become popular to obtain increased statistical power and validated findings. Methods that aggregate transformed $p$-value evidence have been widely used in genomic settings, among which Fisher's and Stouffer's methods are the most popular ones. In practice, raw data and $p$-values of DE evidence are often not available in genomic studies that are to be combined. Instead, only the detected DE gene lists under a certain $p$-value threshold (e.g., DE genes with $p$-value${}<0.001$) are reported in journal publications. The truncated $p$-value information makes the aforementioned meta-analysis methods inapplicable and researchers are forced to apply a less efficient vote counting method or naïvely drop the studies with incomplete information. The purpose of this paper is to develop effective meta-analysis methods for such situations with partially censored $p$-values. We developed and compared three imputation methods - mean imputation, single random imputation and multiple imputation - for a general class of evidence aggregation methods of which Fisher's and Stouffer's methods are special examples. The null distribution of each method was analytically derived and subsequent inference and genomic analysis frameworks were established. Simulations were performed to investigate the type I error, power and the control of false discovery rate (FDR) for (correlated) gene expression data. The proposed methods were applied to several genomic applications in colorectal cancer, pain and liquid association analysis of major depressive disorder (MDD). The results showed that imputation methods outperformed existing naïve approaches. Mean imputation and multiple imputation methods performed the best and are recommended for future applications.
△ Less
Submitted 19 January, 2015;
originally announced January 2015.
-
Hypothesis setting and order statistic for robust genomic meta-analysis
Authors:
Chi Song,
George C. Tseng
Abstract:
Meta-analysis techniques have been widely developed and applied in genomic applications, especially for combining multiple transcriptomic studies. In this paper we propose an order statistic of $p$-values ($r$th ordered $p$-value, rOP) across combined studies as the test statistic. We illustrate different hypothesis settings that detect gene markers differentially expressed (DE) 'in all studies,"…
▽ More
Meta-analysis techniques have been widely developed and applied in genomic applications, especially for combining multiple transcriptomic studies. In this paper we propose an order statistic of $p$-values ($r$th ordered $p$-value, rOP) across combined studies as the test statistic. We illustrate different hypothesis settings that detect gene markers differentially expressed (DE) 'in all studies," "in the majority of studies"' or "in one or more studies," and specify rOP as a suitable method for detecting DE genes "in the majority of studies." We develop methods to estimate the parameter $r$ in rOP for real applications. Statistical properties such as its asymptotic behavior and a one-sided testing correction for detecting markers of concordant expression changes are explored. Power calculation and simulation show better performance of rOP compared to classical Fisher's method, Stouffer's method, minimum $p$-value method and maximum $p$-value method under the focused hypothesis setting. Theoretically, rOP is found connected to the naïve vote counting method and can be viewed as a generalized form of vote counting with better statistical properties. The method is applied to three microarray meta-analysis examples including major depressive disorder, brain cancer and diabetes. The results demonstrate rOP as a more generalizable, robust and sensitive statistical framework to detect disease-related markers.
△ Less
Submitted 31 July, 2014;
originally announced July 2014.
-
A New Test for One-Way ANOVA with Functional Data and Application to Ischemic Heart Screening
Authors:
**-Ting Zhang,
Ming-Yen Cheng,
Chi-Jen Tseng,
Hau-Tieng Wu
Abstract:
We propose and study a new global test, namely the $F_{\max}$-test, for the one-way ANOVA problem in functional data analysis. The test statistic is taken as the maximum value of the usual pointwise $F$-test statistics over the interval the functional responses are observed. A nonparametric bootstrap method is employed to approximate the null distribution of the test statistic and to obtain an est…
▽ More
We propose and study a new global test, namely the $F_{\max}$-test, for the one-way ANOVA problem in functional data analysis. The test statistic is taken as the maximum value of the usual pointwise $F$-test statistics over the interval the functional responses are observed. A nonparametric bootstrap method is employed to approximate the null distribution of the test statistic and to obtain an estimated critical value for the test. The asymptotic random expression of the test statistic is derived and the asymptotic power is studied. In particular, under mild conditions, the $F_{\max}$-test asymptotically has the correct level and is root-$n$ consistent in detecting local alternatives. Via some simulation studies, it is found that in terms of both level accuracy and power, the $F_{\max}$-test outperforms the Globalized Pointwise F (GPF) test of \cite{Zhang_Liang:2013} when the functional data are highly or moderately correlated, and its performance is comparable with the latter otherwise. An application to an ischemic heart real dataset suggests that, after proper manipulation, resting electrocardiogram (ECG) signals can be used as an effective tool in clinical ischemic heart screening, without the need of further stress tests as in the current standard procedure.
△ Less
Submitted 27 September, 2013;
originally announced September 2013.
-
An adaptively weighted statistic for detecting differential gene expression when combining multiple transcriptomic studies
Authors:
Jia Li,
George C. Tseng
Abstract:
Global expression analyses using microarray technologies are becoming more common in genomic research, therefore, new statistical challenges associated with combining information from multiple studies must be addressed. In this paper we will describe our proposal for an adaptively weighted (AW) statistic to combine multiple genomic studies for detecting differentially expressed genes. We will also…
▽ More
Global expression analyses using microarray technologies are becoming more common in genomic research, therefore, new statistical challenges associated with combining information from multiple studies must be addressed. In this paper we will describe our proposal for an adaptively weighted (AW) statistic to combine multiple genomic studies for detecting differentially expressed genes. We will also present our results from comparisons of our proposed AW statistic to Fisher's equally weighted (EW), Tippett's minimum p-value (minP) and Pearson's (PR) statistics. Due to the absence of a uniformly powerful test, we used a simplified Gaussian scenario to compare the four methods. Our AW statistic consistently produced the best or near-best power for a range of alternative hypotheses. AW-obtained weights also have the additional advantage of filtering discordant biomarkers and providing natural detected gene categories for further biological investigation. Here we will demonstrate the superior performance of our proposed AW statistic based on a mix of power analyses, simulations and applications using data sets for multi-tissue energy metabolism mouse, multi-lab prostate cancer and lung cancer.
△ Less
Submitted 28 January, 2013; v1 submitted 16 August, 2011;
originally announced August 2011.