Search | arXiv e-print repository

Optimal Transport for Latent Integration with An Application to Heterogeneous Neuronal Activity Data

Authors: Yubai Yuan, Babak Shahbaba, Norbert Fortin, Keiland Cooper, Qing Nie, Annie Qu

Abstract: Detecting dynamic patterns of task-specific responses shared across heterogeneous datasets is an essential and challenging problem in many scientific applications in medical science and neuroscience. In our motivating example of rodent electrophysiological data, identifying the dynamical patterns in neuronal activity associated with ongoing cognitive demands and behavior is key to uncovering the n… ▽ More Detecting dynamic patterns of task-specific responses shared across heterogeneous datasets is an essential and challenging problem in many scientific applications in medical science and neuroscience. In our motivating example of rodent electrophysiological data, identifying the dynamical patterns in neuronal activity associated with ongoing cognitive demands and behavior is key to uncovering the neural mechanisms of memory. One of the greatest challenges in investigating a cross-subject biological process is that the systematic heterogeneity across individuals could significantly undermine the power of existing machine learning methods to identify the underlying biological dynamics. In addition, many technically challenging neurobiological experiments are conducted on only a handful of subjects where rich longitudinal data are available for each subject. The low sample sizes of such experiments could further reduce the power to detect common dynamic patterns among subjects. In this paper, we propose a novel heterogeneous data integration framework based on optimal transport to extract shared patterns in complex biological processes. The key advantages of the proposed method are that it can increase discriminating power in identifying common patterns by reducing heterogeneity unrelated to the signal by aligning the extracted latent spatiotemporal information across subjects. Our approach is effective even with a small number of subjects, and does not require auxiliary matching information for the alignment. In particular, our method can align longitudinal data across heterogeneous subjects in a common latent space to capture the dynamics of shared patterns while utilizing temporal dependency within subjects. △ Less

Submitted 27 June, 2024; originally announced July 2024.

arXiv:2112.03186 [pdf, other]

Statistical implications of relaxing the homogeneous mixing assumption in time series Susceptible-Infectious-Removed models

Authors: Luis D. J. Martinez Lomeli, Michelle N. Ngo, Jon Wakefield, Babak Shahbaba, Vladimir N. Minin

Abstract: Infectious disease epidemiologists routinely fit stochastic epidemic models to time series data to elucidate infectious disease dynamics, evaluate interventions, and forecast epidemic trajectories. To improve computational tractability, many approximate stochastic models have been proposed. In this paper, we focus on one class of such approximations -- time series Susceptible-Infectious-Removed (T… ▽ More Infectious disease epidemiologists routinely fit stochastic epidemic models to time series data to elucidate infectious disease dynamics, evaluate interventions, and forecast epidemic trajectories. To improve computational tractability, many approximate stochastic models have been proposed. In this paper, we focus on one class of such approximations -- time series Susceptible-Infectious-Removed (TSIR) models. Infectious disease modeling often starts with a homogeneous mixing assumption, postulating that the rate of disease transmission is proportional to a product of the numbers of susceptible and infectious individuals. One popular way to relax this assumption proceeds by raising the number of susceptible and/or infectious individuals to some positive powers. We show that when this technique is used within the TSIR models they cannot be interpreted as approximate SIR models, which has important implications for TSIR-based statistical inference. Our simulation study shows that TSIR-based estimates of infection and mixing rates are systematically biased in the presence of non-homogeneous mixing, suggesting that caution is needed when interpreting TSIR model parameter estimates when this class of models relaxes the homogeneous mixing assumption. △ Less

Submitted 6 December, 2021; originally announced December 2021.

Comments: 16 pages of the main text, 6 figures and 3 tables in the main text

arXiv:2004.09065 [pdf, other]

Optimal Experimental Design for Mathematical Models of Hematopoiesis

Authors: Luis Martinez Lomeli, Abdon Iniguez, Babak Shahbaba, John S Lowengrub, Vladimir Minin

Abstract: The hematopoietic system has a highly regulated and complex structure in which cells are organized to successfully create and maintain new blood cells. Feedback regulation is crucial to tightly control this system, but the specific mechanisms by which control is exerted are not completely understood. In this work, we aim to uncover the underlying mechanisms in hematopoiesis by conducting perturbat… ▽ More The hematopoietic system has a highly regulated and complex structure in which cells are organized to successfully create and maintain new blood cells. Feedback regulation is crucial to tightly control this system, but the specific mechanisms by which control is exerted are not completely understood. In this work, we aim to uncover the underlying mechanisms in hematopoiesis by conducting perturbation experiments, where animal subjects are exposed to an external agent in order to observe the system response and evolution. Develo** a proper experimental design for these studies is an extremely challenging task. To address this issue, we have developed a novel Bayesian framework for optimal design of perturbation experiments. We model the numbers of hematopoietic stem and progenitor cells in mice that are exposed to a low dose of radiation. We use a differential equations model that accounts for feedback and feedforward regulation. A significant obstacle is that the experimental data are not longitudinal, rather each data point corresponds to a different animal. This model is embedded in a hierarchical framework with latent variables that capture unobserved cellular population levels. We select the optimum design based on the amount of information gain, measured by the Kullback-Leibler divergence between the probability distributions before and after observing the data. We evaluate our approach using synthetic and experimental data. We show that a proper design can lead to better estimates of model parameters even with relatively few subjects. Additionally, we demonstrate that the model parameters show a wide range of sensitivities to design options. Our method should allow scientists to find the optimal design by focusing on their specific parameters of interest and provide insight to hematopoiesis. Our approach can be extended to more complex models where latent components are used. △ Less

Submitted 30 June, 2020; v1 submitted 20 April, 2020; originally announced April 2020.

arXiv:1306.6103 [pdf, other]

doi 10.1162/NECO_a_00631

A Semiparametric Bayesian Model for Detecting Synchrony Among Multiple Neurons

Authors: Babak Shahbaba, Bo Zhou, Shiwei Lan, Hernando Ombao, David Moorman, Sam Behseta

Abstract: We propose a scalable semiparametric Bayesian model to capture dependencies among multiple neurons by detecting their co-firing (possibly with some lag time) patterns over time. After discretizing time so there is at most one spike at each interval, the resulting sequence of 1's (spike) and 0's (silence) for each neuron is modeled using the logistic function of a continuous latent variable with a… ▽ More We propose a scalable semiparametric Bayesian model to capture dependencies among multiple neurons by detecting their co-firing (possibly with some lag time) patterns over time. After discretizing time so there is at most one spike at each interval, the resulting sequence of 1's (spike) and 0's (silence) for each neuron is modeled using the logistic function of a continuous latent variable with a Gaussian process prior. For multiple neurons, the corresponding marginal distributions are coupled to their joint probability distribution using a parametric copula model. The advantages of our approach are as follows: the nonparametric component (i.e., the Gaussian process model) provides a flexible framework for modeling the underlying firing rates; the parametric component (i.e., the copula model) allows us to make inference regarding both contemporaneous and lagged relationships among neurons; using the copula model, we construct multivariate probabilistic models by separating the modeling of univariate marginal distributions from the modeling of dependence structure among variables; our method is easy to implement using a computationally efficient sampling algorithm that can be easily extended to high dimensional problems. Using simulated data, we show that our approach could correctly capture temporal dependencies in firing rates and identify synchronous neurons. We also apply our model to spike train data obtained from prefrontal cortical areas in rat's brain. △ Less

Submitted 3 March, 2014; v1 submitted 25 June, 2013; originally announced June 2013.

Journal ref: Neural Computation, September 2014, Vol. 26, No. 9, Pages 2025-2051

arXiv:1006.5170 [pdf, other]

Bayesian Gene Set Analysis

Authors: Babak Shahbaba, Robert Tibshirani, Catherine M. Shachaf, Sylvia K. Plevritis

Abstract: Gene expression microarray technologies provide the simultaneous measurements of a large number of genes. Typical analyses of such data focus on the individual genes, but recent work has demonstrated that evaluating changes in expression across predefined sets of genes often increases statistical power and produces more robust results. We introduce a new methodology for identifying gene sets that… ▽ More Gene expression microarray technologies provide the simultaneous measurements of a large number of genes. Typical analyses of such data focus on the individual genes, but recent work has demonstrated that evaluating changes in expression across predefined sets of genes often increases statistical power and produces more robust results. We introduce a new methodology for identifying gene sets that are differentially expressed under varying experimental conditions. Our approach uses a hierarchical Bayesian framework where a hyperparameter measures the significance of each gene set. Using simulated data, we compare our proposed method to alternative approaches, such as Gene Set Enrichment Analysis (GSEA) and Gene Set Analysis (GSA). Our approach provides the best overall performance. We also discuss the application of our method to experimental data based on p53 mutation status. △ Less

Submitted 26 June, 2010; originally announced June 2010.

arXiv:1003.2390 [pdf, other]

Bayesian Nonparametric Variable Selection as an Exploratory Tool for Finding Genes that Matter

Authors: Babak Shahbaba

Abstract: High-throughput scientific studies involving no clear a'priori hypothesis are common. For example, a large-scale genomic study of a disease may examine thousands of genes without hypothesizing that any specific gene is responsible for the disease. In these studies, the objective is to explore a large number of possible factors (e.g. genes) in order to identify a small number that will be considere… ▽ More High-throughput scientific studies involving no clear a'priori hypothesis are common. For example, a large-scale genomic study of a disease may examine thousands of genes without hypothesizing that any specific gene is responsible for the disease. In these studies, the objective is to explore a large number of possible factors (e.g. genes) in order to identify a small number that will be considered in follow-up studies that tend to be more thorough and on smaller scales. For large-scale studies, we propose a nonparametric Bayesian approach based on random partition models. Our model thus divides the set of candidate factors into several subgroups according to their degrees of relevance, or potential effect, in relation to the outcome of interest. The model allows for a latent rank to be assigned to each factor according to the overall potential importance of its corresponding group. The posterior expectation or mode of these ranks is used to set up a threshold for selecting potentially relevant factors. Using simulated data, we demonstrate that our approach could be quite effective in finding relevant genes compared to several alternative methods. We apply our model to two large-scale studies. The first study involves transcriptome analysis of infection by human cytomegalovirus (HCMV). The objective of the second study is to identify differentially expressed genes between two types of leukemia. △ Less

Submitted 29 February, 2012; v1 submitted 11 March, 2010; originally announced March 2010.

arXiv:math/0703292 [pdf, ps, other]

Nonlinear Models Using Dirichlet Process Mixtures

Authors: Babak Shahbaba, Radford M. Neal

Abstract: We introduce a new nonlinear model for classification, in which we model the joint distribution of response variable, y, and covariates, x, non-parametrically using Dirichlet process mixtures. We keep the relationship between y and x linear within each component of the mixture. The overall relationship becomes nonlinear if the mixture contains more than one component. We use simulated data to co… ▽ More We introduce a new nonlinear model for classification, in which we model the joint distribution of response variable, y, and covariates, x, non-parametrically using Dirichlet process mixtures. We keep the relationship between y and x linear within each component of the mixture. The overall relationship becomes nonlinear if the mixture contains more than one component. We use simulated data to compare the performance of this new approach to a simple multinomial logit (MNL) model, an MNL model with quadratic terms, and a decision tree model. We also evaluate our approach on a protein fold classification problem, and find that our model provides substantial improvement over previous methods, which were based on Neural Networks (NN) and Support Vector Machines (SVM). Folding classes of protein have a hierarchical structure. We extend our method to classification problems where a class hierarchy is available. We find that using the prior information regarding the hierarchical structure of protein folds can result in higher predictive accuracy. △ Less

Submitted 10 March, 2007; originally announced March 2007.

MSC Class: 62H30

arXiv:q-bio/0605015 [pdf, ps, other]

Gene Function Classification Using Bayesian Models with Hierarchy-Based Priors

Authors: Babak Shahbaba, Radford M. Neal

Abstract: We investigate the application of hierarchical classification schemes to the annotation of gene function based on several characteristics of protein sequences including phylogenic descriptors, sequence based attributes, and predicted secondary structure. We discuss three Bayesian models and compare their performance in terms of predictive accuracy. These models are the ordinary multinomial logit… ▽ More We investigate the application of hierarchical classification schemes to the annotation of gene function based on several characteristics of protein sequences including phylogenic descriptors, sequence based attributes, and predicted secondary structure. We discuss three Bayesian models and compare their performance in terms of predictive accuracy. These models are the ordinary multinomial logit (MNL) model, a hierarchical model based on a set of nested MNL models, and a MNL model with a prior that introduces correlations between the parameters for classes that are nearby in the hierarchy. We also provide a new scheme for combining different sources of information. We use these models to predict the functional class of Open Reading Frames (ORFs) from the E. coli genome. The results from all three models show substantial improvement over previous methods, which were based on the C5 algorithm. The MNL model using a prior based on the hierarchy outperforms both the non-hierarchical MNL model and the nested MNL model. In contrast to previous attempts at combining these sources of information, our approach results in a higher accuracy rate when compared to models that use each data source alone. Together, these results show that gene function can be predicted with higher accuracy than previously achieved, using Bayesian models that incorporate suitable prior information. △ Less

Submitted 10 May, 2006; originally announced May 2006.

Showing 1–8 of 8 results for author: Shahbaba, B