-
BIISQ: Bayesian nonparametric discovery of Isoforms and Individual Specific Quantification
Authors:
Derek Aguiar,
Li-Fang Cheng,
Bianca Dumitrascu,
Fantine Mordelet,
Athma A Pai,
Barbara E Engelhardt
Abstract:
Most human protein-coding genes can be transcribed into multiple possible distinct mRNA isoforms. These alternative splicing patterns encourage molecular diversity and dysregulation of isoform expression plays an important role in disease etiology. However, isoforms are difficult to characterize from short-read RNA-seq data because they share identical subsequences and exist in tissue- and sample-…
▽ More
Most human protein-coding genes can be transcribed into multiple possible distinct mRNA isoforms. These alternative splicing patterns encourage molecular diversity and dysregulation of isoform expression plays an important role in disease etiology. However, isoforms are difficult to characterize from short-read RNA-seq data because they share identical subsequences and exist in tissue- and sample-specific frequencies. Here, we develop BIISQ, a Bayesian nonparametric model to discover Isoforms and Individual Specific Quantification from RNA-seq data. BIISQ does not require known isoform reference sequences but instead estimates isoform composition directly with an isoform catalog shared across samples. We develop a stochastic variational inference approach for efficient and robust posterior inference and demonstrate superior precision and recall for short read RNA-seq simulations and simulated short read data from PacBio long read sequencing when compared to state-of-the-art isoform reconstruction methods. BIISQ achieves the most significant gains for longer (in terms of exons) isoforms and isoforms that are lowly expressed (over 500% more transcripts correctly inferred at low coverage in simulations). Finally, we estimate isoforms in the GEUVADIS RNA-seq data, identify genetic variants that regulate transcript ratios, and demonstrate variant enrichment in functional elements related to mRNA splicing regulation.
△ Less
Submitted 23 March, 2017;
originally announced March 2017.
-
TIGRESS: Trustful Inference of Gene REgulation using Stability Selection
Authors:
Anne-Claire Haury,
Fantine Mordelet,
Paola Vera-Licona,
Jean-Philippe Vert
Abstract:
Inferring the structure of gene regulatory networks (GRN) from gene expression data has many applications, from the elucidation of complex biological processes to the identification of potential drug targets. It is however a notoriously difficult problem, for which the many existing methods reach limited accuracy. In this paper, we formulate GRN inference as a sparse regression problem and investi…
▽ More
Inferring the structure of gene regulatory networks (GRN) from gene expression data has many applications, from the elucidation of complex biological processes to the identification of potential drug targets. It is however a notoriously difficult problem, for which the many existing methods reach limited accuracy. In this paper, we formulate GRN inference as a sparse regression problem and investigate the performance of a popular feature selection method, least angle regression (LARS) combined with stability selection. We introduce a novel, robust and accurate scoring technique for stability selection, which improves the performance of feature selection with LARS. The resulting method, which we call TIGRESS (Trustful Inference of Gene REgulation using Stability Selection), was ranked among the top methods in the DREAM5 gene network reconstruction challenge. We investigate in depth the influence of the various parameters of the method and show that a fine parameter tuning can lead to significant improvements and state-of-the-art performance for GRN inference. TIGRESS reaches state-of-the-art performance on benchmark data. This study confirms the potential of feature selection techniques for GRN inference. Code and data are available on http://cbio.ensmp.fr/~ahaury. Running TIGRESS online is possible on GenePattern: http://www.broadinstitute.org/cancer/software/genepattern/.
△ Less
Submitted 6 May, 2012;
originally announced May 2012.
-
ProDiGe: PRioritization Of Disease Genes with multitask machine learning from positive and unlabeled examples
Authors:
Fantine Mordelet,
Jean-Philippe Vert
Abstract:
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-consuming and expensive. Efficient computational methods are therefore needed to pr…
▽ More
Elucidating the genetic basis of human diseases is a central goal of genetics and molecular biology. While traditional linkage analysis and modern high-throughput techniques often provide long lists of tens or hundreds of disease gene candidates, the identification of disease genes among the candidates remains time-consuming and expensive. Efficient computational methods are therefore needed to prioritize genes within the list of candidates, by exploiting the wealth of information available about the genes in various databases. Here we propose ProDiGe, a novel algorithm for Prioritization of Disease Genes. ProDiGe implements a novel machine learning strategy based on learning from positive and unlabeled examples, which allows to integrate various sources of information about the genes, to share information about known disease genes across diseases, and to perform genome-wide searches for new disease genes. Experiments on real data show that ProDiGe outperforms state-of-the-art methods for the prioritization of genes in human diseases.
△ Less
Submitted 1 June, 2011;
originally announced June 2011.
-
A bagging SVM to learn from positive and unlabeled examples
Authors:
Fantine Mordelet,
Jean-Philippe Vert
Abstract:
We consider the problem of learning a binary classifier from a training set of positive and unlabeled examples, both in the inductive and in the transductive setting. This problem, often referred to as \emph{PU learning}, differs from the standard supervised classification problem by the lack of negative examples in the training set. It corresponds to an ubiquitous situation in many applications s…
▽ More
We consider the problem of learning a binary classifier from a training set of positive and unlabeled examples, both in the inductive and in the transductive setting. This problem, often referred to as \emph{PU learning}, differs from the standard supervised classification problem by the lack of negative examples in the training set. It corresponds to an ubiquitous situation in many applications such as information retrieval or gene ranking, when we have identified a set of data of interest sharing a particular property, and we wish to automatically retrieve additional data sharing the same property among a large and easily available pool of unlabeled data. We propose a conceptually simple method, akin to bagging, to approach both inductive and transductive PU learning problems, by converting them into series of supervised binary classification problems discriminating the known positive examples from random subsamples of the unlabeled set. We empirically demonstrate the relevance of the method on simulated and real data, where it performs at least as well as existing methods while being faster.
△ Less
Submitted 5 October, 2010;
originally announced October 2010.
-
SIRENE: Supervised Inference of Regulatory Networks
Authors:
Fantine Mordelet,
Jean-Philippe Vert
Abstract:
Living cells are the product of gene expression programs that involve the regulated transcription of thousands of genes. The elucidation of transcriptional regulatory networks in thus needed to understand the cell's working mechanism, and can for example be useful for the discovery of novel therapeutic targets. Although several methods have been proposed to infer gene regulatory networks from ge…
▽ More
Living cells are the product of gene expression programs that involve the regulated transcription of thousands of genes. The elucidation of transcriptional regulatory networks in thus needed to understand the cell's working mechanism, and can for example be useful for the discovery of novel therapeutic targets. Although several methods have been proposed to infer gene regulatory networks from gene expression data, a recent comparison on a large-scale benchmark experiment revealed that most current methods only predict a limited number of known regulations at a reasonable precision level. We propose SIRENE, a new method for the inference of gene regulatory networks from a compendium of expression data. The method decomposes the problem of gene regulatory network inference into a large number of local binary classification problems, that focus on separating target genes from non-targets for each TF. SIRENE is thus conceptually simple and computationally efficient. We test it on a benchmark experiment aimed at predicting regulations in E. coli, and show that it retrieves of the order of 6 times more known regulations than other state-of-the-art inference methods.
△ Less
Submitted 27 February, 2008;
originally announced February 2008.