-
Extraction and integration of genetic networks from short-profile omic datasets
Authors:
Jacopo Iacovacci,
Alina Peluso,
Timothy Ebbels,
Markus Ralser,
Robert Charles Glen
Abstract:
Mass-spectrometry technologies are widely used in the fields of ionomics and metabolomics to simultaneously profile at the genome scale intracellular concentrations of e.g. amino acids or elements. Short profiles of molecular or sub-molecular features are intrinsically non-Gaussian and may reveal patterns of correlations that reflect the system nature of the cell biochemistry and biology. Here we…
▽ More
Mass-spectrometry technologies are widely used in the fields of ionomics and metabolomics to simultaneously profile at the genome scale intracellular concentrations of e.g. amino acids or elements. Short profiles of molecular or sub-molecular features are intrinsically non-Gaussian and may reveal patterns of correlations that reflect the system nature of the cell biochemistry and biology. Here we introduce two profile similarity measures that enforce information from the empirical covariance matrix of the data, the Mahalanobis cosine and the hybrid-Mahalanobis cosine. We evaluate the performance of these similarity measures in the task of inferring and integrating genetic networks from omics data by analysing experimental datasets derived from the ionome and the metabolome of the model organism S. cerevisiae, and several large curated databases of genetic annotations. The proposed covariance-based similarity measures can in general recover known and predicted associations between genes better than the commonly used Pearson's correlation and the standard cosine similarity. The choice of which of the two measures to recommend depends upon whether the focus is on extracting genetic associations at a global or local genetic network scale.
△ Less
Submitted 10 November, 2020; v1 submitted 30 October, 2019;
originally announced October 2019.
-
Validating the Validation: Reanalyzing a large-scale comparison of Deep Learning and Machine Learning models for bioactivity prediction
Authors:
Matthew C. Robinson,
Robert C. Glen,
Alpha A. Lee
Abstract:
Machine learning methods may have the potential to significantly accelerate drug discovery. However, the increasing rate of new methodological approaches being published in the literature raises the fundamental question of how models should be benchmarked and validated. We reanalyze the data generated by a recently published large-scale comparison of machine learning models for bioactivity predict…
▽ More
Machine learning methods may have the potential to significantly accelerate drug discovery. However, the increasing rate of new methodological approaches being published in the literature raises the fundamental question of how models should be benchmarked and validated. We reanalyze the data generated by a recently published large-scale comparison of machine learning models for bioactivity prediction and arrive at a somewhat different conclusion. We show that the performance of support vector machines is competitive with that of deep learning methods. Additionally, using a series of numerical experiments, we question the relevance of area under the receiver operating characteristic curve as a metric in virtual screening, and instead suggest that area under the precision-recall curve should be used in conjunction with the receiver operating characteristic. Our numerical experiments also highlight challenges in estimating the uncertainty in model performance via scaffold-split nested cross validation.
△ Less
Submitted 9 June, 2019; v1 submitted 28 May, 2019;
originally announced May 2019.
-
Target Fishing: A Single-Label or Multi-Label Problem?
Authors:
Avid M. Afzal,
Hamse Y. Mussa,
Richard E. Turner,
Andreas Bender,
Robert C. Glen
Abstract:
According to Cobanoglu et al and Murphy, it is now widely acknowledged that the single target paradigm (one protein or target, one disease, one drug) that has been the dominant premise in drug development in the recent past is untenable. More often than not, a drug-like compound (ligand) can be promiscuous - that is, it can interact with more than one target protein. In recent years, in in silico…
▽ More
According to Cobanoglu et al and Murphy, it is now widely acknowledged that the single target paradigm (one protein or target, one disease, one drug) that has been the dominant premise in drug development in the recent past is untenable. More often than not, a drug-like compound (ligand) can be promiscuous - that is, it can interact with more than one target protein. In recent years, in in silico target prediction methods the promiscuity issue has been approached computationally in different ways. In this study we confine attention to the so-called ligand-based target prediction machine learning approaches, commonly referred to as target-fishing. With a few exceptions, the target-fishing approaches that are currently ubiquitous in cheminformatics literature can be essentially viewed as single-label multi-classification schemes; these approaches inherently bank on the single target paradigm assumption that a ligand can home in on one specific target. In order to address the ligand promiscuity issue, one might be able to cast target-fishing as a multi-label multi-class classification problem. For illustrative and comparison purposes, single-label and multi-label Naive Bayes classification models (denoted here by SMM and MMM, respectively) for target-fishing were implemented. The models were constructed and tested on 65,587 compounds and 308 targets retrieved from the ChEMBL17 database. SMM and MMM performed differently: for 16,344 test compounds, the MMM model returned recall and precision values of 0.8058 and 0.6622, respectively; the corresponding recall and precision values yielded by the SMM model were 0.7805 and 0.7596, respectively. However, at a significance level of 0.05 and one degree of freedom McNemar test performed on the target prediction results returned by SMM and MMM for the 16,344 test ligands gave a chi-squared value of 15.656, in favour of the MMM approach.
△ Less
Submitted 23 November, 2014;
originally announced November 2014.