-
A picture guide to cancer progression and monotonic accumulation models: evolutionary assumptions, plausible interpretations, and alternative uses
Authors:
Ramon Diaz-Uriarte,
Iain G. Johnston
Abstract:
Cancer progression and monotonic accumulation models were developed to discover dependencies in the irreversible acquisition of binary traits from cross-sectional data. They have been used in computational oncology and virology but also in widely different problems such as malaria progression. These methods have been applied to predict future states of the system, identify routes of feature acquis…
▽ More
Cancer progression and monotonic accumulation models were developed to discover dependencies in the irreversible acquisition of binary traits from cross-sectional data. They have been used in computational oncology and virology but also in widely different problems such as malaria progression. These methods have been applied to predict future states of the system, identify routes of feature acquisition, and improve patient stratification, and they hold promise for evolutionary-based treatments. New methods continue to be developed.
But these methods have shortcomings, which are yet to be systematically critiqued, regarding key evolutionary assumptions and interpretations. After an overview of the available methods, we focus on why inferences might not be about the processes we intend. Using fitness landscapes, we highlight difficulties that arise from bulk sequencing and reciprocal sign epistasis, from conflating lines of descent, path of the maximum, and mutational profiles, and from ambiguous use of the idea of exclusivity. We examine how the previous concerns change when bulk sequencing is explicitly considered, and underline opportunities for addressing dependencies due to frequency-dependent selection. This review identifies major standing issues, and should encourage the use of these methods in other areas with a better alignment between entities and model assumptions.
△ Less
Submitted 30 June, 2024; v1 submitted 11 December, 2023;
originally announced December 2023.
-
Global epistasis on fitness landscapes
Authors:
Juan Diaz-Colunga,
Abigail Skwara,
Karna Gowda,
Ramon Diaz-Uriarte,
Mikhail Tikhonov,
Djordje Bajic,
Alvaro Sanchez
Abstract:
Epistatic interactions between mutations add substantial complexity to adaptive landscapes, and are often thought of as detrimental to our ability to predict evolution. Yet, patterns of global epistasis, in which the fitness effect of a mutation is well-predicted by the fitness of its genetic background, may actually be of help in our efforts to reconstruct fitness landscapes and infer adaptive tr…
▽ More
Epistatic interactions between mutations add substantial complexity to adaptive landscapes, and are often thought of as detrimental to our ability to predict evolution. Yet, patterns of global epistasis, in which the fitness effect of a mutation is well-predicted by the fitness of its genetic background, may actually be of help in our efforts to reconstruct fitness landscapes and infer adaptive trajectories. Microscopic interactions between mutations, or inherent nonlinearities in the fitness landscape, may cause global epistasis patterns to emerge. In this brief review, we provide a succinct overview of recent work about global epistasis, with an emphasis on building intuition about why it is often observed. To this end, we reconcile simple geometric reasoning with recent mathematical analyses, using these to explain why different mutations in an empirical landscape may exhibit different global epistasis patterns - ranging from diminishing to increasing returns. Finally, we highlight open questions and research directions.
△ Less
Submitted 7 October, 2022;
originally announced October 2022.
-
From genotypes to organisms: State-of-the-art and perspectives of a cornerstone in evolutionary dynamics
Authors:
Susanna Manrubia,
José A. Cuesta,
Jacobo Aguirre,
Sebastian E. Ahnert,
Lee Altenberg,
Alejandro V. Cano,
Pablo Catalán,
Ramon Diaz-Uriarte,
Santiago F. Elena,
Juan Antonio García-Martín,
Paulien Hogeweg,
Bhavin S. Khatri,
Joachim Krug,
Ard A. Louis,
Nora S. Martin,
Joshua L. Payne,
Matthew J. Tarnowski,
Marcel Weiß
Abstract:
Understanding how genotypes map onto phenotypes, fitness, and eventually organisms is arguably the next major missing piece in a fully predictive theory of evolution. We refer to this generally as the problem of the genotype-phenotype map. Though we are still far from achieving a complete picture of these relationships, our current understanding of simpler questions, such as the structure induced…
▽ More
Understanding how genotypes map onto phenotypes, fitness, and eventually organisms is arguably the next major missing piece in a fully predictive theory of evolution. We refer to this generally as the problem of the genotype-phenotype map. Though we are still far from achieving a complete picture of these relationships, our current understanding of simpler questions, such as the structure induced in the space of genotypes by sequences mapped to molecular structures, has revealed important facts that deeply affect the dynamical description of evolutionary processes. Empirical evidence supporting the fundamental relevance of features such as phenotypic bias is mounting as well, while the synthesis of conceptual and experimental progress leads to questioning current assumptions on the nature of evolutionary dynamics-cancer progression models or synthetic biology approaches being notable examples. This work delves into a critical and constructive attitude in our current knowledge of how genotypes map onto molecular phenotypes and organismal functions, and discusses theoretical and empirical avenues to broaden and improve this comprehension. As a final goal, this community should aim at deriving an updated picture of evolutionary processes soundly relying on the structural properties of genotype spaces, as revealed by modern techniques of molecular and functional analysis.
△ Less
Submitted 17 March, 2021; v1 submitted 2 February, 2020;
originally announced February 2020.
-
Asterias: a parallelized web-based suite for the analysis of expression and aCGH data
Authors:
Andreu Alibes,
Edward R. Morrissey,
Andres Canada,
Oscar M. Rueda,
David Casado,
Patricio Yankilevich,
Ramon Diaz-Uriarte
Abstract:
Asterias (\url{http://www.asterias.info}) is an integrated collection of freely-accessible web tools for the analysis of gene expression and aCGH data. Most of the tools use parallel computing (via MPI). Most of our applications allow the user to obtain additional information for user-selected genes by using clickable links in tables and/or figures. Our tools include: normalization of expression…
▽ More
Asterias (\url{http://www.asterias.info}) is an integrated collection of freely-accessible web tools for the analysis of gene expression and aCGH data. Most of the tools use parallel computing (via MPI). Most of our applications allow the user to obtain additional information for user-selected genes by using clickable links in tables and/or figures. Our tools include: normalization of expression and aCGH data; converting between different types of gene/clone and protein identifiers; filtering and imputation; finding differentially expressed genes related to patient class and survival data; searching for models of class prediction; using random forests to search for minimal models for class prediction or for large subsets of genes with predictive capacity; searching for molecular signatures and predictive genes with survival data; detecting regions of genomic DNA gain or loss. The capability to send results between different applications, access to additional functional information, and parallelized computation make our suite unique and exploit features only available to web-based applications.
△ Less
Submitted 22 October, 2006;
originally announced October 2006.
-
Variable selection from random forests: application to gene expression data
Authors:
Ramon Diaz-Uriarte,
Sara Alvarez de Andres
Abstract:
Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations, and returns measures of variable importance. Thus, it is important to understand the performance of random forest with microarray data and its use for gen…
▽ More
Random forest is a classification algorithm well suited for microarray data: it shows excellent performance even when most predictive variables are noise, can be used when the number of variables is much larger than the number of observations, and returns measures of variable importance. Thus, it is important to understand the performance of random forest with microarray data and its use for gene selection.
We first show the effects of changes in parameters of random forest on the prediction error. Then we present an approach for gene selection that uses measures of variable importance and error rate, and is targeted towards the selection of small sets of genes. Using simulated and real microarray data, we show that the gene selection procedure yields small sets of genes while preserving predictive accuracy.
Availability: All code is available as an R package, varSelRF, from CRAN, http://cran.r-project.org/src/contrib/PACKAGES.html, or from the supplementary material page.
Supplementary information: http://ligarto.org/rdiaz/Papers/rfVS/randomForestVarSel.html
△ Less
Submitted 22 June, 2005; v1 submitted 16 March, 2005;
originally announced March 2005.
-
Molecular Signatures from Gene Expression Data
Authors:
Ramon Diaz-Uriarte
Abstract:
Motivation: ``Molecular signatures'' or ``gene-expression signatures'' are used to predict patients' characteristics using data from coexpressed genes. Signatures can enhance understanding about biological mechanisms and have diagnostic use. However, available methods to search for signatures fail to address key requirements of signatures, especially the discovery of sets of tightly coexpressed…
▽ More
Motivation: ``Molecular signatures'' or ``gene-expression signatures'' are used to predict patients' characteristics using data from coexpressed genes. Signatures can enhance understanding about biological mechanisms and have diagnostic use. However, available methods to search for signatures fail to address key requirements of signatures, especially the discovery of sets of tightly coexpressed genes. Results: After suggesting an operational definition of signature, we develop a method that fulfills these requirements, returning sets of tightly coexpressed genes with good predictive performance. This method can also identify when the data are inconsistent with the hypothesis of a few, stable, easily interpretable sets of coexpressed genes. Identification of molecular signatures in some widely used data sets is questionable under this simple model, which emphasizes the needed for further work on the operationalization of the biological model and the assessment of the stability of putative signatures. Availability: The code (R with C++) is available from http://www.ligarto.org/rdiaz/Software/Software.html under the GNU GPL.
△ Less
Submitted 8 October, 2004; v1 submitted 30 January, 2004;
originally announced January 2004.