A Cross Entropy test allows quantitative statistical comparison of t-SNE and UMAP representations
Authors:
Carlos P. Roca1,
Oliver T. Burton,
Julika Neumann,
Samar Tareen,
Carly E. Whyte,
Stéphanie Humblet-Baron,
Adrian Liston
Abstract:
The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of highly complex single cell datasets. Despite the ubiquity of these approaches and the clear need for quantitative comparison of single cell datasets, t-SNE and UMAP…
▽ More
The advent of high dimensional single cell data in the biomedical sciences has necessitated the development of dimensionality-reduction tools. t-SNE and UMAP are the two most frequently used approaches, allowing clear visualisation of highly complex single cell datasets. Despite the ubiquity of these approaches and the clear need for quantitative comparison of single cell datasets, t-SNE and UMAP have largely remained data visualisation tools due to the lack of robust statistical approaches available. Here, we have derived a statistical test for evaluating the difference between dimensionality-reduced datasets, using the Kolmogorov-Smirnov test on the distributions of cross entropy of single cells within each dataset. As the approach uses the interrelationship of single cells for comparison, the resulting statistic is robust and capable of distinguishing between true biological variation and rotational symmetry generation during dimensionality reduction. Further, the test provides a valid distance between single cell datasets, allowing the organisation of multiple samples into a dendrogram for quantitative comparison of complex datasets. These results demonstrate the largely untapped potential of dimensionality-reduction tools for biomedical data analysis beyond visualisation.
△ Less
Submitted 8 December, 2021;
originally announced December 2021.