-
Random Expert Sampling for Deep Learning Segmentation of Acute Ischemic Stroke on Non-contrast CT
Authors:
Sophie Ostmeier,
Brian Axelrod,
Benjamin Pulli,
Benjamin F. J. Verhaaren,
Abdelkader Mahammedi,
Yongkai Liu,
Christian Federau,
Greg Zaharchuk,
Jeremy J. Heit
Abstract:
Purpose: Multi-expert deep learning training methods to automatically quantify ischemic brain tissue on Non-Contrast CT Materials and Methods: The data set consisted of 260 Non-Contrast CTs from 233 patients of acute ischemic stroke patients recruited in the DEFUSE 3 trial. A benchmark U-Net was trained on the reference annotations of three experienced neuroradiologists to segment ischemic brain t…
▽ More
Purpose: Multi-expert deep learning training methods to automatically quantify ischemic brain tissue on Non-Contrast CT Materials and Methods: The data set consisted of 260 Non-Contrast CTs from 233 patients of acute ischemic stroke patients recruited in the DEFUSE 3 trial. A benchmark U-Net was trained on the reference annotations of three experienced neuroradiologists to segment ischemic brain tissue using majority vote and random expert sampling training schemes. We used a one-sided Wilcoxon signed-rank test on a set of segmentation metrics to compare bootstrapped point estimates of the training schemes with the inter-expert agreement and ratio of variance for consistency analysis. We further compare volumes with the 24h-follow-up DWI (final infarct core) in the patient subgroup with full reperfusion and we test volumes for correlation to the clinical outcome (mRS after 30 and 90 days) with the Spearman method. Results: Random expert sampling leads to a model that shows better agreement with experts than experts agree among themselves and better agreement than the agreement between experts and a majority-vote model performance (Surface Dice at Tolerance 5mm improvement of 61% to 0.70 +- 0.03 and Dice improvement of 25% to 0.50 +- 0.04). The model-based predicted volume similarly estimated the final infarct volume and correlated better to the clinical outcome than CT perfusion. Conclusion: A model trained on random expert sampling can identify the presence and location of acute ischemic brain tissue on Non-Contrast CT similar to CT perfusion and with better consistency than experts. This may further secure the selection of patients eligible for endovascular treatment in less specialized hospitals.
△ Less
Submitted 7 September, 2023;
originally announced September 2023.
-
Non-inferiority of Deep Learning Acute Ischemic Stroke Segmentation on Non-Contrast CT Compared to Expert Neuroradiologists
Authors:
Sophie Ostmeier,
Brian Axelrod,
Benjamin F. J. Verhaaren,
Soren Christensen,
Abdelkader Mahammedi,
Yongkai Liu,
Benjamin Pulli,
Li-Jia Li,
Greg Zaharchuk,
Jeremy J. Heit
Abstract:
To determine if a convolutional neural network (CNN) deep learning model can accurately segment acute ischemic changes on non-contrast CT compared to neuroradiologists. Non-contrast CT (NCCT) examinations from 232 acute ischemic stroke patients who were enrolled in the DEFUSE 3 trial were included in this study. Three experienced neuroradiologists independently segmented hypodensity that reflected…
▽ More
To determine if a convolutional neural network (CNN) deep learning model can accurately segment acute ischemic changes on non-contrast CT compared to neuroradiologists. Non-contrast CT (NCCT) examinations from 232 acute ischemic stroke patients who were enrolled in the DEFUSE 3 trial were included in this study. Three experienced neuroradiologists independently segmented hypodensity that reflected the ischemic core on each scan. The neuroradiologist with the most experience (expert A) served as the ground truth for deep learning model training. Two additional neuroradiologists (experts B and C) segmentations were used for data testing. The 232 studies were randomly split into training and test sets. The training set was further randomly divided into 5 folds with training and validation sets. A 3-dimensional CNN architecture was trained and optimized to predict the segmentations of expert A from NCCT. The performance of the model was assessed using a set of volume, overlap, and distance metrics using non-inferiority thresholds of 20%, 3ml, and 3mm. The optimized model trained on expert A was compared to test experts B and C. We used a one-sided Wilcoxon signed-rank test to test for the non-inferiority of the model-expert compared to the inter-expert agreement. The final model performance for the ischemic core segmentation task reached a performance of 0.46+-0.09 Surface Dice at Tolerance 5mm and 0.47+-0.13 Dice when trained on expert A. Compared to the two test neuroradiologists the model-expert agreement was non-inferior to the inter-expert agreement, p < 0.05. The CNN accurately delineates the hypodense ischemic core on NCCT in acute ischemic stroke patients with an accuracy comparable to neuroradiologists.
△ Less
Submitted 7 September, 2023; v1 submitted 24 November, 2022;
originally announced November 2022.
-
USE-Evaluator: Performance Metrics for Medical Image Segmentation Models with Uncertain, Small or Empty Reference Annotations
Authors:
Sophie Ostmeier,
Brian Axelrod,
Jeroen Bertels,
Fabian Isensee,
Maarten G. Lansberg,
Soren Christensen,
Gregory W. Albers,
Li-Jia Li,
Jeremy J. Heit
Abstract:
Performance metrics for medical image segmentation models are used to measure the agreement between the reference annotation and the predicted segmentation. Usually, overlap metrics, such as the Dice, are used as a metric to evaluate the performance of these models in order for results to be comparable. However, there is a mismatch between the distributions of cases and difficulty level of segment…
▽ More
Performance metrics for medical image segmentation models are used to measure the agreement between the reference annotation and the predicted segmentation. Usually, overlap metrics, such as the Dice, are used as a metric to evaluate the performance of these models in order for results to be comparable. However, there is a mismatch between the distributions of cases and difficulty level of segmentation tasks in public data sets compared to clinical practice. Common metrics fail to measure the impact of this mismatch, especially for clinical data sets that include low signal pathologies, a difficult segmentation task, and uncertain, small, or empty reference annotations. This limitation may result in ineffective research of machine learning practitioners in designing and optimizing models. Dimensions of evaluating clinical value include consideration of the uncertainty of reference annotations, independence from reference annotation volume size, and evaluation of classification of empty reference annotations. We study how uncertain, small, and empty reference annotations influence the value of metrics for medical image segmentation on an in-house data set regardless of the model. We examine metrics behavior on the predictions of a standard deep learning framework in order to identify metrics with clinical value. We compare to a public benchmark data set (BraTS 2019) with a high-signal pathology and certain, larger, and no empty reference annotations. We may show machine learning practitioners, how uncertain, small, or empty reference annotations require a rethinking of the evaluation and optimizing procedures. The evaluation code was released to encourage further analysis of this topic. https://github.com/SophieOstmeier/UncertainSmallEmpty.git
△ Less
Submitted 7 September, 2023; v1 submitted 26 September, 2022;
originally announced September 2022.
-
Diffusion-Weighted Magnetic Resonance Brain Images Generation with Generative Adversarial Networks and Variational Autoencoders: A Comparison Study
Authors:
Alejandro UngrĂa Hirte,
Moritz Platscher,
Thomas Joyce,
Jeremy J. Heit,
Eric Tranvinh,
Christian Federau
Abstract:
We show that high quality, diverse and realistic-looking diffusion-weighted magnetic resonance images can be synthesized using deep generative models. Based on professional neuroradiologists' evaluations and diverse metrics with respect to quality and diversity of the generated synthetic brain images, we present two networks, the Introspective Variational Autoencoder and the Style-Based GAN, that…
▽ More
We show that high quality, diverse and realistic-looking diffusion-weighted magnetic resonance images can be synthesized using deep generative models. Based on professional neuroradiologists' evaluations and diverse metrics with respect to quality and diversity of the generated synthetic brain images, we present two networks, the Introspective Variational Autoencoder and the Style-Based GAN, that qualify for data augmentation in the medical field, where information is saved in a dispatched and inhomogeneous way and access to it is in many aspects restricted.
△ Less
Submitted 24 June, 2020;
originally announced June 2020.
-
Stream-subhalo interactions in the Aquarius simulations
Authors:
Robyn E. Sanderson,
Carlos Vera-Ciro,
Amina Helmi,
Joren Heit
Abstract:
We perform the first self-consistent measurement of the rate of interactions between stellar tidal streams created by disrupting satellites and dark subhalos in a cosmological simulation of a Milky-Way-mass galaxy. Using a retagged version of the Aquarius A dark-matter-only simulation, we selected 18 streams of tagged star particles that appear thin at the present day and followed them from the po…
▽ More
We perform the first self-consistent measurement of the rate of interactions between stellar tidal streams created by disrupting satellites and dark subhalos in a cosmological simulation of a Milky-Way-mass galaxy. Using a retagged version of the Aquarius A dark-matter-only simulation, we selected 18 streams of tagged star particles that appear thin at the present day and followed them from the point their progenitors accrete onto the main halo, recording in each snapshot the characteristics of all dark-matter subhalos passing within several distance thresholds of any tagged star particle in each stream. We considered distance thresholds corresponding to constant impact parameters (1, 2, and 5 kpc), as well as those proportional to the region of influence of each subhalo (one and two times its half-mass radius $r_{1/2}$). We then measured the age and present-day, phase-unwrapped length of each stream in order to compute the interaction rate in different mass bins and for different thresholds, and compared these to analytic predictions from the literature. We measure a median rate of $1.5^{+3.0}_{-1.1}\ (9.1^{+17.5}_{-7.1},\ 61.8^{+211}_{-40.6})$ interactions within 1 (2, 5) kpc of the stream per 10 kpc of stream length per 10 Gyr. Resolution effects (both time and particle number) affect these estimated rates by lowering them.
△ Less
Submitted 19 August, 2016;
originally announced August 2016.