Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Katherine Andriole

Report of the Medical Image De-Identification Task Group -- Best Practices and Recommendations

Apr 01, 2023

David A. Clunie, Adam Flanders, Adam Taylor, Brad Erickson, Brian Bialecki, David Brundage, David Gutman, Fred Prior, J Anthony Seibert, John Perry(+9 more)

Abstract:This report addresses the technical aspects of de-identification of medical images of human subjects and biospecimens, such that re-identification risk of ethical, moral, and legal concern is sufficiently reduced to allow unrestricted public sharing for any purpose, regardless of the jurisdiction of the source and distribution sites. All medical images, regardless of the mode of acquisition, are considered, though the primary emphasis is on those with accompanying data elements, especially those encoded in formats in which the data elements are embedded, particularly Digital Imaging and Communications in Medicine (DICOM). These images include image-like objects such as Segmentations, Parametric Maps, and Radiotherapy (RT) Dose objects. The scope also includes related non-image objects, such as RT Structure Sets, Plans and Dose Volume Histograms, Structured Reports, and Presentation States. Only de-identification of publicly released data is considered, and alternative approaches to privacy preservation, such as federated learning for artificial intelligence (AI) model development, are out of scope, as are issues of privacy leakage from AI model sharing. Only technical issues of public sharing are addressed.

* 131 pages

Via

Access Paper or Ask Questions

An Overview and Case Study of the Clinical AI Model Development Life Cycle for Healthcare Systems

Mar 26, 2020

Charles Lu, Julia Strout, Romane Gauriau, Brad Wright, Fabiola Bezerra De Carvalho Marcruz, Varun Buch, Katherine Andriole

Figure 1 for An Overview and Case Study of the Clinical AI Model Development Life Cycle for Healthcare Systems

Figure 2 for An Overview and Case Study of the Clinical AI Model Development Life Cycle for Healthcare Systems

Abstract:Healthcare is one of the most promising areas for machine learning models to make a positive impact. However, successful adoption of AI-based systems in healthcare depends on engaging and educating stakeholders from diverse backgrounds about the development process of AI models. We present a broadly accessible overview of the development life cycle of clinical AI models that is general enough to be adapted to most machine learning projects, and then give an in-depth case study of the development process of a deep learning based system to detect aortic aneurysms in Computed Tomography (CT) exams. We hope other healthcare institutions and clinical practitioners find the insights we share about the development process useful in informing their own model development efforts and to increase the likelihood of successful deployment and integration of AI in healthcare.

* Accepted for oral presentation at ICLR 2020, AI for Affordable Healthcare workshop

Via

Access Paper or Ask Questions

Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

Nov 14, 2019

Neil Deshmukh, Selin Gumustop, Romane Gauriau, Varun Buch, Bradley Wright, Christopher Bridge, Ram Naidu, Katherine Andriole, Bernardo Bizzo

Figure 1 for Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

Figure 2 for Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

Figure 3 for Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

Figure 4 for Semi-Supervised Natural Language Approach for Fine-Grained Classification of Medical Reports

Abstract:Although machine learning has become a powerful tool to augment doctors in clinical analysis, the immense amount of labeled data that is necessary to train supervised learning approaches burdens each development task as time and resource intensive. The vast majority of dense clinical information is stored in written reports, detailing pertinent patient information. The challenge with utilizing natural language data for standard model development is due to the complex nature of the modality. In this research, a model pipeline was developed to utilize an unsupervised approach to train an encoder-language model, a recurrent network, to generate document encodings; which then can be used as features passed into a decoder-classifier model that requires magnitudes less labeled data than previous approaches to differentiate between fine-grained disease classes accurately. The language model was trained on unlabeled radiology reports from the Massachusetts General Hospital Radiology Department (n=218,159) and terminated with a loss of 1.62. The classification models were trained on three labeled datasets of head CT studies of reported patients, presenting large vessel occlusion (n=1403), acute ischemic strokes (n=331), and intracranial hemorrhage (n=4350), to identify a variety of different findings directly from the radiology report data; resulting in AUCs of 0.98, 0.95, and 0.99, respectively, for the large vessel occlusion, acute ischemic stroke, and intracranial hemorrhage datasets. The output encodings are able to be used in conjunction with imaging data, to create models that can process a multitude of different modalities. The ability to automatically extract relevant features from textual data allows for faster model development and integration of textual modality, overall, allowing clinical reports to become a more viable input for more encompassing and accurate deep learning models.

* Accepted for IEEE publication & presented at MIT URTC

Via

Access Paper or Ask Questions

Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks

Sep 13, 2018

Hoo-Chang Shin, Neil A Tenenholtz, Jameson K Rogers, Christopher G Schwarz, Matthew L Senjem, Jeffrey L Gunter, Katherine Andriole, Mark Michalski

Figure 1 for Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks

Figure 2 for Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks

Figure 3 for Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks

Figure 4 for Medical Image Synthesis for Data Augmentation and Anonymization using Generative Adversarial Networks

Abstract:Data diversity is critical to success when training deep learning models. Medical imaging data sets are often imbalanced as pathologic findings are generally rare, which introduces significant challenges when training deep learning models. In this work, we propose a method to generate synthetic abnormal MRI images with brain tumors by training a generative adversarial network using two publicly available data sets of brain MRI. We demonstrate two unique benefits that the synthetic images provide. First, we illustrate improved performance on tumor segmentation by leveraging the synthetic images as a form of data augmentation. Second, we demonstrate the value of generative models as an anonymization tool, achieving comparable tumor segmentation results when trained on the synthetic data versus when trained on real subject data. Together, these results offer a potential solution to two of the largest challenges facing machine learning in medical imaging, namely the small incidence of pathological findings, and the restrictions around sharing of patient data.

* Accepted for 2018 Workshop on Simulation and Synthesis in Medical Imaging - SASHIMI2018

Via

Access Paper or Ask Questions

Fully-Automated Analysis of Body Composition from CT in Cancer Patients Using Convolutional Neural Networks

Aug 11, 2018

Christopher P. Bridge, Michael Rosenthal, Bradley Wright, Gopal Kotecha, Florian Fintelmann, Fabian Troschel, Nityanand Miskin, Khanant Desai, William Wrobel, Ana Babic(+8 more)

Figure 1 for Fully-Automated Analysis of Body Composition from CT in Cancer Patients Using Convolutional Neural Networks

Figure 2 for Fully-Automated Analysis of Body Composition from CT in Cancer Patients Using Convolutional Neural Networks

Figure 3 for Fully-Automated Analysis of Body Composition from CT in Cancer Patients Using Convolutional Neural Networks

Figure 4 for Fully-Automated Analysis of Body Composition from CT in Cancer Patients Using Convolutional Neural Networks

Abstract:The amounts of muscle and fat in a person's body, known as body composition, are correlated with cancer risks, cancer survival, and cardiovascular risk. The current gold standard for measuring body composition requires time-consuming manual segmentation of CT images by an expert reader. In this work, we describe a two-step process to fully automate the analysis of CT body composition using a DenseNet to select the CT slice and U-Net to perform segmentation. We train and test our methods on independent cohorts. Our results show Dice scores (0.95-0.98) and correlation coefficients (R=0.99) that are favorable compared to human readers. These results suggest that fully automated body composition analysis is feasible, which could enable both clinical use and large-scale population studies.

Via

Access Paper or Ask Questions