Artificial Intelligence May Fall Short When Analyzing Data Across Multiple Health Systems
Study shows deep learning models must be carefully tested across multiple environments before being put into clinical practice.
Artificial intelligence (AI) tools trained to detect pneumonia on chest X-rays suffered significant decreases in performance when tested on data from outside health systems, according to a study conducted at the Icahn School of Medicine at Mount and published in a special issue of PLOS Medicine on machine learning and health care. These findings suggest that artificial intelligence in the medical space must be carefully tested for performance across a wide range of populations; otherwise, the deep learning models may not perform as accurately as expected.
As interest in the use of computer system frameworks called convolutional neural networks (CNN) to analyze medical imaging and provide a computer-aided diagnosis grows, recent studies have suggested that AI image classification may not generalize to new data as well as commonly portrayed.
Researchers at the Icahn School of Medicine at Mount Sinai assessed how AI models identified pneumonia in 158,000 chest X-rays across three medical institutions: the National Institutes of Health; The Mount Sinai Hospital; and Indiana University Hospital. Researchers chose to study the diagnosis of pneumonia on chest X-rays for its common occurrence, clinical significance, and prevalence in the research community.
In three out of five comparisons, CNNs’ performance in diagnosing diseases on X-rays from hospitals outside of its own network was significantly lower than on X-rays from the original health system. However, CNNs were able to detect the hospital system where an X-ray was acquired with a high-degree of accuracy, and cheated at their predictive task based on the prevalence of pneumonia at the training institution. Researchers found that the difficulty of using deep learning models in medicine is that they use a massive number of parameters, making it challenging to identify specific variables driving predictions, such as the types of CT scanners used at a hospital and the resolution quality of imaging.
“Our findings should give pause to those considering rapid deployment of artificial intelligence platforms without rigorously assessing their performance in real-world clinical settings reflective of where they are being deployed,” says senior author Eric Oermann, MD, Instructor in Neurosurgery at the Icahn School of Medicine at Mount Sinai. “Deep learning models trained to perform medical diagnosis can generalize well, but this cannot be taken for granted since patient populations and imaging techniques differ significantly across institutions.”
“If CNN systems are to be used for medical diagnosis, they must be tailored to carefully consider clinical questions, tested for a variety of real-world scenarios, and carefully assessed to determine how they impact accurate diagnosis,” says first author John Zech, a medical student at the Icahn School of Medicine at Mount Sinai.
This research builds on papers published earlier this year in the journals Radiology and Nature Medicine, which laid the framework for applying computer vision and deep learning techniques, including natural language processing algorithms, for identifying clinical concepts in radiology reports for CT scans.
About the Mount Sinai Health System
The Mount Sinai Health System is New York City's largest integrated delivery system, encompassing eight hospitals, a leading medical school, and a vast network of ambulatory practices throughout the greater New York region. Mount Sinai's vision is to produce the safest care, the highest quality, the highest satisfaction, the best access and the best value of any health system in the nation. The Health System includes approximately 7,480 primary and specialty care physicians; 11 joint-venture ambulatory surgery centers; more than 410 ambulatory practices throughout the five boroughs of New York City, Westchester, Long Island, and Florida; and 31 affiliated community health centers. The Icahn School of Medicine is one of three medical schools that have earned distinction by multiple indicators: ranked in the top 20 by U.S. News & World Report's "Best Medical Schools", aligned with a U.S. News & World Report's "Honor Roll" Hospital, No. 12 in the nation for National Institutes of Health funding, and among the top 10 most innovative research institutions as ranked by the journal Nature in its Nature Innovation Index. This reflects a special level of excellence in education, clinical practice, and research. The Mount Sinai Hospital is ranked No. 18 on U.S. News & World Report's "Honor Roll" of top U.S. hospitals; it is one of the nation's top 20 hospitals in Cardiology/Heart Surgery, Gastroenterology/GI Surgery, Geriatrics, Nephrology, and Neurology/Neurosurgery, and in the top 50 in six other specialties in the 2018-2019 "Best Hospitals" issue. Mount Sinai's Kravis Children's Hospital also is ranked nationally in five out of ten pediatric specialties by U.S. News & World Report. The New York Eye and Ear Infirmary of Mount Sinai is ranked 11th nationally for Ophthalmology and 44th for Ear, Nose, and Throat. Mount Sinai Beth Israel, Mount Sinai St. Luke's, Mount Sinai West, and South Nassau Communities Hospital are ranked regionally.