global
Variáveis
Utilitários
ESTILOS PERSONALIZADOS

The Initial Steps of Multimodal AI in Radiology

Radiology - Volume 309, Number 1 - https://doi.org/10.1148/radiol.232372

Download PDF

See also the article by Khader et al in this issue.

Until recently, the predominant benefit of artificial intelligence (AI) in health care was improving the accuracy of medical image interpretation, a unimodal task. Now, transformer models have enabled multimodal AI, expanding the dimensions of inputs to include electronic health record data and unstructured text, genomic, and sensor data. The concept is tantalizing for radiologists because it mimics their multifaceted decision-making process. But as with any emerging technology, challenges in costs, domain specialization, and system integration lie ahead 1.

In this issue of Radiology , Khader et al 2 introduce a step in this direction. The authors developed a transformer-based AI model that integrates two modes of patient data, specifically chest radiographs and some clinical parameters, to improve diagnostic performance. With a substantial training set across two data sets, the model was trained to diagnose up to 25 diseases.

Using data from the publicly available Medical Information Mart for Intensive Care (MIMIC) database, the model achieved a mean area under the receiver operating characteristic curve (AUC) of 0.77 when both data types were used. In comparison, the performance was statistically significantly lower when only chest radiographs were used (mean AUC, 0.70) and when only clinical parameters were considered (mean AUC, 0.72) ( P < .001) 2. These results show that integrating both imaging and nonimaging data can yield modestly better diagnostic performance.

Multimodal AI is not an entirely new concept in the field of medical imaging; however, its efficient large-scale application in radiology represents a new frontier 2. Prior studies have largely focused on either imaging or clinical data, but rarely both. Thus, this work serves as an example of how combining these two modalities has the potential to improve accuracy. As physicians, we know that having access to pertinent information can benefit our decision-making and patient outcomes, particularly in radiology 3. This is a solid rationale for training AI models with a broader context.

While the results are promising, it is crucial to acknowledge that this bimodal model, based on retrospective data, only scratches the surface, both from the standpoint of the two layers of data it includes and the many dimensions it does not. The clinical parameters used were minimal, and unstructured text from electronic records was not included, despite recently being established by large language models as a data-rich source. The images used two-dimensional data (radiographs), but a considerable part of medical data is three-dimensional (CT and MRI). Additionally, the model’s effectiveness in real-world scenarios outside of the MIMIC data set still needs further exploration, limiting the immediate dissemination of this technique into broader clinical practice. Moreover, the incremental improvement in the mean AUC to 0.77, while statistically significant, is not necessarily clinically impactful, especially when compared with the aspirational benchmark of an AUC of 0.95 or higher for diagnostic performance. This suggests that, while the study is a step in the right direction, it represents a relatively early phase in the development of truly multimodal AI models for health care.

Yet, it does map the future direction for the field, exploiting the capabilities of generative AI to integrate the many layers of orthogonal data to understand more deeply each unique individual. As we are enabled to assimilate input longitudinal data from notes, laboratory tests, genomic data, biosensors, environmental data, and multiple imaging modalities at once, along with the corpus of medical knowledge, the likelihood of achieving maximal accuracy is enhanced. We need to acknowledge that the cumulative multilayered data for an individual patient now exceeds what physicians can process. On the other hand, having radiologists in the loop for oversight, to fold in their experience and wisdom, is absolutely essential.

Recently, Google released Med-PaLM M, a multimodal version of Med-PaLM with inputs of medical images, clinical notes, and genomic data 4. While the feasibility and power of such generative AI in health care is starting to be demonstrated, we need compelling prospective study validation, with independent replication, that using such models promotes the accuracy of medical diagnoses and patient outcomes. Furthermore, there are logistical challenges and substantial costs for implementing these AI systems, especially the need for postimplementation surveillance of their performance.

In summary, the study by Khader et al 2 shows promising early exploratory evidence that the future of AI in radiology could benefit from integrating additional layers of patient data. Also, the authors deposited their code in a public repository so that other investigators can evaluate this new technology and the robustness of this model. Future models that incorporate a broader range of data types will likely lead to a higher level of diagnostic accuracy. This is an important direction of future AI medical research that stands as a hypothesis now and hopefully, through extensive and rigorous work, will be fully proven.

Felipe Kitamura is a neuroradiologist, director of Applied Innovation and AI at Dasa, affiliated professor at Universida
Felipe Kitamura is a neuroradiologist, director of Applied Innovation and AI at Dasa, affiliated professor at Universidade Federal de São Paulo, associate scholar at the Stanford AIMI Center, visiting professor at the Mayo Clinic, member of the RSNA Artificial Intelligence Steering Committee, co-chair of the SIIM Machine Learning Education Subcommittee, consultant for MD.ai, Kaggle Competitions Master, and RSNA Honored Educator Awardee. His interest is in innovation to improve patient care.
Eric Topol is founder and director of the Scripps Research Translational Institute, a professor of molecular medicine, a
Eric Topol is founder and director of the Scripps Research Translational Institute, a professor of molecular medicine, and executive vice-president of Scripps Research. He has published more than 1200 peer-reviewed articles, including more than 300 000 citations, was elected to the National Academy of Medicine, and is one of the top 10 most cited researchers in medicine. His principal scientific focus has been on individualized medicine using genomic, digital, and AI tools.