global
Variáveis
Utilitários
ESTILOS PERSONALIZADOS

Best Practices for the Safe Use of Large Language Models and Other Generative AI in Radiology

Radiology - Volume 316, Number 3 - https://doi.org/10.1148/radiol.241516

Download PDF

Summary

As large language models (LLMs) and other generative artificial intelligence (AI) models are rapidly integrated into radiology workflows, unique pitfalls threatening their safe use have emerged. Problems with AI are often identified only after public release, highlighting the need for preventive measures to mitigate negative impacts and ensure safe, effective deployment into clinical settings. This article summarizes best practices for the safe use of LLMs and other generative AI models in radiology, focusing on three key areas that can lead to pitfalls if overlooked: regulatory issues, data privacy, and bias. To address these areas and minimize risk to patients, radiologists must examine all potential failure modes and ensure vendor transparency. These best practices are based on the best available evidence and the experiences of leaders in the field. Ultimately, this article provides actionable guidelines for radiologists, radiology departments, and vendors using and integrating generative AI into radiology workflows, offering a framework to prevent these problems. © RSNA, 2025

Key points

  • Summary Best practices for the safe use of large language models and other generative artificial intelligence in radiology require vendor transparency and addressing potential pitfalls in three key areas: regulation, data privacy, and bias. Essentials
  • Best practices for the safe use of large language models and other generative artificial intelligence (AI) in radiology focus on addressing potential pitfalls in three key areas: regulatory issues, data privacy, and bias.
  • Evolving regulatory frameworks require examining potential failure modes of generative AI beyond statistical metrics: similarity between synthetic and real-world data, output reproducibility, robustness against hallucinations, and the effects of human-computer interaction.
  • Data privacy risks of generative AI tools in radiology can be mitigated through transparent discussions about the origin of the training data and how vendors will use the data resulting from clinical use.
  • Generative AI biases from homogeneous training data that exclude historically underserved populations may lead to performance variations among demographic groups, perpetuating harmful stereotypes and misinformation.
  • Vendor transparency is a key practical way for radiologists to identify potential data privacy issues and mitigate the use of biased generative AI tools in clinical practice.

Introduction

Since the release of OpenAI’s ChatGPT in November 2022 1, large language models (LLMs) and generative artificial intelligence (AI) have transformed the technology landscape. In medicine and radiology, LLMs have generated excitement for their potential to automate complex, labor-intensive tasks 25 ( Fig 1 ). Early examples of LLMs in radiology include answering patient questions 69, translating radiology reports into plain language 10,11, and extracting data from radiology reports 12,13. Generative AI models have been explored for creating hyperrealistic medical images to replace traditional image processing techniques, such as cinematic rendering 14. One study evaluated a generative AI tool that produced radiology reports from chest radiograph images and found that the AI-generated reports were similar in accuracy and quality to radiologist reports while providing higher quality than teleradiology preliminary reports 15. A randomized trial also demonstrated that LLMs can improve physician diagnostic reasoning 16, highlighting their potential to improve clinical care.

Diagram shows potential use cases of large language models (LLMs) and other generative artificial intelligence (GenAI) m

Figure 1: Diagram shows potential use cases of large language models (LLMs) and other generative artificial intelligence (GenAI) models in radiology. Potential use cases include answering patient questions, patient-friendly radiology report generation, data extraction, image processing, radiology report generation (geared toward other physicians), and synthetic data generation. All visual elements used are under appropriate use agreements. “Patient-Friendly Radiology Reports,” “Data Extraction,” and “Radiology Report Generation” visuals are used under a Creative Commons license (CC0 1.0 DEED) from rawpixel.com , rawpixel.com , and openclipart.org , respectively; “Image Processing Techniques” visual is reprinted, with permission, from reference 14 ; “Synthetic Data Generation” visual is used under a Creative Commons license (CC BY-SA 3.0 DEED [ https://creativecommons.org/licenses/by/3.0/us/deed.en ] ) from Wikimedia.org .

Despite this excitement, LLMs have pitfalls that threaten their safe use in radiology, including performance drift 17 and bias 18,19, which align with previous concerns regarding earlier AI applications in medicine 2024. Unique problems associated with LLMs have emerged, such as hallucinations, where LLMs generate convincing but factually incorrect text or replies 7,25. As generative AI models become integrated into clinical practice 2628, radiologists and vendors must remain aware of the potential drawbacks of these technologies to prevent unintended consequences in real-world clinical practice.

In response to the rapid evolution of generative AI models and the unique needs of radiologists, this article describes best practices and pitfalls in using generative AI models in radiology in three key areas: regulatory issues, data privacy, and bias. Although recent articles have provided broad overviews of LLMs and generative AI in radiology 29,30, we focus on these potential pitfalls to guide radiologists and aligned stakeholders in safely evaluating and using generative AI models in radiology. By providing a list of best practices to avoid these pitfalls, we offer actionable guidelines for radiologists, radiology departments, and vendors engaged in using and integrating generative AI models into radiology workflows.

Regulatory Issues

Our review of regulatory issues surrounding LLMs and other generative AI tools in radiology will cover the following: (a) the need for generative AI–specific regulatory frameworks, (b) navigating a continually evolving regulatory landscape, and (c) the need for these regulatory frameworks to look beyond statistical metrics of accuracy. We present the efforts of the U.S. Food and Drug Administration (FDA) as an illustrative example; the FDA’s policies discussed herein are meant to be descriptive rather than prescriptive because specific regulatory guidelines will differ according to country or region. Generative AI regulatory frameworks are nascent but are being developed by the European Union and individual countries, including the United States, the United Kingdom, Australia, and Singapore 31. In addition to the FDA examples, we direct readers to the European Union’s AI Act 32, a comprehensive legal framework for the use of AI in high-risk settings across all sectors (not just health care); this framework will likely be adopted by the European Parliament 33 and complements the medical-specific FDA examples, continuing to be updated as generative AI technologies evolve 34.

Need for Generative AI–specific Regulatory Frameworks

Although generic frameworks for the regulation of AI have been proposed 35, such as the White House Office of Science and Technology Policy’s “Blueprint for an AI Bill of Rights” 36, the FDA advocates for a medicine-specific approach (precise regulation) 35 by considering AI-based software as medical devices 37. However, even these “precise” regulatory policies do not readily apply to generative AI tools. Nongenerative AI models in radiology that acquire, process, or analyze medical images are considered software as medical devices 38 and have historically been classified by the FDA according to intended use 39. In contrast, generative AI tools do not fit these historical categories 25 because of their unique risks and technical and functional characteristics, which demand a new regulatory paradigm. For example, LLMs differ from previous deterministic deep learning tools in radiology because of their scale, complexity, stochastic behavior, versatile capabilities as generalist foundation models 24, and reliance on massive datasets of often nontransparent origin 40. Despite the attention that generative AI tools have received in medicine, generative AI models currently lack precise regulatory frameworks, an area in which policymakers, radiologists, researchers, and vendors should collaborate to develop.

Navigating a Continually Evolving Regulatory Landscape

The regulatory landscape of AI in medicine is evolving, with much of the global discussion led by the FDA 38,40. The FDA has acknowledged that its “traditional paradigm of medical device regulation was not designed for adaptive artificial intelligence and machine learning technologies” 41. The FDA guidelines for computer-aided detection and diagnosis date back to 2012 42, and the concept of software as a medical device, as defined by the International Medical Device Regulators Forum, dates back to at least 2013 43. However, the new wave of AI tools over the past decade has necessitated updated regulation. In response, the FDA has published iterative regulatory frameworks 37,4447 and collaborated internationally with agencies in Canada and the United Kingdom 45. These recommendations have been summarized in detail elsewhere 38,40. With the rapid development of generative AI tools, regulatory frameworks from the FDA and other agencies will continue to evolve. There is currently no specific FDA guidance for the use of generative AI tools in medicine.

Although definitive FDA guidelines for regulating generative AI tools have not been established, current general recommendations for AI-based software as a medical device provide practical insights. To meet the FDA classification criteria for a medical device, software must analyze medical images, signals, or patterns 48. Many generative AI models may not qualify as medical devices under the latest FDA guidance for clinical decision support (CDS) software from September 2022 48. That guidance outlines four criteria for nondevice CDS software, defined as software that provides recommendations to health care providers rather than specific outputs or directives ( Fig 2 ) and (a) does not acquire, process, or analyze medical images, signals, or patterns; (b) displays, analyzes, or prints medical information normally communicated between health care professionals; (c) provides recommendations to a health care professional, rather than providing a specific output or directive; and (d) provides the basis of recommendations so that a health care professional does not rely primarily on these recommendations to make a decision. As noted by Goodman et al 25, this guidance “provides an unintentional ‘roadmap’ for how LLMs could avoid FDA regulation;” if software does not meet one or more of the four criteria, it would not fall under FDA oversight. For example, an LLM-based tool that summarizes a patient’s medical chart or simplifies a radiology report 10 would be considered “nondevice CDS” because it summarizes patient data (ie, general information provision, which meets one of the four nondevice CDS criteria) and uses text data. However, if an LLM-based tool were used to generate a risk score for a specific disease, it could be classified as a medical device subject to FDA oversight. This ambiguity regarding whether generative AI tools in radiology qualify as medical devices remains unresolved. Radiologists, researchers, and vendors must be aware of the FDA’s regulatory criteria in their clinical deployment, development, and commercialization of generative AI tools. This ambiguity will likely grow as we enter the era of multimodal LLMs, such as OpenAI’s GPT-4o 49, which can accept image and text inputs, and Google’s Med-Gemini 50, which can handle multiple medical data modalities, including radiology images, pathology images, text, and genomics data. Questions also persist about how the FDA will regulate generative AI models integrated into previously cleared and approved medical devices.

U.S. Food and Drug Administration (FDA) guidance for determining whether a clinical decision support medical software is

Figure 2: U.S. Food and Drug Administration (FDA) guidance for determining whether a clinical decision support medical software is considered a device that falls under device regulation. This infographic has modified content released by the FDA at .

The unique technical aspects of generative AI models introduce regulatory challenges. For example, LLMs are often overshadowed by newer models: OpenAI’s GPT-3.5 was released in November 2022, followed by GPT-4 in March 2023 51 and GPT-4o in May 2024 49; any regulation on one generation of an LLM may not apply to a new one. Even without formal upgrades, LLMs exhibit performance drift 17,52, whereby their performance fluctuates over time, necessitating continuous monitoring. Open-source or locally deployed generative AI models may prevent these performance drifts by allowing users to control weights and biases. This is an advantage over proprietary (closed-source) models. As the FDA clears static versions of software, model updates and changes in LLM behavior pose ongoing regulatory challenges 38 ( Fig 3 ). For example, if a medical device software builds a user interface on top of a pre-existing LLM (ie, a GPT wrapper 53), has its underlying LLM undergone an unexpected update, and is it still technically FDA-cleared? Similarly, techniques for optimizing generative AI models, such as prompt engineering, hyperparameter adjustments (eg, adjusting temperature to control output randomness and creativity) 29, and retrieval-augmented generation 54 may complicate regulation. These could be considered “updates” to software because they affect performance and potential risks of model outputs. Radiologists should be aware of potential regulatory nuances related to the technical features of generative AI models to effectively interpret and engage with guidelines from the FDA and other regulatory organizations.

Technical features of large language models (LLMs) and other generative artificial intelligence (AI) models that present

Figure 3: Technical features of large language models (LLMs) and other generative artificial intelligence (AI) models that present unique regulatory challenges to these technologies in radiology. Temperature refers to a parameter that controls the randomness or creativity of the outputs generated by an LLM or other generative AI model.

Amid the uncertain future of regulatory policies for generative AI models, a potential regulatory framework could distinguish between generalist foundation models and their specific end-user applications. For example, a generalist medical foundation model like Google’s Med-Gemini 50 is unlikely to be feasibly regulated. In contrast, a specific end-user application, such as radiology report summarization built on Med-Gemini, can be feasibly regulated in a manner similar to more traditional medical devices regarding performance, repeatability, and bias. End-user applications are more easily regulated on their specific outputs compared with foundation models, representing a feasible path for regulation of LLMs and generative AI in real-world clinical use.

In the absence of clear guidance for regulating generative AI models, there is an opportunity for discussion among the FDA, other government agencies, radiologists, researchers, and vendors. The FDA has invited feedback on proposed regulatory frameworks for AI-based software as a medical device 37. The National Institutes of Health’s Advanced Research Projects Agency for Health has recently launched the Chatbot Accuracy and Reliability Evaluation Exploration Topic 55 to fund the development of approaches for evaluating and improving LLM-based chatbots for patient-facing applications to inform regulatory guidance. As generative AI evolves, opportunities to engage with the FDA and other regulatory bodies will likely continue. Radiologists involved in developing these technologies for clinical workflows should be prepared to participate in these discussions.

The Need for Regulatory Frameworks for Generative AI Tools in Radiology to Look beyond Accuracy

FDA clearance of AI-based software as medical devices has typically relied on standard diagnostic performance measures like area under the receiver operating characteristic curve, accuracy, sensitivity, and specificity 38,42. However, these statistical metrics can be misleading for AI tools, especially for generative AI models, which have unique failure modes that require evaluation. This includes measures of similarity between synthetic and real-world data, output reproducibility, robustness against hallucinations, and the effects of human-computer interaction ( Table 1 ). Even metrics of similarity between machine-generated and human-generated texts, such as BLEU 9 (Bilingual Evaluation Understudy) 56 and ROUGE (Recall-Oriented Understudy for Gisting Evaluation) 57, have been shown to be misleading regarding the medical accuracy of LLM-generated radiology reports 58. Because LLMs are probabilistic by design, they do not produce the same response to the same prompt, leading to inconsistent outputs 7,17,52,59. This makes reproducibility an important diagnostic performance attribute for these tools in clinical practice 25. Similarly, the robustness of LLMs to hallucinations is important to evaluate because of the insidious nature of hallucinated details in otherwise accurate text outputs. Specific forms of LLM hallucinations 25 include sycophancy, where an LLM tailors text output to user expectations indicated in the prompt. For example, providing a leading differential diagnosis can cause an LLM to emphasize that particular diagnosis in its response. For “complete-the-narrative” errors, an LLM summarizes a clinical narrative, such as a radiology report, by adding a brief but clinically meaningful error, like describing a single radiologic finding that never existed. Further technical details on these pitfalls have been described elsewhere 52,60.

Table 1: Performance Metrics beyond Standard Statistical Metrics * with Examples and Challenges for LLMs and Other Generative AI Models in Radiology

Performance MetricExamplesChallenges
Quantitative measures of similarity between synthetic and real-world data, which can be evaluated using metrics such as the structural similarity index measureBLEU (56) and ROUGE (57) for comparing similarity of LLM-generated text to human-generated textThese metrics may not capture medical accuracy and can be misleading regarding the quality of synthesized data
Reproducibility of outputsEvaluating the consistency of outputs from an LLM to the same prompt over three separate sessionsPerformance drift (17,52) makes it difficult to reliably know whether reproducibility of model outputs will hold in the future
Robustness to hallucinationsSycophancy, in which an LLM tailors text output to user expectations indicated in the prompt (eg, providing a leading differential diagnosis can cause an LLM to emphasize that particular diagnosis in its response)Evaluating hallucinations requires domain experts to evaluate the veracity of the generated text, image, or other data outputs
Human-computer interaction effectsEvaluating the effect of an LLM or other generative AI tool on a radiologist’s productivity and cognitive loadThere is no one-size-fits-all set of metrics for these human-computer interaction effects; each metric will have to be tailored for a specific use case and clinical setting

Note.—AI = artificial intelligence, BLEU = Bilingual Evaluation Understudy, LLM = large language model, ROUGE = Recall-Oriented Understudy for Gisting Evaluation.

Standard statistical metrics refers to standard methods used to evaluate AI models statistically, such as area under the receiver operating characteristic curve, sensitivity, and specificity.

Another important consideration for evaluating and regulating generative AI models is interaction with human end-users, also known as human-computer interaction. How do these tools perform in real-world practice? It is crucial to evaluate their effect on human performance relevant to specific tasks. For example, if an LLM is used to predraft responses to electronic medical record inbox questions from patients or other health care providers, reductions in cognitive burden and time spent on these questions would be important factors to evaluate. A recent prospective study found that using LLMs to draft responses to patient questions in an electronic medical record inbox reduced measures of cognitive burden but did not change the time spent on the task 61. Evaluating physician diagnostic performance is also important, given the risk of automation bias (overreliance on AI) and potentially worsened performance. Recent studies on the effect of AI-based tools to augment physician interpretations of medical images have shown variable results 6265. Factors influencing AI’s effectiveness include medical specialty, model performance, whether a model is fair or biased 62,64, and the need for personalized clinician-AI collaboration 65.

Generative AI models in radiology may also be patient-facing to improve their health literacy and empower patients with accessible information for informed health care decisions. For such use cases, regulatory considerations must focus on the patient experience and depend on each clinical use case and deployment setting. Early examples of patient-centric metrics for LLMs include the readability of text (eg, radiology reports or answers to common patient questions) generated by an LLM 8,10,66, with American Medical Association recommendations suggesting readability at the sixth-grade level or lower 67. In addition, evaluating the empathy or bedside manner of LLM-generated responses is important 68. Another issue related to health equity and inclusion is assessing LLM performance across different languages 69,70 because many languages other than English are spoken in the United States 71.

Finally, as generative AI models expand from text generation to other modalities, such as images 14, new evaluation paradigms will be needed. Although quantitative metrics, such as the structural similarity index, can compare generated images with real-world images, they do not always reflect human perception or accurately portray such concepts as anatomy 14 ( Fig 4 ). As these models in radiology evolve into imaging and other modalities, new regulatory frameworks and supporting research will be essential.

Example of a hyperrealistic image produced by a generative artificial intelligence (AI) model (Midjourney, version 5.2;

Figure 4: Example of a hyperrealistic image produced by a generative artificial intelligence (AI) model (Midjourney, version 5.2; https://www.midjourney.com ) according to the following prompt: “A single human deltoid muscle, anatomic realistic, sectioned showing muscle tissue, black background, macro photography, ultra detailed, hyper realism, 8k, specular reflection, shiny–v 5.2–style raw.” The arrows show anatomically inaccurate tendon insertions. Reprinted, with permission, from reference 14 .

Data Privacy

This section discusses privacy concerns related to training LLMs and other generative AI models with patient data and the resulting privacy implications. Protecting patient data privacy is crucial in developing AI tools in radiology 40,7274, especially with generative AI models 29,40. Generative AI models are trained on large datasets of often uncertain origin 75, raising concerns about the potential inclusion of copyrighted material 76 and its unintentional reproduction. Similar concerns exist for generative AI models in radiology, particularly regarding training datasets that may contain sensitive patient data, such as protected health information or records of patients who have opted out of data use for research. Data privacy concerns are amplified by the unique vulnerabilities of LLMs, which are prone to jailbreak attacks that remove models’ guardrails and safety measures through prompt engineering. These attacks enable LLMs to extract verbatim memorized data from their training 77. Malicious jailbreak attacks could extract patient data from an LLM, posing a tremendous risk to patient privacy. Image generative AI models, such as Stable Diffusion, have shown the ability to memorize training images and output them during image generation 78. Because a Stable Diffusion adaptation for chest radiograph synthesis has been developed 79, such models could inadvertently generate images from the training dataset that violate patient privacy.

Beyond patient data privacy risks related to the models and their training data, there is a risk associated with using them on new, unexposed data. Although ChatGPT-3.5 and ChatGPT-4 have been applied to numerous radiology-related tasks 6,10,11,80, these are proprietary models owned by OpenAI that require users to send their data directly to the company via its application programming interface. This precludes the use of patient data that has not been deidentified 74 according to the Health Insurance Portability and Accountability Act (HIPAA) 81,82. Some companies can meet HIPAA requirements through robust security measures and a signed business associate agreement. Deploying open-source models locally or on secure cloud servers can enable the entry of patient data into generative AI models while ensuring data security. The need to deidentify text and images is challenging because manual deidentification of medical data is labor-intensive, and automated methods are not fully effective 8386. Furthermore, depending on local or institutional policies, sending even deidentified data to an LLM or other generative AI tool may not be allowed; popular public medical datasets may be prohibited from being shared with third parties, including online services such as ChatGPT 87. Prohibitions on sending data to external parties may arise from concerns that companies such as OpenAI store data to update and improve their models 88. Users of patient data or anonymized research data must understand any applicable data restrictions or policies, such as HIPAA 29. An alternative to using commercial third-party LLMs for radiology report disease label extraction is locally deploying open-source LLMs that protect patient data privacy 74. Similar approaches with locally deployed generative AI models may allow for compliance with HIPAA and other data privacy regulations, even when working with patient data that have not been deidentified. Federated learning enables decentralized model training without sharing sensitive data 72, providing a practical solution for safely using LLMs with patient data.

Radiologists considering purchasing commercial products that use LLMs or other generative AI models should ask for transparency regarding data use 29. A clear data use agreement should be provided, along with the vendors’ willingness to answer questions about it. Vendors should also provide a standard data privacy policy outlining protective measures for patient data and plans to address potential data breaches. Radiologists can inquire about measures taken by vendors to prevent security risks like jailbreaking attacks. The same principles for selecting any radiology AI software 89,90 will apply to generative AI models, and radiologists should engage allied partners, such as the local IT or informatics team, when evaluating potential generative AI tools for clinical deployment. For more on ensuring data security when using LLMs clinically, which is beyond the scope of this article, interested readers can refer to the relevant literature 9194.

Bias

AI models in radiology can produce biased predictions 22,23,95, threatening their safe and equitable use. Bias often stems from homogeneous training datasets that exclude historically underserved and underrepresented populations 23,9698. In addition, LLMs and other generative AI models have demonstrated bias in generating text and other modality outputs in radiology-related tasks 18,19,99,100, likely because of biased training datasets. LLMs produce different outputs for radiology report simplification 100 or differential diagnosis generation from a clinical vignette 19 by simply changing the race or sex mentioned in the prompt, likely reflecting biases in the training data. Commercial LLMs from major technology companies such as Google and OpenAI have also demonstrated a tendency to propagate inaccurate and harmful race-based medical tropes 18, raising concerns about their potential to contribute to medical misinformation and perpetuate bias. Omiye and Lester et al 18 showed that four commercial LLMs generated debunked race-based content in response to questions about race-based medicine or racial misconceptions, such as asserting that race should be used when calculating kidney function and lung capacity or that Black and White people have differences in skin thickness ( Fig 5 ). In another study evaluating generative AI image generation, OpenAI’s DALL-E 2 model was prompted to generate images of a “radiologist,” amplifying real-world demographic disparities in the radiology workforce (eg, 80% of images depicted men compared with 71% male representation in the workforce), and showing radiologists from minority groups in less formal attire than White men 99. These early examples of bias in generative AI models in radiology and medicine show risks of directly harming patients through differential bias in clinical recommendations and indirectly harming patients and physicians by perpetuating stereotypes. Radiologists using these tools must be aware of the potential harm from underlying biases in these models to avoid introducing or perpetuating health inequities.

Large language model (LLM) outputs to questions and scenarios that check for race-based medicine or misconceptions about

Figure 5: Large language model (LLM) outputs to questions and scenarios that check for race-based medicine or misconceptions about race and medicine and health care. For each question and model evaluated, the rating represents the number of runs (out of five total runs) that had concerning race-based responses. Red correlates with increasing number of concerning race-based responses. The specific models evaluated include Bard, now called Gemini (Google; May 18 and August 3, 2023, versions); ChatGPT (OpenAI; May 12 and August 3, 2023, versions); Claude (Anthropic; May 15 and August 3, 2023, versions); and GPT-4 (OpenAI). eGFR = estimated glomerular filtration rate. The chart is reproduced and the caption adapted from reference 18 under a Creative Commons Attribution 4.0 International License ( https://creativecommons.org/licenses/by/4.0/deed.en ) .

Strategies to mitigate biases in AI models are well described for radiology 101104 and apply to generative AI models. The key strategy is to collect data representing demographic groups reflecting the underlying populations. Radiology imaging datasets used for AI model training have often omitted demographic variables 24,97, which should be considered standard metadata for any AI model—generative or otherwise—to enhance transparency regarding potential bias. Federated learning 72 can train models on large diverse datasets without central data aggregation, potentially reducing biases. Evaluating bias and algorithmic fairness should be routine in assessing generative AI models, although established frameworks and definitions for bias are still lacking. The unique nature of training generative AI models necessitates novel methods to reduce bias, such as targeted reinforcement learning with feedback from physician end-users. However, as Zack and Lehman et al 19 note, such methods may be limited to open-source LLMs because proprietary ones are typically not modifiable. Early work shows promise for measuring biases using standardized frameworks, such as EquityMedQA, a set of medical question–answering datasets designed to identify biases 105, and “stress testing” LLMs with modified prompts, such as swapping races in clinical vignettes 18,19.

As with identifying potential data privacy issues in commercial generative AI models, calling for vendor transparency is a key practical way for radiologists to mitigate the use of biased tools in clinical practice. This includes engaging with vendors to discuss dataset information used in product training and evaluation, as well as subanalyses of relevant demographic groups to identify potential biases. Radiologists should feel empowered to request these data, which are increasingly being solicited at a professional society level by groups such as the American College of Radiology’s Data Science Institute through its Transparent-AI program 106. As with other AI tools 107, radiologists must conduct postimplementation surveillance of generative AI models to evaluate biases, especially because of the unpredictable nature of performance drift 17,52, which could worsen biases over time.

Clinical Implications of These Pitfalls

The aforementioned theoretical details are crucial for understanding the safe use of LLMs and other generative AI tools, but they must be contextualized to specific clinical use cases and settings. Previous work has described the range of radiology clinical use cases for applied technologies 30, including tasks before imaging, such as protocol automation 108; tasks after imaging, such as making imaging diagnoses 109,110; and patient-facing tasks, such as answering questions about imaging 7. Each task has unique safety implications requiring radiologists to use their domain knowledge to anticipate and prepare. For example, a generative AI tool that makes imaging diagnoses can have its outputs checked by a human radiologist as a safeguard against errors or hallucinations, whereas an LLM that answers patient questions does not naturally have an expert to verify its responses (referred to as expert-in-the-loop); thus, the latter may require more extensive robustness and safety testing before deployment. Each use case will be unique, but the goal should remain consistent: to do the right thing for patients.

Future Directions

Tables 2 – 4 summarize best practices and future directions for implementation of generative AI models in real-world settings to avoid pitfalls in three key areas: regulation, data privacy, and bias. We provide future directions in the form of open questions to guide further study and policy development.

Table 2: Best Practices and Future Directions for Avoiding Pitfalls Related to Regulatory Issues in LLMs and Other Generative AI Models in Radiology

Best Practice RecommendationFuture Directions (Open Questions)
LLMs and other generative AI models in radiology should undergo “precise regulation” to account for radiology’s unique workflows and the unique pitfalls of generative AI technology while being guided by “general regulation” principlesWhat specific regulatory frameworks will apply to generative AI models in radiology?
Regulatory guidance and recommendations from the FDA should be closely followed because they are frequently being updatedDo FDA definitions of SaMD need to be revised or expanded to accommodate for generative AI models?
Radiologists should be aware of the FDA’s definitions of clinical decision support software and SaMD to guide the evaluation, development, and clinical use of generative AI models according to FDA and other local and institutional regulatory guidelinesWhat standard metrics should be used to evaluate generative AI models in radiology beyond mere accuracy?
Radiologists should be aware of the potential technical limitations of generative AI models that may inform FDA regulatory frameworksHow should regulation of generative AI models account for unique technical features of these tools, such as stochasticity?
Radiologists are encouraged to engage in discussion and collaboration with the FDA and other regulatory bodies to develop regulatory frameworks for generative AI models
Generative AI models should be evaluated for measures beyond standard diagnostic performance, including repeatability, hallucinations, and human-centric measures (relevant to both physicians and patients)

Table 3: Best Practices and Future Directions for Avoiding Pitfalls Related to Data Privacy in LLMs and Other Generative AI Models in Radiology

Best Practice RecommendationFuture Directions (Open Questions)
When training models with radiology patient data, prevent the inadvertent inclusion of prohibited data (eg, protected health information or data from patients who have opted out of research)What are technically robust methods to prevent the inadvertent inclusion of prohibited patient data during training and development?
Engineer safeguards to prevent jailbreaking that could result in patient data extractionHow can multiple privacy-preserving techniques, such as differential privacy, data encryption, and secure multiparty computation, be combined to ensure data security?
Ensure compliance with local and institutional data privacy policies; exercise similar caution with public datasets, which may prohibit sharing data with third partiesWhat are the potential pitfalls in applying techniques such as the ones listed above in medical data? How can we best translate these techniques in ways that will best work for the unique privacy issues in patient data?
As alternatives to third-party commercial models, consider using open-source models deployed locally or in secure cloud environments to maintain patient data privacyHow susceptible are generative AI models trained on radiology patient data to malicious jailbreaking attempts? Do these susceptibilities differ from those of domain-agnostic generative AI models?
When purchasing commercial products, ask vendors for transparency about data use policies and measures to prevent data jailbreakingHow well do open-source models perform compared with third-party commercial models for varying radiology tasks of clinical interest? What is the trade-off in performance versus security of these various model choices?
Involve interdisciplinary teams in discussions about potential product purchases; sample questions to ask vendors, including specific language and terms to look for in contracts, are outlined elsewhere (29)
Consider methods and tools to address data privacy concerns, such as federated learning (72), that allow for training on multicenter data without sharing outside a center’s firewall; these strategies have shown success in computer vision tasks in radiology (111–113) and are promising for generative AI models in the field

Table 4: Best Practices and Future Directions for Avoiding Pitfalls Related to Bias in LLMs and Other Generative AI Models in Radiology

Best Practice RecommendationFuture Directions (Open Questions)
Bias should be recognized in multiple forms, including erroneous patient care recommendations and indirect perpetuation of stereotypesHow diverse are the datasets used to train generative AI models in radiology?
Datasets used to train and evaluate models should be diverse and represent the underlying populations from which they are drawnWhat are the best ways to identify and measure biases in generative AI models in radiology?
Radiologists should lead the development of frameworks to evaluate for bias and fairness of generative AI in radiology, which are not yet establishedWhat technical methods can be used to mitigate biases in generative AI models in radiology?
Vendors should be asked to provide transparency regarding their models’ potential for biases, including training dataset demographic characteristics and any bias evaluations that have been performed
Consider implementation of techniques and tools during model evaluation that can help identify bias concerns, such as stress-testing LLMs with modified prompts by changing the race or sex (18,19); federated learning strategies (72) that allow for training on multicenter data without requiring sharing of data outside of a center’s firewall can help mitigate bias by increasing the diversity of training data; these strategies have proven effective in computer vision radiology tasks in radiology (111–113) and are promising for generative AI models in the field

Note.—AI = artificial intelligence, FDA = U.S. Food and Drug Administration, LLM = large language model, SaMD = software as a medical device.

Note.—AI = artificial intelligence, LLM = large language model.

Conclusion

As large language models (LLMs) and other generative artificial intelligence (AI) models are integrated into radiology workflows, unique pitfalls threatening their safe use have emerged. These pitfalls are often identified after public release. To mitigate the negative impacts of LLMs and other generative AI models, it is essential to prevent potential pitfalls in three key areas: regulation, data privacy, and bias. We hope these best practices will serve as actionable guidelines for radiologists, radiology departments, and vendors interacting with generative AI models to ensure their safe and effective use.