ChatGPT Is Shaping the Future of Medical Writing But Still Requires Human Judgment

See also the article by Biswas and the editorial by Shen et al.
In this issue of Radiology , the article by Biswas 2 explains a breakthrough in artificial intelligence (AI) text generation called ChatGPT. This AI-based writing assistant can generate text for medical writing, including scientific articles. What is fascinating is that the article is a self-reference. It was almost all written using ChatGPT, except for the introduction, cautions, and edits by the human author. The article headings reflect the exact prompts inputted by the human author. For example, one header prompt that created the next paragraph written by ChatGPT was "Will chatGPT replace human medical writer?" This makes it an article that talks about itself (meta-text), proving the capabilities of ChatGPT as an assistive tool in medical writing 2.
The Biswas article explains key concepts such as medical writing, natural language processing, generative pretraining transformer (GPT), and ChatGPT. Medical writing is the process of creating any text-based communication involving medical knowledge. Medical writing includes, but is not limited to, scientific articles, radiologic reports, regulatory documents, and patient educational materials. Natural language processing is an interdisciplinary field of computer science and linguistics that develops ways for machines to interpret text. More simply, it helps computers communicate with humans in our natural language (written, spoken, etc). Recent advances in AI, namely the transformers 3 (a deep learning architecture), have increased the capabilities of natural language processing to classify, translate, and generate text. GPT is a large transformer model 4 trained on an extensive text database (often called a corpus).
The training process of GPT involves asking the model to predict the next element in a sentence, which enables GPT to generate random free text, mimicking the training text (and being biased by it). ChatGPT is a more complex model that requires two more steps. After pretraining GPT with a large corpus, GPT is trained to predict the answer to a given question. That process is called supervised learning. Supervised learning is done using a data set of paired questions and answers created by humans. Next, GPT generates many answers for every question, and human labelers rank the best outputs. These data are used to train another model using a different machine learning paradigm called reinforcement learning. A reinforcement learning model rewards desired behaviors and punishes undesired ones, allowing the model to learn through trial and error. ChatGPT has arguably created the most realistic text to resemble human writing in many domains, including natural, social, and formal sciences 5.
The Biswas 2 article lists many possible uses for ChatGPT in medical writing. These include real-time assistance or creating drafts for clinical trial protocols, study reports, regulatory documents, and patient-facing materials, as well as translation of medical information into a myriad of languages. The possibilities seem endless, even outside medicine, and the benefit of making writing more efficient drew the world’s attention, with a million users in 5 days. Also in this issue of Radiology , Shen et al 6 provide a more detailed description of possible use cases for ChatGPT.
They Who Love Roses Must Endure the Thorns
Before ChatGPT, GPT had been considered a breakthrough with its own limitations in health care 7. Although the experience with ChatGPT might be astonishing, it also has many limitations. Biswas 2 notes several concerns. These include ethical concerns about authorship and accountability for AI-generated content. Editing by human authors is necessary to prevent plagiarism (discussed more herein).
There are also multiple legal issues to consider. When AI-generated content is used for commercial purposes, care must be taken that there is no copyright infringement. Human authors must ensure that the use of AI-generated text complies with any relevant regulations or laws. There are currently no laws addressing the use of AI in the medical literature. Medicolegal issues must also be considered; for example, errors in radiologic reports created by AI could lead to lawsuits. This leads to questions of accountability.
A lack of original thought is also a concern. As ChatGPT is based on prior data fed to it, it will lead to repetitive text generation and, thus, lack the creativity and originality of human authors. Inaccuracy, bias, and transparency are other noted issues. Regarding transparency, the role of AI in the writing process must be clear.
Many of these concerns are also recognized by OpenAI, the research company that created ChatGPT, noting that “ChatGPT sometimes writes plausible-sounding but incorrect or nonsensical answers” 5. This means that health care professionals using ChatGPT who are unaware of its limitations could potentially harm patients. Many examples have been posted on social media in the past weeks showing compelling medical arguments written by ChatGPT, including references to scientific articles that allegedly would prove the point. Yet, a closer look at the references indicates that the journals and the authors exist but the title of the article does not. And the digital object identifier (DOI) links to another unrelated article. The bottom line is that we must be cognizant that ChatGPT sometimes writes incorrect answers. Using it to expedite the writing process is helpful—provided we check the accuracy of the text and references.
OpenAI also states that ChatGPT sometimes “responds to harmful instructions or exhibits biased behavior” 5. We should be aware of that.
Another concern in using ChatGPT is plagiarism. Because the model is trained on publicly available text on the internet, it might (and frequently will) copy phrases or even entire sentences from other documents. I am curious to learn if scientific journals will see an increase in the plagiarism percentage of received manuscripts because of this limitation. A workaround is checking the plagiarism percentage using validated tools and paraphrasing segments of ChatGPT’s output.
Some authors have included ChatGPT in the author list of scientific articles 8, a practice seen with skepticism by some journals because drafting the work is only one of four required International Committee of Medical Journal Editors criteria for authorship 9. Another criterion is the “agreement to be accountable for all aspects of the work,” which is not in the skill set of ChatGPT.
Despite these limitations, ChatGPT has raised the bar, bringing writing support to the next level. With proper use, writers can benefit from this tool. However, instead of having humans in the loop, humans should be in charge—with AI in the loop—which was the motto of the Stanford University 2022 Human-Centered AI (HAI) Fall Conference 10.
Looking toward the future, better large language models will likely be developed, reducing their monetary costs and mitigating some of their current limitations. Their use will eventually become widespread and integrated into pervasive text editing software. There is a hypothetical future in which the article you are currently reading will be one of the last to be written without the help of AI.
I encourage readers to test ChatGPT at their own discretion.



