Authors: Tongyao Wang, Juan Mu, Jialing Chen, Chia-Chin Lin
Categories: Article, ChatGPT, Education, Generative artificial intelligence, Nursing, Tracheostomy
Source: International Journal of Nursing Studies Advances
The release of ChatGPT for general use in 2023 by OpenAI has significantly expanded the possible applications of generative artificial intelligence in the healthcare sector, particularly in terms of information retrieval by patients, medical and nursing students, and healthcare personnel.
To compare the performance of ChatGPT-3.5 and ChatGPT-4.0 to clinical nurses on answering questions about tracheostomy care, as well as to determine whether using different prompts to pre-define the scope of the ChatGPT affects the accuracy of their responses.
Cross-sectional study.
The data collected from the ChatGPT was collected using the ChatGPT-3.5 and 4.0 using access provided by the University of Hong Kong. The data from the clinical nurses working in mainland China was collected using the Qualtrics survey program.
No participants were needed for collecting the ChatGPT responses. A total of 272 clinical nurses, with 98.5 % of them working in tertiary care hospitals in mainland China, were recruited using a snowball sampling approach.
We used 43 tracheostomy care-related questions in a multiple-choice format to evaluate the performance of ChatGPT-3.5, ChatGPT-4.0, and clinical nurses. ChatGPT-3.5 and GPT-4.0 were both queried three times with the same questions by different no prompt, patient-friendly prompt, and act-as-nurse prompt. All responses were independently graded by two qualified otorhinolaryngology nurses on a 3-point accuracy scale (correct, partially correct, and incorrect). The Chi-squared test and Fisher exact test with post-hoc Bonferroni adjustment were used to assess the differences in performance between the three groups, as well as the differences in accuracy between different prompts.
ChatGPT-4.0 showed significantly higher accuracy, with 64.3 % of responses rated as ‘correct’, compared to 60.5 % in ChatGPT-3.5 and 36.7 % in clinical nurses (X ^2^ = 74.192, p < .001). Except for the ‘care for the tracheostomy stoma and surrounding skin’ domain (X^2^ = 6.227, p = .156), scores from ChatGPT-3.5 and -4.0 were significantly better than nurses’ on domains related to airway humidification, cuff management, tracheostomy tube care, suction techniques, and management of complications. Overall, ChatGPT-4.0 consistently performed well in all domains, achieving over 50 % accuracy in each domain. Alterations to the prompt had no impact on the performance of ChatGPT-3.5 or -4.0.
ChatGPT may serve as a complementary medical information tool for patients and physicians to improve knowledge in tracheostomy care.
ChatGPT-4.0 can answer tracheostomy care questions better than most clinical nurses. There is no reason nurses should not be using it.
Keywords: Generative artificial intelligence, ChatGPT, Education, Tracheostomy, Nursing
Tracheostomy is a surgical procedure in which an incision in the anterior wall of the trachea is made in order to provide airway patency (Khanum et al., 2022). Poor management of the tracheostomy tube following the procedure is associated with serious risks, such as tracheal stenosis, tracheoesophageal fistula and infection, hemorrhage, and airway obstruction (McGrath et al., 2013). A survey from the National Confidential Enquiry into Patient Outcome and Death showed that there was high morbidity and mortality in patients living with a tracheostomy due to preventable complications (Wilkinson et al., 2015). Post-tracheostomy complications were common in clinical settings, with more than 21 % of patients experiencing complications while in intensive care and 24.3 % while on the medical or surgical ward (Wilkinson et al., 2015). This highlights the importance of tracheostomy care competency. Clinical nurses both on the ward and in the intensive care unit should be able to demonstrate competency on tracheostomy care knowledge and skills.
Tracheostomy care competency include airway humidification, wound care, stoma care, suction, and prevention and management of complications, such as airway obstruction and hemorrhage. For patients living with a tracheostomy tube and their family members, health professionals, particularly clinical nurses, are also responsible for educating them about tracheostomy care. Many patients and family caregivers continue to experience post-traumatic distress caused by having or taking care of a tracheostomy tube (Wang et al., 2022). Notably, patients and their caregivers are increasingly turning to online resources for health information. The Health Information National Trends Survey revealed that 74.9 % of Americans relied on the internet to find health-related information (Winston, 2021). Thus, it is imperative to ensure that the web resources used by healthcare providers and patients offer reliable, accurate, and easily comprehensible tracheostomy care instructions.
On November 30th, 2022, OpenAI launched ChatGPT, which is a text-based chatbot powered by a large language model and based on the Generative Pre-trained Transformer (GPT) architecture. A large language model, such as OpenAI's GPT-3.5 and GPT-4.0, refers to a type of artificial intelligence (AI) model that can understand and generate natural text responses. A report indicated that by August 2023, ChatGPT had an estimated 100 million active users, and it had the fastest-growing user base in history for a consumer application, with 1 million users registered in 5 days (Brandl and Ellis, 2023). ChatGPT has become a popular destination for users seeking information, assistance, and answers to their questions. The large language model - driven chatbot can handle a wide range of inquiries across various topics, making it a valuable tool for users to access information efficiently and effectively (OpenAI, 2023c). The GPT model has revolutionized the knowledge search technique by utilizing AI to generate conversational responses that are natural and contextually relevant. This is achieved through pre-training on large datasets and fine-tuning the model to understand context, grasp linguistic nuances, and generate human-like responses (OpenAI, 2023a). By leveraging the power of machine learning, these large language models have transformed the way users interact with search engines and other information retrieval systems, creating a more engaging and interactive experience. For instance, Microsoft announced on March 14th, 2023 that the new Bing search engine, Bing.com, had been running on GPT-4.0 for 5 weeks (Mehdi, 2023) and later also introduced the Microsoft 365 Copilot in Microsoft Apps, which allow users to give instructions using natural language prompts on March 16th (Spataro, 2023).
Prompt engineering is a technique to design effective questions or prompts to get desired information or responses from an individual or a system. For example, instead of asking, “How are you feeling?”, a nurse using prompt engineering would say, “On a scale of 1 to 10, how would you rate your pain level right now?” In a generative AI model, prompt engineering is the process of using a text-based instruction to describe how the AI model performs its tasks (Lawton, 2023). A prompt can be used to improve clarity, simplicity, context, specificity, and role-playing (Macready, 2023). For instance, to ask for a simple response, you can use the prompt of “answer the question using fewer than 10 words”, or to set the scope of the ChatGPT's response, a role-play prompt can be used, such as “Pretend you're a human resources officer working at Microsoft,” before asking it to generate three questions when someone is interviewing for a software engineer position. Effective prompt engineering is helpful to narrow the performance of the ChatGPT, which was originally built for general purposes.
ChatGPT has also sparked considerable interest in the medical community, with preliminary research yielding promising results, indicating its potential as a tool for clinical care assistance. Studies have shown that ChatGPT has successfully passed the United States Medical Licensing Exam and demonstrated satisfactory answers to banks of answers for specialty training oral exam questions (Ali et al., 2023a; Bernstein et al., 2023; Kung et al., 2023). However, these findings need to be interpreted with extreme caution. The large language model for GPT-4.0 was trained on data up to early 2023. All ChatGPT models have been trained on vast amounts of medical literature, making it capable of answering medical questions accurately up to early 2023.
There has been no evaluation of ChatGPT's performance on newly-developed survey or exam questions, which were outside of its training scope. In evaluating ChatGPT's effectiveness, limitations, and potential biases in exam performance, it would be more impartial to use a newly-developed survey comprising materials that were not included in its training dataset than established exam question banks. In 2023, we developed and validated the first comprehensive survey on examining nurses’ performance on tracheostomy care. The accuracy of the ChatGPT responses to tracheostomy care questions had yet to be determined. We hypothesized that the ChatGPT was capable of answering queries on tracheostomy care.
We aimed to first evaluate and compare the performance of two large language models, namely OpenAI's ChatGPT-3.5 and -4.0, with clinical nurses in responding to queries related to tracheostomy care. The second objective of our study was to determine if initial chatbot prompting affected the accuracy of answers generated by ChatGPT.
This was a cross-sectional study.
The participants were large language models GPT-3.5, GPT-4.0, and clinical nurses working in hospitals located in mainland China, who were recruited using a convenience sampling approach during the study period from August 25th, 2023, to October 6th, 2023.
The basic characteristic survey for clinical nurse participants included demographic data (age, birth sex, and education) and variables related to clinical experiences (years of working experience, professional certification, classification of hospital, and type of specialty nurse).
The authors developed and validated a tracheostomy care practice survey Chinese version by reviewing the clinical guidelines and consensus statements on tracheostomy care (Bodenham et al., 2018; Dawson, 2014; De Leyn et al., 2007; McGrath et al., 2012; Mitchell et al., 2013; Mussa et al., 2021; NSW Agency for Clinical Innovation, 2021; Rovira et al., 2021). A brief description of the 43 questions is shown in Table 1. The tracheostomy care practice survey comprises 43 items, including 16 multiple-choice questions and 27 select-all-that-apply questions. The survey was developed specifically to examine nurses’ knowledge of taking care of patients with tracheostomy tubes. The survey assessed six (1) airway humidification – 8 items, (2) cuff management – 8 items, (3) management of tracheostomy tube – 7 items, (4) care for the tracheostomy stoma and surrounding skin – 4 items, (5) suction technique – 9 items, and (6) prevention and management of common postoperative complications – 7 items. The content validity was evaluated by having nine nurse specialists with more than 10 years clinical working experience complete the survey, and the average content validity index was 0.922. The Cronbach's alpha values for domains of the tracheostomy care practice survey ranged from 0.852 to 0.983, indicating a satisfactory reliability of the measurement. Each response was graded using a 3-point accuracy 1 = correct, 2 = partially correct (correct but incomplete), and 3 = incorrect responses that were mixed with correct and incorrect answers or were completely incorrect answers.
We used the sample-to-item ratio for determining the sample size for a descriptive study (Memon et al., 2020). With a ratio of 5-to-1 and total of 43 items, the estimated sample size was 215.
The study described has been carried out in accordance with the Code of Ethics of the World Medical Association (Declaration of Helsinki). Ethical approval was obtained from the Beijing Civil Aviation General Hospital (IRB approval number 2023-L-K-03). We refrained from identifying the full name of the hospital where the clinical nurse participants were employed in order to uphold ethical standards. Instead, they were just required to provide the classification of the hospital.
The study survey was distributed using snowball sampling to nurses from five cities in China in the form of an electronic survey, via Qualtrics. The completion of the survey was voluntary and no incentive was provided to any of the study participant. Through personal networks, the authors solicited collaboration from four tertiary hospital nurses and asked them to distribute to their online work group in WeChat, the most commonly used social media in China. The study promotion posters were also shared and posted on their WeChat social media platform by the authors and invited nurses.
The conversation with ChatGPT was initiated by typing a survey question with or without a prompt directly into the chat. To better inform the ChatGPT what to do, users can use ‘prompt,’ which are text-based instructions written in natural language (Macready, 2023). We gave ChatGPT-3.5 and -4.0 three different ‘no prompt’, ‘patient-friendly prompt’, and ‘act-as-nurse prompt”, before being queried the 43 survey questions from the tracheostomy care practice survey. Each query led to one question-and-answer thread. Details of each prompt were described as follows.
The responses from clinical nurses were transferred into IBM SPSS Version 28.0.1.0 for scoring and further analysis. Percentages of correct responses provided by ChatGPT-3.5, ChatGPT-4.0, and clinical nurses were calculated independently for each category and the full survey according to the scoring criteria. Incomplete survey from 180 participates were not included in the analysis. Using the Chi-square test, differences in performance among the three groups were determined. Using the Chi-square test or Fisher exact test, the proportions of correct answers were compared by prompt types to determine whether initial ChatGPT prompting was associated with the differences in their grades. A post-hoc Bonferroni correction was applied to each possible pair-wise comparison. P-values were 2-tailed, and a p-value less than 0.05 was considered statistically significant.
This study evaluated the capacity to answer multiple-choice clinical practice questions on tracheostomy care by two large language models (ChatGPT-3.5 and ChatGPT-4) and clinical nurses in mainland China. Meanwhile, we also engineered the ChatGPT with three types of prompts before using them to generate answers. From ChatGPT, we collected a total of (43 questions x 3 prompts x 2 large language models =) 258 responses. From clinical nurses, as shown in Table 2, there were 453 participants who accessed the survey, and 272 completed the clinical practice questions on tracheostomy care. Most participants were from tertiary hospitals, and their clinical roles were nurses, senior nurses, and senior nurse managers or above. The majority were female, with bachelor's degree or above. Average working experience was 7.9 years with only 20.6 % of the participants working in the intensive care units or specialties of ear, nose throat or surgical otolaryngology.
Table 3 shows the numbers and percentages of the correct, partially correct and incorrect responses from ChatGPT-3.5 and 4.0 on three prompts, and clinical nurses. The chi-square test results showed that the differences in grading proportions among the three groups (ChatGPT-3.5, ChatGPT-4.0, and clinical nurses) were statistically significant. Post-hoc pairwise comparisons were performed using the Bonferroni correction to account for the results and it showed that in terms of accuracy, compared to ChatGPT-3.5 and nurses, ChatGPT-4.0 has a significantly higher correct rate, while the difference between the responses from ChatGPT-3.5 and -4.0 was not statistically significant.
Table 3 shows a detailed sub-analysis of the grading proportions across six tracheostomy care domains. With the exception of the 'airway humidification' domain, both ChatGPT models achieved an accuracy of above 50 % across all other tracheostomy care domains. In the domain of airway humidification, ChatGPT-4.0 exhibited marginally-superior performance compared to ChatGPT-3.5. However, nurses' performance in five out of the six domains fell below 50 % accuracy, with the exception of ‘care for the tracheostomy stoma and surrounding skin’. The Chi-Square test/Fisher Exact test results showed that, except for the ‘care for the tracheostomy stoma and surrounding skin ' domain, the difference in grading proportions in all other domains among three groups was statistically significant.
All post-hoc comparisons were conducted with Bonferroni corrections. In the domain of 'airway humidification', the proportions of correct responses from ChatGPT-3.5, ChatGPT-4.0, and clinical nurses were statistically significant. Statistical differences among responses from ChatGPT-3.5, ChatGPT-4.0, and clinical nurses were also observed in the domains of ‘cuff management', ‘management of tracheostomy tube after tracheotomy’, and ‘suction technique through tracheostomy site’, with ChatGPT-4.0 performing the best. Notably, in the domain of ‘prevention and management of common postoperative complications’, the post hoc analysis showed that there was a significant difference between the two ChatGPT models and clinical nurses on the proportions of correct responses and the proportions of partially correct responses, while there was no significant difference between the ChatGPT-3.5 and ChatGPT-4.0.
On Chi-Squared analysis, both ChatGPT-3.5 or ChatGPT-4.0 had no difference on performance among using different prompts. In individual domain of tracheostomy care, there was no significant variation on ChatGPT-3.5 and -4.0′s performance when using different prompts domain.
We evaluated the performance of ChatGPT-3.5, ChatGPT-4.0, and clinical nurses in responding to questions about tracheostomy care. To the best of our knowledge, this is the first study to assess the accuracy of responses to tracheostomy care queries generated by a large language model powered chatbot in comparison with clinical nurses, and the first to explore the utility of chatbots in clinical nursing application. We found that ChatGPT-4.0 has the potential to provide customized answers to questions about tracheostomy care.
Compared to ChatGPT, nurses performed relatively poorly at answering questions about tracheostomy care. Nurses showed higher percentages of “incorrect” ratings in all domains of tracheostomy care. ChatGPT-4.0 was more accurate in addressing tracheostomy care-related queries than both ChatGPT-3.5 and clinical nurses. This result confirmed earlier research by Ali et al. (2023b), Lim et al. (2023), and Raimondi et al. (2023) that highlighted ChatGPT-4.0′s superiority over other large language model chatbots in neurosurgery exams, myopia-related queries, and ophthalmology tests, respectively. Compared to previous models, ChatGPT-4.0′s advantages may be attributed to the larger training dataset, smarter algorithms, greater generalization, and stronger understanding on the context it receives (OpenAI, 2023b). It is critical to highlight that those answers from ChatGPT are not always accurate; this could be caused by limitations within its training database (Yeo et al., 2023), as no measure has been developed on assessing tracheostomy care.
Across the six question domains, it is notable that both ChatGPT models and clinical nurses exhibited good performance when addressing queries related to the prevention and management domain. This is different from a previous study by Lim et al. (2023), who showed ChatGPT's performance was good on topics of myopia pathogenesis, risk factors, clinical presentation, diagnosis, and prognosis, except for treatment and prevention. The difference could be explained by the fact that Lim et al.'s study required a chatbot to respond to open-ended questions, which allowed chatbots to provide more comprehensive responses based on its memory from large amount of medical literature, whereas our study asked chatbots to answer only multiple-choice questions from an established survey on tracheostomy care.
In addition, there was no statistically significant difference on ChatGPT's accuracy among different prompts in answering multiple-choice questions. This lack of difference is likely due to the fact that ChatGPT could not elaborate its responses while responding to multiple-choice format questions. With open-ended questions, ChatGPT can give as much information as possible by retrieving from its trained database. Our findings echo a prior study by Campbell et al. (2023) that found ChatGPT provided appropriate answers to most questions about obstructive sleep apnea regardless of prompting. Since many first-time users may not be aware that ChatGPT can be fine-tuned with different styles of initial prompts, it is reassuring to know that patients who query ChatGPT are likely to receive a clinically-appropriate response. However, all the tested prompts were simple phrases and did not alter the ChatGPT's performance. Prompts with more advanced instructions might lead to significant modification in its function.
The finding that ChatGPT-4.0 has the capacity to answer questions about tracheostomy care has clinical implications. Patients who are confused about their trach care at home could search Google or ask ChatGPT the same question. Google will give patients a list of references that they need to read through before getting to an ‘answer’, whether it is correct or incorrect, while ChatGPT can potentially do the same job within seconds and without overloading patients (Wang and Voss, 2022). ChatGPT could help promote the accessibility of health information by reducing the cognitive demand of searching for health information, as it generates targeted and personalized answers. With the advanced computing and memory power of generative AI/ChatGPT, ChatGPT could potentially be non-inferior in performing patient education on a specialized topic, compared to nurses who have not been trained on such a topic.
There are potential limitations of the study. First, the study focused specifically on tracheostomy care, so the findings may not be generalizable to other areas of healthcare or other medical topics. Second, the study included clinical nurses from mainland China only, which may not represent the broader population of clinical nurses worldwide. We did not collect the information on the clinical nurse participant's local protocol on tracheostomy care. Different regions and countries may have varying levels of knowledge, protocols, or experience in tracheostomy care. Third, the study used multiple-choice and select-all-that-apply survey questions for evaluating the performance of ChatGPT and clinical nurses, which may not capture the full range of their knowledge and understanding. Fourth, although the study examined the effect of different prompts on ChatGPTs’ accuracy, it may not have explored all possible prompt variations that could have influenced the AI's performance.
In this study, we have demonstrated the potential of ChatGPT-4.0 as a tool for providing customized answers to multiple-choice questions about tracheostomy care in a more efficient and personalized way, compared to traditional search engines. Further research is needed to explore the full potential of ChatGPT in the clinical nursing setting in answering open-ended questions and to assess the effectiveness of more advanced prompts in enhancing its performance.
During the preparation of this work, the author(s) used [QuillBot AI / Paraphrasing tool] in order to identity grammar issues. After using this tool/service, the author(s) reviewed and edited the content as needed and take(s) full responsibility for the content of the publication.
The study received no funding.
Tongyao Wang: Writing – review & editing, Writing – original draft, Visualization, Validation, Project administration, Methodology, Formal analysis, Data curation. Juan Mu: Writing – review & editing, Resources, Formal analysis, Data curation. Jialing Chen: Writing – review & editing, Writing – original draft, Formal analysis, Data curation. Chia-Chin Lin: Methodology, Resources, Supervision, Writing – review & editing.
We have no conflicts of interest to disclose