Authors: Yair Bannett (1Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California;), Fatma Gunturkun (2Stanford Quantitative Sciences Unit, Stanford, California;), Malvika Pillai (3Veterans Affairs Palo Alto Health Care System, Palo Alto, California;; 4Biomedical Informatics Research Center, Stanford University School of Medicine, Stanford, California;), Jessica E. Herrmann (5Stanford University School of Medicine, Stanford, California), Ingrid Luo (2Stanford Quantitative Sciences Unit, Stanford, California;), Lynne C. Huffman (1Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California;), Heidi M. Feldman (1Division of Developmental-Behavioral Pediatrics, Stanford University School of Medicine, Stanford, California;)
Categories: Article
Source: Pediatrics
Authors: Yair Bannett, Fatma Gunturkun, Malvika Pillai, Jessica E. Herrmann, Ingrid Luo, Lynne C. Huffman, Heidi M. Feldman
To assess the accuracy of a large language model (LLM) in measuring clinician adherence to practice guidelines for monitoring side effects after prescribing medications for children with attention-deficit/hyperactivity disorder (ADHD).
Retrospective population-based cohort study of electronic health records. Cohort included children aged 6 to 11 years with ADHD diagnosis and 2 or more ADHD medication encounters (stimulants or nonstimulants prescribed) between 2015 and 2022 in a community-based primary health care network (n = 1201). To identify documentation of side effects inquiry, we trained, tested, and deployed an open-source LLM (LLaMA) on all clinical notes from ADHD-related encounters (ADHD diagnosis or ADHD medication prescription), including in-clinic/telehealth and telephone encounters (n = 15 628 notes). Model performance was assessed using holdout and deployment test sets, compared with manual medical record review.
The LLaMA model accurately classified notes that contained side effects inquiry (sensitivity = 87.2, specificity = 86.3, area under curve = 0.93 on holdout test set). Analyses revealed no model bias in relation to patient sex or insurance. Mean age (SD) at first prescription was 8.8 (1.6) years; characteristics were mostly similar across patients with and without documented side effects inquiry. Rates of documented side effects inquiry were lower for telephone encounters than for in-clinic/telehealth encounters (51.9% vs 73.0%, P < .001). Side effects inquiry was documented in 61.4% of encounters after stimulant prescriptions and 48.5% of encounters after nonstimulant prescriptions (P = .041).
Deploying an LLM on a variable set of clinical notes, including telephone notes, offered scalable measurement of quality of care and uncovered opportunities to improve psychopharmacological medication management in primary care.
Accurate measurement of clinical practice is a necessary and critical component of improving health care quality and health outcomes in a learning health system.^1^ However, traditional methods for capturing clinical practice, such as medical record reviews, are time-consuming, labor-intensive, and unconducive to real-time improvement efforts.^2,3^ Large language models (LLMs) are a type of artificial intelligence (AI) that offers opportunities to capture clinical practice at scale by automatically analyzing free-text information from clinical notes in the electronic health records (EHRs). Here, we focus on a highly prevalent childhood condition—attention-deficit/hyperactivity disorder (ADHD)—which has long-standing clinical practice guidelines, as a test case for leveraging an LLM to assess quality of care.
ADHD is a prevalent neurodevelopmental disorder estimated to affect 10% of US children.^4^ Most children with ADHD are treated by their primary care pediatrician (PCP).^5,6^ PCPs frequently prescribe medications, including stimulants and nonstimulants, as part of the management of ADHD. Evidence-based clinical practice guidelines for primary care management of ADHD, published by the American Academy of Pediatrics (AAP), encourage PCPs to monitor benefits and side effects when prescribing medications.^7–9^ The few studies that have assessed PCP adherence to AAP guidelines for ADHD medication management have been limited in scope, analyzing available EHR structured data (eg, prescriptions, encounter dates) or performing labor-intensive medical record reviews on a small sample of patients.^10–12^
The lack of scalable access to essential components of clinical care that are documented as free text in the EHR has resulted in national quality metrics for ADHD—Healthcare Effectiveness Data and Information Set (HEDIS) measures—that capture only the frequency of in-office follow-up of children prescribed ADHD medications.^13^ The HEDIS metrics have no evidence base (ie, guidelines do not include a recommended frequency of follow-up), and they do not capture any guideline-based treatment recommendations, such as monitoring for medication side effects.^14^ These poor metrics have limited our ability to evaluate the impact of guideline-concordant care on patient outcomes. Furthermore, a focus on information abstracted only from in-office visits may significantly underestimate the frequency of encounters. Current ADHD medication management in primary care takes place in multiple communication routes with families, including through telehealth and telephone encounters and secure messaging systems, which have been shown to contribute to achieving symptom improvement.^15^ LLMs are well suited to synthesize this vast amount of clinical text and provide a comprehensive assessment of clinical care.
In this study, we aimed to assess the accuracy of an open-source LLM in measuring the extent to which primary care clinicians document inquiring about ADHD medication side effects in clinical notes from in-clinic/telehealth and telephone encounters. If successful, such a model can become an integral part of robust quality measurement, highlighting avenues for near-real-time improvement efforts and ultimately leading to improved patient outcomes.
We present our study in accordance with the MINimum Information for Medical AI Reporting (MINIMAR) frame-work and the Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) reporting guidelines.^16,17^ This study was approved by the Stanford University School of Medicine Institutional Review Board.
Packard Children’s Health Alliance (PCHA) is a community-based pediatric health care network in Northern California affiliated with Stanford Children’s Health and Lucile Packard Children’s Hospital. PCHA has 25 pediatric primary care clinics, grouped into 11 practices.
This was a retrospective population-based cohort study. We extracted structured and unstructured (free-text) EHR data (2015–2022) from 11 community primary care practices of all clinic/telehealth/telephone ADHD-related encounters for patients aged 6 to 11 years. We defined an ADHD-related encounter as an encounter with an ADHD visit diagnosis or with an ADHD medication prescribed, including stimulants (methylphenidate and amphetamines) and nonstimulants (alpha-agonists and atomoxetine). ADHD diagnoses and prescriptions for ADHD medications were identified based on codes and concepts from the Observational Medical Outcomes Partnership Common Data Model (OMOP CDM), which allows for common format and nomenclature across databases (see Supplemental Methods).
We established criteria for the study cohort to focus on documentation of medication side effects. The study cohort included patients with at least (1) 1 ADHD diagnosis, (2) 2 ADHD medication encounters (ie, encounters in which the PCP prescribed stimulants or nonstimulants), and (3) 1 available clinical note from an ADHD-related encounter that occurred within 3 months after an ADHD-related medication prescription. The final study cohort comprised 1201 patients (see Supplemental Figure 1 for study flowchart).
We used structured data in the EHR to describe the following patient patient age (at encounter of interest), sex, race/ethnicity (Asian/non-Hispanic, Black/non-Hispanic, Hispanic, white/non-Hispanic, other/non-Hispanic, unknown), and medical insurance at first ADHD medication encounter (private/public).
We extracted clinical notes from all in-clinic, telehealth, and telephone ADHD-related encounters conducted after an ADHD medication was prescribed (4486 in-clinic/telehealth notes and 11 142 telephone notes; total notes, n = 15 628).
To create a “ground truth” for documentation of side effects inquiry, we developed annotation guidelines used by 2 clinicians (Y.B. and J.E.H.) who performed independent medical record review and annotation of a sample of clinical notes using the CLAMP software.^18^ We sampled and annotated 2 to 26 clinical notes per patient from medication encounters (in-clinic/telehealth and telephone encounters) for a sample of 119 patients (n = 501 notes). Inter-annotator agreement (IAA) was assessed using Cohen’s kappa statistic.^19^ After confirming high agreement (IAA = 0.86) for the first 84 notes, annotators divided the annotation of the remaining notes (n = 417).
Eligible encounters included ADHD-related encounters occurring within 3 months after any ADHD medication prescription, a timeframe in which monitoring of medication side effects is expected. For each eligible encounter, we assessed the presence or absence of side effects inquiry in all associated clinical notes. Side effects inquiry was considered present if any note from the encounter included mentions of medication side effects, whether reporting their absence (eg, “no weight loss”) or presence (eg, “reduced appetite”).
We developed a binary classification pipeline based on the open-source Large Language Model Meta AI, trained on 13 billion parameters (LLaMA13B), to classify notes as containing or not containing documentation of side effects inquiry (see Supplemental Figure 2 and Supplemental Methods describing data extraction and model architecture). Figure 1 illustrates the model training and deployment workflow. The annotated set of 501 notes was subdivided with an 20 ratio into train (n = 411) and holdout test (n = 90) sets. The train set was used for model development and hyperparameter tuning, while the holdout test set was set aside to evaluate model performance compared with ground truth labels using various metrics including sensitivity, specificity, and area under the receiver operating characteristic curve (AUC). Model thresholds were selected to maximize sensitivity and minimize the false negative rate on the training set because we wanted to avoid wrongly classifying notes to suggest clinicians were not adhering to practice guidelines when they did adhere.
We report 95% CIs for AUC, which captures the discrimination power of the model and is not influenced by threshold selection. After confirming acceptable model performance on the holdout test set, we deployed the model on the remaining unannotated notes from all ADHD-related encounters for the study cohort (deployment notes, n = 15 127). We then sampled and annotated 363 notes (IAA = 0.93) to assess model performance in the deployment test set. The code is available at https://github.com/ybannett/NLP_ADHD_PTBM.
For the error analysis, misclassified notes in the holdout and deployment test sets were reviewed to understand model errors and potential reasoning behind misclassifications. For the fairness analysis, we used a combined dataset including holdout and deployment test sets and applied the classification parity approach to assess whether model outcomes are roughly equal across several patient subgroups (patient sex and insurance type).^20^ The fairness analysis was conducted to rule out model bias that may perpetuate disparities in care, if such disparities exist.^21,22^ We did not examine race/ethnicity data in the fairness analysis due to a large percentage of missing data and colinearity between insurance type and race/ethnicity in our data.^23^
Continuous variables were summarized by mean and SD, and categorical variables were presented as counts and percentages. The balance of demographic/clinical characteristics between patients with and without documentation of side effects inquiry was assessed using absolute standardized differences (ASDs). ASD values of 0.2, 0.5, and 0.8 correspond to small, medium, and large differences between the groups, respectively. The intraclass correlation coefficient (ICC) quantified the proportion of total variance attributable to differences across and within practices. A low ICC (<0.5) indicates high variability. Clinicians who saw fewer than 5 patients with ADHD were excluded to ensure reliable estimates. Generalized linear mixed-effects models were used to analyze differences in SE inquiry rates between (1) telephone encounters and in-clinic/telehealth encounters and (2) encounters associated with stimulant prescriptions, nonstimulant prescriptions, or both. Patient ID was included as a random effect to estimate the overall effects of encounter type or prescription type on SE inquiry rates while adjusting for potential correlations between encounters from the same patient. Only race/ethnicity had missing data (28.8%). All analyses were conducted using Python version 3.11.5.
The LLaMA model achieved excellent performance in classifying notes that contain side effects inquiry, as compared with ground truth labels. In the holdout test set (n = 90 notes), the model achieved sensitivity of 87.2%, specificity of 86.3%, and an AUC of 0.93 (95% CI, 0.88–0.99). In the deployment test set (n = 363 notes), the model achieved sensitivity of 88.0%, specificity of 90.1%, and an AUC of 0.92 (95% CI, 0.89–0.94).
We investigated potential reasons for model misclassifications. False positive classifications included documentations that were not clearly related to the prescribed ADHD medication (eg, review of systems) or documented side effects from other prescribed medications (eg, acne medication). False negative classifications included abbreviated documentation (eg, “follow up add meds weight loss”) or documentation of nonspecific symptoms (eg, “giggly and disruptive on new med”).
The model showed consistent performance (presented with AUC and 95% CI) across patient subgroups, including patient sex and insurance type (Figure 2).
The study cohort comprised 1201 patients aged 6 to 11 years who had an ADHD diagnosis and at least 2 ADHD-related medications prescribed by their PCP. Table 1 presents patient characteristics stratified by documented side effects inquiry (yes/no), as classified by our model. The median age at first ADHD diagnosis and first ADHD-related medication prescription was 9 years. Of 1201 patients, 73.8% (n = 886) were male, and 73.9% (n = 888) were privately insured. The cohort primarily consisted of white/non-Hispanic patients (n = 532, 44.3%) or unknown race/ethnicity patients (n = 346, 28.8%). Side effects inquiry was documented for 86.0% (n = 1033) of patients within 3 months of at least 1 ADHD-related medication prescription. Patients with side effects inquiry had a slightly higher median number of ADHD-related encounters (6 vs 4 encounters). There were minimal differences between the groups for patient sex, race/ethnicity, insurance type, and observation window. Figure 3 illustrates the time to side effects inquiry after each patient’s first medication prescription. In half of the patients, PCPs documented inquiring about medication side effects within less than 2 months of the first prescription and, in 65% of patients, within 3 months of the first prescription.
Overall, side effects inquiry was documented in 60.0% (n = 5981) of ADHD-related encounters within 3 months of a medication prescription and was highly variable across primary care practices (range, 23.4%–78.6% of encounters; ICC = 0.4). Variation in documentation of side effects inquiry was higher across clinicians within practices than across practices (clinician-level ICC = 0.11, Supplemental Figure 3). More than 50% of ADHD-related encounters were completed by telephone in 7 of 11 practices (Figure 4). However, the proportion of telephone encounters with documented side effects inquiry was significantly lower at 51.9% compared with 73.1% for in-clinic/telehealth encounters (P < .001), with relatively high inquiry rates (49%–55% of telephone encounters) observed in only 2 practices. When examining encounters by the type of ADHD medication prescribed, side effects inquiry was documented in 61.4% (n = 5532) of ADHD-related encounters within 3 months after stimulants were prescribed compared with 48.6% (n = 202) of encounters after nonstimulants and 46.8% (n = 247) of encounters after both stimulants and nonstimulants were prescribed (P = .041 and P < .001, respectively; Figure 5).
In this study, we have shown that deploying an LLM on different types of clinical notes can effectively assess adherence to AAP practice guidelines for 1 aspect of medication management of children with ADHD. Fairness analysis did not indicate any evidence of model bias. By deploying this model, we uncovered specific targets for improvement in PCP medication management, including underutilization of telephone encounters to monitor medication side effects, and low rates of side effects management when prescribing nonstimulants. After external validation in other health organizations, this novel approach can be used to assess clinical care at scale and to inform quality improvement efforts for a wide range of medical conditions.
The high model performance rates established in this study demonstrate that LLMs offer an accurate and reliable method of analyzing the vast amount of unstructured EHR data that represents the documentation trail of current pediatric care, which includes significant between-visit management of chronic conditions (eg, telephone encounters, secure messaging). Our findings are promising in that LLMs allow efficient data mining and interpretation of large volumes of clinical data that, thus far, have been inaccessible through existing methodologies for medical record data extraction. Additionally, the automaticity and speed of LLMs offer the opportunity to develop near-real-time feedback through dashboards for clinicians and health organizations on their care—an effective method to improve clinical practice.^24^
When applying an LLM to assess documentation of clinical care, it is critical to assess AI model bias, given its potential to perpetuate health disparities.^21,22^ We therefore incorporated a fairness analysis into our study. In this case, we confirmed that the model performance did not differ significantly across patient subgroups, including patient sex and insurance type.
Several targets for quality improvement were uncovered by deploying the LLM on all clinical notes. While all practices had high utilization of telephone encounters for ADHD medication management, only 2 of the 11 practices used telephone encounters to regularly inquire about medication side effects. This finding serves as a learning opportunity for other practices. A simple intervention that adopts a standardized medication refill form, currently used in the 2 practices that documented side effects inquiry in approximately 50% of telephone encounters, can positively impact the quality of medication management. Another finding was the lower rates of side effects inquiry in children prescribed nonstimulants, as compared with stimulants. Here too, an intervention that targets potential knowledge gaps related to side effects profile of these less frequently prescribed medications can promote high-quality care. The high variation in care identified across and within practices presents valuable information when planning and evaluating quality improvement efforts at the practice and the clinician level.
This study complements our previous study that focused on clinician adherence to AAP guidelines in recommending nonpharmacological behavioral treatment for young children with ADHD, in which we demonstrated the use of LLMs in the evaluation of quality of care.^25^ These 2 studies are the first, to our knowledge, that provide objective support for the successful use of AI to provide a comprehensive evaluation of ADHD management, a prevalent neurobehavioral condition that is predominantly managed in primary care. This novel application of AI to assess quality of care presents an opportunity to overcome long-standing barriers to data extraction that have thus far prevented a comprehensive assessment of clinical care provided for children with ADHD and the extent to which variation in care affects patient outcomes. Future implementation of such algorithms could allow primary care clinicians to receive feedback on between-visit management (eg, telephone encounters for medication management) and to receive real-time decision support (eg, prompts for inquiring about specific side effects)—2 areas of need in primary care that have been recently identified as areas that could benefit from AI.^26^
Our cohort identification relied on multiple medication prescriptions with a primary indication for ADHD. Although this approach facilitates standardized implementation in any health care system, it introduces some level of misclassification error. Our assessment of quality of care in this study relies on clinician documentation in the EHR. It is possible that, in some cases, clinicians inquired about side effects without documenting their inquiries in the EHR. Furthermore, our study focused on medication management of ADHD by PCPs, and we did not have information on prescriptions provided to children outside the examined network (eg, by child psychiatrists). We are currently expanding this approach to include patient vital signs data (ie, weight, blood pressure, pulse) and documentation of patient care through the EHR secure messaging system, which were not available in the dataset created for this study. Because we were interested in measuring clinician adherence to guidelines, our outcome included any mention of side effects inquiry; we did not have information on rates of side effects presence and absence. We are currently examining the use of another model to answer this separate research question and use case. Finally, although the study was conducted in a large network of primary care practices with a diverse population (including 15% Hispanic, similar to the US census), its generalizability and the accuracy of model classifications need to be assessed in other health care networks.
Given the high prevalence of ADHD and the potential harm that can be caused to children and adolescents if psychopharmacological medication side effects are not considered and documented as part of clinical care, this study carries significant implications. By leveraging recent advances in the ability of LLMs to accurately classify large bodies of text, our novel approach offers an efficient way to assess the quality of ADHD medication management. Such rapid transformation of clinical data into knowledge can foster learning health systems, in which clinicians and health care organizations receive timely and actionable feedback on the care that they provide. After replication of this approach in other health care systems, developed LLMs can be made available for use across health organizations and contribute to enhancing evidence-based and equitable health care delivery for children with ADHD and other medical conditions, ultimately improving patient outcomes.