Authors: Neguine Rezaii, Daisy Hochberg, Megan Quimby, Bonnie Wong, Michael Brickhouse, Alexandra Touroutoglou, Bradford C Dickerson, Phillip Wolff
Categories: Original Article, data-driven classification, generative artificial intelligence, primary progressive aphasia, theory-driven classification, verb frequency, AcademicSubjects/MED00310, AcademicSubjects/SCI01870
Source: Brain
Neurodegenerative dementia syndromes, such as primary progressive aphasias (PPA), have traditionally been diagnosed based, in part, on verbal and non-verbal cognitive profiles. Debate continues about whether PPA is best divided into three variants and regarding the most distinctive linguistic features for classifying PPA variants.
In this cross-sectional study, we initially harnessed the capabilities of artificial intelligence and natural language processing to perform unsupervised classification of short, connected speech samples from 78 pateints with PPA. We then used natural language processing to identify linguistic features that best dissociate the three PPA variants.
Large language models discerned three distinct PPA clusters, with 88.5% agreement with independent clinical diagnoses. Patterns of cortical atrophy of three data-driven clusters corresponded to the localization in the clinical diagnostic criteria. In the subsequent supervised classification, 17 distinctive features emerged, including the observation that separating verbs into high- and low-frequency types significantly improved classification accuracy. Using these linguistic features derived from the analysis of short, connected speech samples, we developed a classifier that achieved 97.9% accuracy in classifying the four groups (three PPA variants and healthy controls).
The data-driven section of this study showcases the ability of large language models to find natural partitioning in the speech of patients with PPA consistent with conventional variants. In addition, the work identifies a robust set of language features indicative of each PPA variant, emphasizing the significance of dividing verbs into high- and low-frequency categories. Beyond improving diagnostic accuracy, these findings enhance our understanding of the neurobiology of language processing.
Keywords: generative artificial intelligence, primary progressive aphasia, data-driven classification, theory-driven classification, verb frequency
See Hillis (https://doi.org/10.1093/brain/awae242) for a scientific commentary on this article.
Language is a vital faculty through which we share our thoughts and feelings, build relationships and pass on collective knowledge. When this faculty falters, challenges arise, as seen in those with primary progressive aphasia (PPA). PPA is a neurological disorder characterized by the gradual erosion of language abilities while initially leaving other cognitive, affective and sensorimotor functions largely spared.^1^ The specific characteristics of the condition vary among individuals but generally fall into one of three non-fluent variant PPA (nfvPPA), characterized by agrammatism and/or effortful, halting speech; semantic variant PPA (svPPA), marked by difficulties in confrontational naming and single-word comprehension; and logopenic variant PPA (lvPPA), distinguished by core deficits in word retrieval and sentence repetition.^2^ Patients are typically classified after a comprehensive clinical assessment, including a battery of confrontational language tests. Despite the widespread use of these diagnostic constructs, debate continues about the specific features of each subtype and their distinctiveness.^3^ Some core features, such as agrammatism, are not well defined or easily quantified,^4^ which is one reason why a sizeable number of patients are classified as ‘mixed’ PPA.^5-9^ Another criticism concerns the nature of the categories themselves and the extent to which they reflect the natural partitioning of language abnormalities of PPA. In addition, there is an alternative hypothesis that patients with PPA, like those with post-stroke aphasia, exhibit a multidimensional spectrum of impairments, with only those at the extremes of certain dimensions being categorically distinct.^10^ Finally, we propose as an additional challenge that the classification of PPA patients into subtypes might depend, at least in part, on the specific types of confrontational tests included in the assessment battery and might not fully reflect linguistic impairments in naturalistic communication.
We performed this study with two major goals. First, we sought to test the hypothesis that the three major PPA variants represent natural categories of aphasic subtypes detectable in the connected speech of patients. We pursued a data-driven approach to discovering natural categories of aphasic subtypes using generative artificial intelligence (AI) large language models (LLMs). Rather than determining whether a neural network can be trained to predict the existing categories of PPA (a process that reinforces prespecified diagnostic constructs), we investigate whether the classic PPA variants emerge from natural correlations within the features of people’s speech. Given that LLMs can process language at multiple levels of abstraction, from the syntactic to the conceptual, they do not need the linguistic features to be prespecified or otherwise coded. The strengths of LLMs can be adapted to discover similarities and differences in the speech of a sample of patients. Thus, using this form of generative AI, clusters of patients with similar language characteristics could be discovered from naturalistic speech samples without the biases inherent in the selection of tests in an assessment battery or the lenses through which clinicians interpret results. Once the clusters of similar patients were discovered, we sought to validate them biologically by analysing the regional atrophy pattern measured with MRI scans for each cluster.
Our second goal was to identify the linguistic features associated with each category of similar PPA patients. Complementing the capabilities of LLMs, we used an automated parser to identify linguistic features that are used less frequently by each of the PPA variants relative to cognitively unimpaired controls. The value of these linguistic features can be tested by determining the degree to which their impairment predicts the different variants of PPA. Beyond revealing linguistic features used to classify patients with PPA, such an analysis might also provide new insights into the nature of the language system and its breakdown in aphasia. Feature-based classification of PPA has been used in prior research. For example, Fraser et al.^11^ used three different classifiers on a large set of lexicosemantic features. They found that such features could be used by a naïve Bayes classifier to classify patients as either svPPA or nfvPPA with 79% accuracy. Likewise, Themistocleous et al.^12^ used a deep neural network on linguistic features derived from picture descriptions to distinguish all three PPA subtypes with 80% accuracy. These models highlight how features from a syntactic parser can be effective in determining different types of PPA, but also their limitations. Improving the performance of PPA classifiers might require moving beyond the types of features provided in a syntactic parser.
Therefore, we also sought to investigate one specific linguistic feature that has received minimal attention in PPA research, which is the category of verbs. Prior research has shown that verb comprehension predominantly engages prefrontal cortical areas.^13-18^ These results are congruent with findings in patients with nfvPPA, who showed deficits in both verb comprehension and verb production, which was mainly assessed through confrontational naming tests.^19-21^ However, verb meanings also activate the left temporoparietal junction, encompassing the posterior lateral temporal cortex and inferior parietal lobule.^22-28^ Although patients with lvPPA exhibit cortical degeneration in the left temporoparietal junction, verb processing deficits have not been reported in this variant.^19^ Studies of patients with post-stroke aphasia demonstrate that individuals with left temporoparietal lesions tend to use high-frequency (mainly ‘light’) verbs. In contrast, those with prefrontal lesions produce more low-frequency (mostly ‘heavy’) verbs.^29-34^ We have reported a similar finding in nfvPPA patients.^35^ With this background, we also tested whether separating high- and low-frequency verbs would improve the performance of the model in classifying PPA variants, hypothesizing that patients with nfvPPA would use more low-frequency verbs than those with lvPPA and svPPA.
Seventy-eight patients with PPA were recruited from an ongoing longitudinal study at the Primary Progressive Aphasia Program in the Frontotemporal Disorders Unit of Massachusetts General Hospital (MGH). All patients underwent a standard clinical evaluation, comprising a structured history obtained from both patient and informant, comprehensive medical, neurological and psychiatric history and examinations, neuropsychological and speech–language assessments and a clinical brain MRI scan.^36^ Ratings on our language scale, called the Progressive Aphasia Severity Scale (PASS), were also included.^37^ PASS uses the best judgment of the clinician and integrates information from the test performance of the patient and an interview with a companion and includes ‘boxes’ for fluency, syntax, word retrieval and expression, repetition, auditory comprehension, single-word comprehension, reading, writing and functional communication. The PASS Sum-of-Boxes (SoB) is the sum of the box scores. The clinical and demographic information on the patients is shown in Table 1.
In addition, we used data from 20 healthy, cognitively unimpaired older adult controls from the Speech and Feeding Disorders Laboratory at the MGH Institute of Health Professions. These individuals had an average age of 65.2 years and an average years of education of 15.8. Fifty per cent of healthy controls were female, and 75% were right-handed. All study participants provided informed consent in accordance with guidelines established by the Mass General Brigham Healthcare System Institutional Review Boards, which govern human subjects research at MGH and specifically approved this study.
The participants were asked to view a drawing of a family at a picnic from the Western Aphasia Battery–Revised^38^ and describe it using as many full sentences as they could. Responses were audio-recorded and later transcribed into text using the Microsoft Dictate application. The transcriptions were then checked manually for accuracy by a research collaborator blind to patient characteristics.
The analyses used the 11-billion-parameter version of T5, which is available through the Huggingface transformer library.^39^ The T0 LLM available through the Huggingface transformer library was used to generate guesses about possible sentences.^39^ Processing was conducted on a Microway Compute Server with four NVIDIA Ampere A100 80 GB GPUs. To measure language similarities across PPA patients, we used a two-step algorithm to implement this strategy. Initially, an LLM named T0 (T-Zero)^40^ predicted each sentence in a patient’s language sample. This prediction was based on the sentence immediately preceding it, combined with all sentences produced by the other patients. This procedure was repeated 10 times. For each sentence, the highest similarity score generated by a second LLM, T5 (T-five),^41^ determined the sentence similarity of the patients. We selected T5 for its pretraining in generating semantic similarity scores on a scale from one to five. We calculated the overall similarity between two patients by averaging these scores across all sentences. This process resulted in a 78 × 78 matrix, representing the similarities among the 78 patients studied. The process of discovering clusters was accomplished using hierarchical clustering. The matrix of similarities generated by T5 was converted into a pairwise dissimilarity matrix using the Kendall rank correlation coefficient. Hierarchical clustering was conducted in R using the AGNES agglomerative clustering,^42^ and the ‘flexible’ weight average linkage approach, with the par.method set to 0.8. The dissimilarity matrix was generated using Pearson correlations, computed using the get_dist() function from the ‘factoextra’ library. The rect.hclust function from the ‘cluster’ package in R was used to cut the hierarchical clustering solution such that the participants were partitioned into one to seven clusters.
Neuroimaging data were available for 66 of the 78 PPA patients. A control sample was used as a reference for quantifying the magnitude of atrophy in our PPA patients. The sample included 24 cognitively normal older adults as control participants, also recruited at MGH (mean age 67.4 ± 4.9 years; 12 females; mean education 15.7 ± 2.3 years). These control participants underwent a neurological and cognitive assessment to confirm the absence of a medical history of neurological or psychiatric conditions, a structured interview of the participant and an informant by a neurologist, neurological examination and a neuropsychological test battery (UDS 3.0), and were determined to be clinically normal with the Clinical Dementia Rating scale (CDR = 0). All controls had normal brain structure based on MRI and low cerebral amyloid based on quantitative analysis of C-labelled Pittsburgh Compound-B (^11^C-PiB) PET data [frontal, lateral, and retrosplenial regions distribution volume ratio (FLR DVR) < 1.2]. MRI data were collected from participants on a Siemens 3 T MAGNETOM Tim Trio scanner using a 12-channel phased-array head coil. Structural MRI data were acquired using a T1-weighted magnetization-prepared rapid gradient echo (MPRAGE) sequence with the following repetition time = 2530 ms, time to echo = 3.48 ms, flip angle = 7°, number of interleaved sagittal slices = 176, field of view = 256 mm and voxel size = 1 mm isotropic. The MPRAGE data of each participant underwent intensity normalization, skull stripping and automated segmentation of cerebral white matter to locate the grey–white boundary via FreeSurfer v.6.0, which is documented and freely available for download online (http://surfer.nmr.mgh.harvard.edu). Defects in the surface topology were corrected,^43^ and the grey–white boundary was deformed outwards using an algorithm designed to obtain an explicit representation of the pial surface. All cortical surface derivatives were inspected visually for technical accuracy and were edited manually when necessary. Cortical thickness was calculated as the closest distance from the grey–white boundary to the grey–CSF boundary at each vertex on the tessellated surface. For each cluster of patients, we performed a whole-cortex, vertex-wise analysis of cortical thickness compared with controls to identify the spatial topography of atrophy in each patient cluster. For this analysis, we registered the thickness data of all participants to fsaverage space and smoothed them geodesically with a full-width half-maximum of 10 mm. The results of these analyses were inspected via maps of statistical significance at each vertex point overlaid on the average cortical surface template. For these exploratory analyses, a statistical threshold of P < 0.01 was used.
The transcribed speech samples of the patients and healthy controls were analysed for linguistic features using the Stanza natural language-processing toolkit.^44^
In addition to the features extracted automatically from the parser, we separated high- and low-frequency verbs based on the logic discussed in the Introduction. We used three different corpora of spoken the Switchboard Dialog Act Corpus,^45^ the Santa Barbara Corpus of Spoken American English^46^ and the Corpus of Contemporary American English (COCA).^47^
Counting the number of occurrences for each language feature requires normalization. Such analyses need to control for the overall size of the language sample and for basic differences in the relative frequency of certain linguistic features. For example, nouns are produced far more frequently than adjectives or number words in ordinary speech in healthy individuals. Hence, determining whether a particular linguistic feature occurs less often than expected in an individual with PPA requires that the frequency of the linguistic feature be assessed relative to benchmarks established in controls. This analysis allows us to assess directly the degree to which the frequencies differ from normal controls, not only from the other variants. The number of occurrences of each linguistic feature is calculated relative to the frequencies observed in healthy controls (for details of the calculation, see Supplementary Material 1).
To assess whether linguistic features were in deficit for particular variants of PPA, three dummy variables were created for each variant. For the nfvPPA dummy variable, patients were coded with ‘1’ if they belonged to the nfvPPA cluster and ‘0’ otherwise. The same approach was taken for lvPPA and svPPA. The classic approach to adjust the significance level for multiple statistical tests is the Bonferroni method, which is α/C, where C is the number of comparisons. This method is widely used for problems with independent multiple comparisons. However, in situations in which the predictors are correlated, the Simes method can be used. In this method, all C P-values are sorted from smallest to largest. An adjusted significance level is calculated for each position of the list based on αk/C, where α is the alpha level (e.g. 0.05) and k is the position in the ordered list. Correcting for multiple comparisons using the Simes method indicated that correlations were significant for P < 0.01278. Table 2 contains all negative correlations that met the Simes correction cut-off.
The phylogenetic tree was produced using the fviz_dend() function available in the factoextra library in R by setting the type argument to ‘phylogenic’, the k argument to ‘3’, and the phylo_layout argument to ‘layout.gem’. In such graphs, the distance between the vertices is less important than how the vertices are connected.^48^
The point-biserial correlations demonstrate that a wide range of linguistic features are associated with the different variants of PPA. Predictive modelling was conducted using multinomial logistic regression based on the multinom function from the nnet package in R. The regression included four nfvPPA, lvPPA, svPPA and healthy control. The input was the raw counts indicating the number of times a patient produced a particular linguistic feature. The healthy controls served as the baseline reference category in the model.
To protect against over-fitting, we used 10-fold cross-validation over 100 repetitions. The performance of the model is the average accuracy across the folds and repetitions. Cross-validation was conducted using the trainControl function in the caret package in R.
Using data from 78 PPA patients, we first evaluated whether the three PPA variant classification system has external validity by using LLMs to uncover implicit divisions in the ways that patients speak when describing the Western Aphasia Battery Picnic Scene (WAB-PS).^38^ Using LLMs on the language samples of patients resulted in a 78 × 78 matrix, representing the similarities among the 78 patients studied. Next, we analysed this matrix for similar clusters of patients using AGNES hierarchical clustering, which constructs a hierarchy of clusters in a bottom-up manner.^42^ A dissimilarity matrix was generated using Pearson correlations. Figure 1 shows the solution that resulted from applying hierarchical clustering to the dissimilarity matrix. The label for each participant represents the PPA variant of the individual as diagnosed independently by expert clinicians using the comprehensive assessment. The clusters generated from this data-driven analysis of connected speech samples agreed with the clinical diagnoses by experts regarding the variant of PPA in 88.5% of cases.
Figure 1 Hierarchical clustering solution of language samples from primary progressive aphasia participants based on text similarity scores and their neuroanatomical associates. Participant labels indicate the primary progressive aphasia (PPA) variant as determined by comprehensive clinical assessments. The first division separates a group primarily composed of nfvPPA patients from the remaining patients, and the second division separates a group of mostly svPPA patients from a group of primarily lvPPA patients. The inset shows the total within-cluster sum-of-squares error as a function of the number of clusters, which demonstrates that clusters beyond three offer minimal reduction in error, implying that the optimal number of clusters is three. The total within-cluster sum of squares was based on the dissimilarity matrix used to generate the hierarchical clustering solution. Cortical surface maps indicate areas where each of the three clusters of PPA patients had significantly greater atrophy than age-matched controls, largely recapitulating the atrophy patterns typically seen in each variant. All between-group comparisons are significant at P < 0.01. lvPPA = logopenic variant PPA; nfvPPA = non-fluent variant PPA; PPA = primary progressive aphasia; svPPA = semantic variant PPA.
Participants were partitioned into one to seven clusters from this hierarchical clustering solution. The total within-cluster sum-of-squares error (SSE) for each partition was calculated from the dissimilarity matrix used to generate the hierarchical clusters. The total SSEs for the different numbers of clusters are plotted in the inset of Fig. 1. The SSEs declined rapidly as the number of clusters increased from one to three, then levelled off, implying a three-cluster solution according to the elbow criterion.^49^
We examined the biological validity of the data-driven clusters of PPA patients identified above using an analysis of regional cortical atrophy of each of the three clusters of patients compared with an age-matched group of control participants (Fig. 1). Patients in the first cluster, mostly made up of svPPA cases, exhibited prominent bilateral (left more than right) atrophy in the temporal pole, inferior, middle and superior temporal regions, insula and orbitofrontal cortex. Patients in the second cluster, largely composed of lvPPA cases, exhibited asymmetrical atrophy in the left anterior temporal cortex, inferior parietal lobule and the superior and middle temporal gyri, with weaker effects in the dorsolateral prefrontal cortex and inferior frontal gyrus. Patients in the third cluster, largely composed of nfvPPA cases, exhibited atrophy in the dorsolateral prefrontal cortex (left more than right), bilateral dorsomedial and mid-cingulate cortex, left inferior frontal gyrus and left inferior temporal gyrus. These atrophy patterns of clusters of patients segregated by the data-driven LLM analysis closely recapitulate the localization of cortical atrophy, which is well established for each of these variants using comprehensive clinical evaluations.
We demonstrated above that an unsupervised AI method using LLMs could differentiate the three major PPA variants from a brief connected speech sample in a manner highly consistent with that of expert clinicians. The findings imply that much of the information needed for classifying variants is present in people’s connected speech. Although successful, this data-driven analysis does not specify the information used by the LLM to compute these similarities. To address this limitation, we used natural language processing methods to identify the linguistic features in the connected speech samples that best distinguish the PPA variants. We performed a clustering analysis on the linguistic features to see whether they provided convergent support for the clusters that emerged from the LLM analysis.
As one hypothesized language feature, we divided verbs into high- and low-frequency groups. Given that there is no single agreed-upon list of heavy and light verbs, we adopted a quantitative approach to detect a potential natural division in verb distribution using these Switchboard, the Santa Barbara Corpus of Spoken American English and COCA. The frequencies from the top 966 most frequent verbs from the COCA, which contained the highest number of verbs compared with other corpora, were converted into a distance matrix by computing the absolute difference in frequency between each verb and every other verb. The resulting distance matrix was submitted to the scikit-learn 1.3.0 k-means clustering after setting the number of clusters from 2 to 50.^50^ The most frequently occurring cluster contained the first 13 ‘be’, ‘have’, ‘do’, ‘go’, ‘say’, ‘know’, ‘get’, ‘think’, ‘see’, ‘come’, ‘want’, ‘make’ and ‘take’ (Fig. 2). Interestingly, these 13 verbs were among the most frequent verbs in the other two corpora (Supplemental Material 2). Despite being only a small group, these high-frequency verbs made up 56% of the verbs produced by patients and 54% of the verbs produced by healthy subjects in this study. In this analysis, these top 13 verbs were classified as high frequency, whereas all other verbs were categorized as low-frequency verbs.
Figure 2 The frequency of the 50 most common verbs in spoken English, according to the corpus of contemporary American English. The distribution is divided into high-frequency and low-frequency verbs. The frequency of the verb ‘be’ (7 025 941) is so high that it is cut off at the frequency of the second highest verb, ‘have’.
Additional linguistic features were identified using the Stanza natural language processing toolkit.^44^ The Stanza parser identified 103 language features. Of these, we retained features with at least three observations for further analysis, resulting in a list of 84 features. These retained features consisted of part-of-speech (n = 27, e.g. noun and preposition), syntactic relationships expressed in universal dependencies (n = 34, e.g. nominal case and clausal complement) and morphosyntactic features (n = 23, e.g. past tense and possessive).
To capture language deficits, we reduced the 84 features to those that appeared significantly less frequently than in the control group for at least one of the variants. This was achieved by combining the frequency counts across the patients in the three variants, resulting in a 3 × 84 (group × linguistic feature) matrix. We applied the extended version of Fisher’s exact test to a 3 (groups) × 84 (features) table to determine the probability of the residuals.^51^ The method indicated that 52 of the original 84 features were associated with significantly negative residuals for at least one of the variants, with family-wise error correction managed by the Simes method.^52^ As a result, the original matrix was reduced further to a 78 (patient) × 52 (feature) frequency matrix. The counts in the 78 × 52 matrix were compared against expected values based on the healthy controls to generate residuals. We retained only the negative residuals in the 78 × 52 matrix and set the positive residuals to zero. Transposing the matrix allowed us to assess all pairwise dissimilarities with respect to Pearson correlations across participants. The dissimilarity matrix was subsequently analysed using agglomerative (AGNES) hierarchical clustering with the flexible weight average linking approach, with the par.method set to 0.6 and assuming one to seven clusters. The total within-cluster SSE was calculated for each number of clusters and plotted in the inset of Fig. 3. The SSEs declined rapidly as the number of clusters increased from one to three, then levelled off, implying that the linguistic features fall into three main clusters. The phylogenetic tree of these features is shown in Fig. 3. The linguistic features in the first cluster (yellow) are associated with nouns, either directly or indirectly, via the modification of nouns through determiners, adjectives or prepositional phrases. Linguistic features in the second cluster are associated with verb phrases, including low-frequency verbs, tense and possessives (grey). The third cluster (red) is dominated by linguistic features concerning clauses. The third cluster also contains linguistic features indicating highly abstract words with austere, template-like semantics, including pronouns and light verbs (e.g. VERB_HighFreq). Figure 3 highlights how the number of cluster-based similarities in the linguistic features matches the number of clusters derived from similarities between participants, providing converging support for a data-driven division of PPA patients into three clusters.
Figure 3 The phylogenetic clustering of the natural language processing-derived linguistic features. The linearized matrix of residuals of features in deficit was submitted to the phylogenetic tree function. The inset shows the scree plot of the total within the sum-of-squares error for different clusters of the linguistic features based on (dis)similarities used to generate the clusters, indicating three as the optimal number of clusters. Yellow features are associated with nouns, grey features are associated with verb phrases, including heavy verbs, and red features concern clauses and highly abstract words, including pronouns and light verbs. Features in ALL CAPS denote part of speech. Title case denotes lexical or inflectional features. Language features in all lowercase signify dependency relationships (for a full description of features, see Supplemental Material 3).
The linguistic feature cluster analysis used features that were in deficit in PPA. However, PPA can also indirectly affect the use of certain language features, leading to increased use of other language features. Figure 4 shows both the positive and negative residuals for the 52 features analysed in the cluster analysis, providing further insight into the nature of the three clusters. Moving from left to right, the first group of linguistic features (nouns) are in greatest deficit for svPPA patients (yellow). The middle group of linguistic features (heavy verbs and verb phrases) are those that are in greatest deficit for lvPPA patients (grey), and the third group of linguistic features (clauses, pronouns and light verbs) are those in greatest deficit nfvPPA patients (red). In general, the linguistic features associated with the first group centre on nouns, as would be expected in svPPA; the linguistic features associated with the second group centre on heavy verbs and verb phrases, a novel finding for lvPPA; and the features in the third group centre on clauses, pronouns and light verbs, which are major constituents of syntax as would be expected in nfvPPA.
Figure 4 Line chart showing the natural language processing-derived residuals of language features of each primary progressive aphasia variant compared with healthy individuals. Each line shows the residual value of language features for each PPA variant relative to healthy controls. The shaded areas show the features clustered together based on the phylogenetic clustering solution in Fig. 3. Features in the left region are in most deficit for svPPA patients, the middle region for lvPPA and the right region for nfvPPA. Features in ALL CAPS denote part of speech. Title case denotes lexical or inflectional features. Language features in all lowercase signify dependency relationships (for a full description of features, see Supplemental Material 3). lvPPA = logopenic variant PPA; nfvPPA = non-fluent variant PPA; PPA = primary progressive aphasia; svPPA = semantic variant PPA.
In addition to revealing deficits in the use of linguistic features, the results also show patterns of relative preservation or, possibly, compensation. Residuals above the zero-line indicate frequency counts that were greater than would be expected from healthy controls. The trade-off between high- and low-frequency verbs is shown in Fig. 5A, r(76) = −0.549, P < 0.001. This pattern helps to explain why previous work has not found verbs to be a predictor of PPA variants. Individuals with PPA do not stop using verbs but rather use different types of verbs. Interestingly, if the distinction between high- and low-frequency verbs is disregarded, there was no evidence of a verb deficit across any variant. Standardized (adjusted) Pearson residuals indicated no evidence that verbs as a whole were used more often in nfvPPA individuals than in healthy individuals, r = 1.012 P = 0.32. In lvPPA, there was evidence that verb usage increased relative to healthy controls, r = 2.54, P = 0.012, directly opposite to the deficit of low-frequency specific verbs that we found. Rather, evidence for verb loss appeared only when the verbs were divided into high- or low-frequency types. For example, when lvPPA patients have difficulty retrieving the verb ‘donate’, they might be able to produce ‘give’. Conversely, patients with nfvPPA produce a higher proportion of low-frequency verbs. A χ^2^ test of independence applied to the raw counts indicated that high/low verb frequency was not independent of patient type (nfvPPA, lvPPA), χ^2^(2) = 5.98, P < 0.001. Supporting this conclusion, Shan and Gerstenberger’s^51^ extended version of Fisher’s exact test indicated that for nfvPPA individuals, the rate of high-frequency verbs was less than in healthy controls, r = −2.79, P = 0.006, and the rate of using low-frequency verbs was greater than in healthy controls, r = 4.47, P < 0.001.^51^ Furthermore, for lvPPA individuals, the rate of using high-frequency verbs was greater than for healthy controls, r = 5.82, P < 0.011, and that of low-frequency verbs was marginally less than for healthy controls, r = −2.108, P = 0.038. This pattern indicates a double dissociation between PPA variant and verb type.
Figure 5 The scatterplots denoting trade-offs between language features. (A) The trade-off between low- and high-frequency verbs. Patients with nfvPPA are localized in the upper left corner, indicating relatively few high-frequency verbs, whereas lvPPA patients tend to be in the lower right corner, indicating relatively few low-frequency verbs. (B) The trade-off between nouns and prepositions. lvPPA = logopenic variant primary progressive aphasia; nfvPPA = non-fluent variant primary progressive aphasia; svPPA = semantic variant primary progressive aphasia.
Figure 5B shows a scatter plot indicating a strong negative correlation between pronoun and noun usage, r(76) = −0.827, P < 0.001. Individuals with lvPPA used high-frequency verbs more often than healthy controls, whereas individuals with nfvPPA used low-frequency verbs more often than healthy controls. An analogous pattern of results was observed between different types of nouns (common and pronouns) and PPA variant. A χ^2^ test of independence applied to the raw counts indicated that noun and pronoun counts were not independent of PPA variant (nfvPPA, svPPA), r = 177.9, P < 0.001. Fisher’s exact test indicated that for nfvPPA individuals, the rate of using common nouns was higher than in healthy controls, r = 6.284, P < 0.001, and lower in svPPA individuals, r = −12.51, P < 0.001. Conversely, for nfvPPA individuals, the rate of using pronouns was less than in healthy controls, r = −6.57, P < 0.001, and for svPPA individuals, the rate of using pronouns was higher than in healthy controls, r = 11.717, P < 0.001. The pattern of noun type (common nouns, pronouns) offers yet another example of a double dissociation between PPA variant and language feature.
Multinomial logistic regression analysis was conducted to identify which features best differentiated the three PPA variants. Variable selection was achieved by correlating the residuals from the previous analyses with dummy variables representing the participant groups. The resulting correlation coefficients were rank ordered. We used this ordered list of linguistic features to add variables to the model in a forward stepwise manner, with the exception of low-frequency verbs, which were included in the model for theoretical reasons discussed in this work. In addition to low-frequency verbs, the variables included high-frequency verbs, determiners, nouns, pronouns, adverbs, verbs, adverbial modifiers, determiners, clause markers, articles, demonstratives, finites, neuter, nominative case and personal/possessive pronouns.
The dependent variable of variant contained four nfvPPA, lvPPA, svPPA and healthy controls. Models were trained on the frequency counts, with healthy controls serving as a reference category. Table 2 shows the logistic coefficient for each predictor, which is the expected amount of change in the logit for each unit change in the predictor. A likelihood ratio test indicated that such a model significantly outperformed a model based on only the constant, χ^2^(51) = 270.25, P < 0.001. The 17-variable model after 10-fold cross-validation resulted in an accuracy of 97.88% (sensitivity/recall = 0.98; specificity = 0.99; precision = 0.98). The few errors were evenly distributed across the classes, as shown in the normalized confusion matrix in Table 3. When high- and low-frequency verbs were combined, the accuracy of the model after 10-fold cross-validation dropped to 74.5% (sensitivity/recall = 0.74; specificity = 0.90; precision = 0.74). A least likelihood test indicated that the model separating high- and low-frequency verbs accounted for more variance than a model that did not distinguish the two types of verbs, χ^2^(3) = 65.39, P < 0.001.
In this study, we leveraged advances in AI and machine learning to classify variants of PPA based on short naturalistic samples of connected speech. Through a data-driven approach, we used LLMs to measure language similarity among pairs of language samples and subsequently applied hierarchical clustering analysis to the resulting similarity matrix. Our analysis revealed the emergence of three main clusters, demonstrating an 88.5% agreement with the independent classification of PPA variants using the international consensus diagnostic criteria. These findings suggest that the three-variant classification of PPA is likely to reflect natural categorical groups measurable from naturalistic connected speech. Furthermore, as expected based on the predominant clinical variants in each of the three data-driven groups, regional brain atrophy in these groups matched well with the atrophy patterns that are well established for the three variants. Once we showed that the canonical PPA variants reflect natural kinds, we sought to identify the linguistic features that maximize the distinction between PPA variants. We used an automated syntactic parser to extract the linguistic features most robustly associated with each PPA variant, which sheds new light on dissociable aspects of impaired versus preserved elements of language in the three PPA variants.
We found that patients with svPPA exhibit deficits in nouns in addition to linguistic features related to noun modification. These linguistic features include determiners and features that are used in conjunction with noun phrases, such as expletives or other types of noun modifiers. Patients with lvPPA showed deficits in using low-frequency verbs in addition to verb features. Patients with lvPPA are also impaired in using features associated with possessive relationships. This category also includes particles, which in this dataset predominantly serve to form phrasal verbs. Finally, patients with nfvPPA exhibit deficits in constructing clauses. This group comprises features related to subordinate clauses, clausal modifiers and features frequently involved in subordinate clauses. Furthermore, grammatical case relationships, which encode the grammatical roles played by noun phrases in sentences, emerge only in the context of clauses. The clause-related features are associated with high-frequency verbs, which, as suggested in the two-level theory of verb meaning (see later), reflect basic clausal structures.^53-55^ Additionally, this third group of features includes pronouns, representing abstract versions of nouns, just as the high-frequency verbs represent abstract versions of low-frequency verbs.
Our a priori approach to separating high- versus low-frequency verbs in the classification model significantly increased the accuracy from 70% to 98%. In addition to improving the classification model, the high/low-frequency verb distinction could inform the neurolinguistic literature on how the category of verbs is processed across the language network. Inconsistent findings about the neurobiological underpinnings of verbs have been reported in the field. Similar inconsistencies exist in the PPA literature, with some studies reporting verb deficits in nfvPPA, whereas others have attributed the deficit to lvPPA, and yet other studies attributing no verb deficit to any PPA variant (for a review, see Thompson and Mack^56^). In this work, we identified a double dissociation between verb frequency and PPA variants, a pattern that could be explained, in part, by the two-level theory of verb meaning.^57-61^ This theory suggests that verbs consist of two separate layers of meaning. One layer is the ‘root’ or unique meaning, which captures idiosyncratic semantic features that (i) distinguish each verb in a given class from all the others; (ii) are often concrete and modality-specific; and (iii) do not interface with grammar.^57-61^ Another layer is the event structure template, which is (i) common to all the verbs in a given class; (ii) composed primarily of schematic predicates and variables for arguments; and (iii) relevant to the grammatical properties of all the verbs in a given class. Event structure templates are represented by a limited set of event types, such as state, result state, manner and instrument, which are defined by primitive predicates. In this model, the basic event structure classifies verbs. For instance, the semantic frame [x ACT
Our findings are consistent with a three-stage model of language production, which posits that language production results from the successive addition of increasingly complex linguistic elements, a concept we refer to as cumulative complexity. The model uses time-honoured categories of three primary groups of linguistic noun-related, verb-related and clause-related features. This approach implies a sequential order in linguistic processing, commencing with the assembly of noun phrases, followed by the construction of verb phrases and culminating in the synthesis of a complete sentence. In this model, no single brain region is dedicated to semantics or syntax. Grammar-related features are distributed throughout the language system. This interpretation is supported by meta-analyses showing that the ordering of cortical processing in the language networks starts in the temporal lobe before moving to the inferior parietal lobule and finally reaching the frontal lobe.^62,63^
Directionality in the language circuit is supported further by the pattern of impairment and potential compensation shown in Fig. 4. If the language network is directional, then compensation should be more likely to occur towards the end of the processing pipeline than at the beginning. Figure 4 shows precisely this pattern, with positive residuals emerging more prominently on the right of the pipeline than on the left. Interestingly, however, some compensation appears earlier in the pipeline. Such compensation could be explained by a recursive process by which outputs from the frontal lobe serve as inputs into the temporal lobe, potentially through the ventral stream via the extreme capsule fibre system/longitudinal inferior–frontal–occipital fasciculus or the uncinate fasciculus.^27,64^ As proposed by Friederici et al.,^27^ the uncinate fasciculus, which connects the frontal operculum and orbitofrontal cortex to the anterior superior temporal gyrus, might be involved in building local syntactic phrases.
Another significant aspect of the model is that damage to specific areas within the language networks, such as the frontoinsular cortex in nfvPPA or the temporal pole in svPPA, does not merely result in deficits in producing certain language features in comparison to healthy individuals. Instead, such damage also triggers a compensatory enhancement of other language features aimed at preserving the effective communication of information. For example, nfvPPA patients who have difficulty using complex syntax use a higher rate of nouns to all words in comparison to healthy controls. This finding aligns with our recent work showing that patients with nfvPPA who have difficulty using long and complex structures use more informative words, such as heavier verbs and more content words in their sentences to sustain sentence information.^35,65-68^ Likewise, patients with lexicosemantic deficits who have difficulty using nouns relative to other words produce more clause subordination than healthy individuals, as we have recently shown.^66^ For example, patients with svPPA exhibit a higher rate of embedding and other complex syntactic structures than healthy speakers.^66,69^
In summary, this study showcases the efficacy of contemporary generative LLMs in using data-driven analysis of a brief connected-speech sample to categorize patients with PPA into one of its three typical variants—a task traditionally accomplished by expert clinicians after exhaustive, specialized assessment typically taking several hours. Leveraging natural language processing for linguistic feature analysis, this approach identified linguistic features enabling robust classification, including those absent in the current diagnostic criteria, such as the pivotal role of verb categories when divided into high- or low-frequency verbs. Our methodology has the potential to refine existing diagnostic standards. For instance, the current criteria for subtyping svPPA and lvPPA include object naming and word retrieval deficiencies, respectively, a differentiation often blurred owing to their resemblance, thus creating challenges in clinical practice. Our findings indicate a more pronounced impairment in noun usage among svPPA patients, whereas lvPPA patients struggle more with retrieving proper nouns, such as the names of people. Such nuanced characterization of language deficits in PPA might facilitate the prediction of the underlying pathology of each syndrome.^70^
Besides offering a robust classification mechanism, our method also unravels insights into the neurobiological mechanisms of verb processing in the language network. The approach in this study can be extended to other neurodegenerative conditions, fostering a more objective, theory-neutral categorization system that could provide further insights into the neurobiology of language and other aspects of cognition critical for communication. Establishing a reliable connection between language characteristics and neural processes requires analysis across a diverse range of language families. For instance, our recent findings indicated that the set of highly frequent verbs in English is less pronounced in Persian, suggesting that certain aspects of verb distribution might be unique to English and not necessarily applicable to all other languages.^71^ Future work is needed to test the generalizability of our results on an entirely different cohort of PPA samples and to extend these evaluations to various forms of discourse, such as storytelling and other types of narrative. Certain language impairments might become apparent only under the specific pressures of a language elicitation task. Connected speech allows individuals to adopt compensatory strategies, using alternative linguistic features when faced with deficits in producing some features. Confrontational naming tasks might underappreciate these subtle dynamics, potentially leading to an inaccurate assessment of deficits, particularly if the frequency of the items presented is not controlled meticulously. Ideally, future studies would also assess the acoustic features of speech to investigate their potential to improve the accuracy of the classification model. Collectively, this approach would establish a solid foundation for a deeper understanding, hence enhanced management of neurodegenerative disorders.
We thank the research participants and their care partners, without whose contributions this work would not have been possible. We also thank the Athinoula A. Martinos Center for biomedical imaging support.
Neguine Rezaii, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA.
Daisy Hochberg, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA.
Megan Quimby, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA.
Bonnie Wong, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA; Department of Psychiatry, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA.
Michael Brickhouse, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA.
Alexandra Touroutoglou, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA; Department of Psychology, Emory University, Atlanta, GA 30322, USA.
Bradford C Dickerson, Frontotemporal Disorders Unit, Department of Neurology, Massachusetts General Hospital, Harvard Medical School, Boston, MA 02114, USA; Athinoula A. Martinos Center for Biomedical Imaging, Harvard Medical School, Boston, MA 02129, USA; Massachusetts Alzheimer's Disease Research Center, Harvard Medical School, Boston, MA 02114, USA.
Phillip Wolff, Department of Psychology, Emory University, Atlanta, GA 30322, USA.
Codes and data used in this work can be accessed upon sending a request to Dr Bradford Dickerson at brad.dickerson@mgh.harvard.edu.
This work was supported by the US National Institute on Deafness and Other Communication Disorders grants R01 DC014296 to B.C.D. and R21 DC019567 to B.C.D. and P.W., National Institute on Aging grants R01 AG081249 to B.C.D. and R21 AG073744 to B.C.D. and P.W., National Institute of Neurological Disorders and Stroke grant RF1 NS131395 to B.C.D., and Alzheimer’s Association grant 23AACSF-1029880 and Massachusetts General Hospital (MGH) Screening Technologies in Primary Care Innovation Fund (PCIF) 2023A063002 to N.R. This research was carried out in part at the Athinoula A. Martinos Center for Biomedical Imaging at the MGH, using resources provided by the Center for Functional Neuroimaging Technologies, P41EB015896, a P41 Biotechnology Resource Grant supported by the National Institute of Biomedical Imaging and Bioengineering (NIBIB), National Institutes of Health. This work also involved the use of instrumentation supported by the National Institutes of Health Shared Instrumentation Grant Program and/or High-End Instrumentation Grant Program, specifically, grant number(s) S10RR021110, S10RR023043 and S10RR023401.
B.C.D. has served as a paid consultant for Acadia, Alector, Arkuda, Biogen, Denali, Eisai, Genentech, Ilios, Lilly, Merck, Takeda and Wave Lifesciences and a paid editor for Elsevier. These relationships are not related to the content of the manuscript. All other authors declare that they have no competing interests.
Supplementary material is available at Brain online.
Codes and data used in this work can be accessed upon sending a request to Dr Bradford Dickerson at brad.dickerson@mgh.harvard.edu.