Authors: Yinghui Geng, Huijun Zhang
Categories: Article, Alzheimer’s disease, MMSE, machine learning, random forest, structural MRI, risk stratification, Diseases, Medical research, Neurology, Neuroscience
Source: Scientific Reports
Authors: Yinghui Geng, Huijun Zhang
Early identification of patients with Alzheimer’s disease (AD) who will experience near-term cognitive decline can support trial enrichment and risk-stratified follow-up. Using the Alzheimer’s Disease Neuroimaging Initiative (ADNI), we developed two prognostic models for 12-month Mini-Mental State Examination (MMSE) decrease (≥ 3 points): (i) a clinical logistic-regression model and (ii) a random-forest model combining clinical variables with MRI-derived volumetric measures. In 306 participants with baseline AD and complete 12-month MMSE (mean age 74.8 years; baseline MMSE 23.1), 131 (42.8%) declined. Five-fold stratified cross-validation with within-fold preprocessing and imputation was used for internal validation. The clinical model achieved an area under the ROC curve (AUC) of 0.755, while the random-forest model achieved an AUC of 0.773 and provided higher net benefit across threshold probabilities of 0.20–0.80 in decision-curve analysis. Risk stratification using pre-specified cut-offs (< 0.25, 0.25–0.50, ≥ 0.50) yielded monotonic observed decline rates (13.2%, 35.3%, 67.2%). These findings suggest that a transparent two-model framework based on ADNI data provides moderate prognostic accuracy and clinically interpretable three-tier risk stratification; however, external validation and local recalibration are required before clinical implementation.
The online version contains supplementary material available at 10.1038/s41598-026-43321-1.
Alzheimer’s disease (AD) remains the most common cause of dementia in older adults. According to the Alzheimer’s Association 2024 report, an estimated 6.9 million U.S. adults aged ≥ 65 years were living with AD dementia in 2024. Without effective preventive or disease-modifying strategies, this number may reach 13–14 million by 2060, indicating a sustained rise in clinical and economic burden on memory clinics and long-term care systems^1^. At the same time, recent high-level reviews emphasise that AD is biologically and clinically heterogeneous—patients follow different atrophy patterns, carry different biomarker profiles, and progress at different speeds—which makes it difficult to identify, at baseline, those who will deteriorate rapidly enough to justify intensified monitoring or trial enrolment^2^.
Longitudinal cohort analyses have shown that even within already diagnosed AD populations, some individuals experience a clinically meaningful decline over 12–18 months (for example, MMSE decrease ≥ 3 points), whereas others remain comparatively stable; such heterogeneity reduces statistical power in AD trials and complicates real-world follow-up planning^3,4^. A pragmatic solution is to develop prediction models that can flag, at baseline, those patients who are likely to decline over the next year, so that they can be prioritised for disease-modifying or non-pharmacological interventions, or simply be scheduled for closer cognitive/functional assessments^3,4^.
The Alzheimer’s Disease Neuroimaging Initiative (ADNI) provides precisely the kind of multimodal, prospectively collected data that support such models. Several 2023–2024 studies using ADNI or ADNI-like cohorts have demonstrated that combining structural MRI, cognitive scales and other clinical features in machine-learning (ML) frameworks (gradient boosting, random forests, or deep multimodal networks) improves prediction of AD progression or conversion compared with single-modality models^3–5,8^. However, most of these studies focused on MCI-to-AD conversion or multi-year trajectories, and only a minority reported short-term (12-month) cognitive decline in patients who already had AD, or translated their model outputs into clinically interpretable risk strata. Moreover, many ML papers reported AUC only, without complementary metrics such as Brier score, calibration slope/intercept, or—most importantly—decision-curve analysis (DCA), even though current reporting guidance considers DCA essential for judging whether a new model offers net benefit over “treat all” or “treat none.”^9,10^.
Methodological updates such as the 2024 TRIPOD + AI statement and the BMJ series on evaluation of clinical prediction models now explicitly recommend two things that are directly relevant to ADNI-based machine-learning (i) any complex or AI-based model should be compared with a simple, clinically plausible baseline model built from routinely available variables (e.g. age, sex, baseline MMSE, ADAS-Cog); and (ii) model performance should be presented together with decision-curve analysis (DCA) across the range of thresholds that are realistic for clinical decision-making^6,7,9,10^. Following this framework, our study—“Machine-learning-based prediction and risk stratification of 12-month cognitive decline in Alzheimer’s an ADNI analysis with model comparison and decision-curve evaluation”—used ADNI participants with complete 12-month MMSE (1) build a transparent clinical baseline comparator; (2) develop and tune an RF model on the full candidate-predictor set; (3) compare discrimination, calibration and Brier score under five-fold cross-validation; and (4) apply DCA to identify probability thresholds (~ 0.20–0.80) with the highest net benefit and to derive a three-tier risk-stratification scheme for memory-clinic use.
We conducted a retrospective prognostic modelling study using the Alzheimer’s Disease Neuroimaging Initiative (ADNI) ADNIMERGE dataset. The baseline visit was defined as the first ADNI visit at which participants fulfilled the AD diagnosis. All candidate predictors were obtained at baseline, and the outcome was assessed 12 months later.
The analysis followed recent guidance for reporting machine-learning prediction models (TRIPOD-AI) and for evaluating clinical prediction models. The ADNI study was approved by the Institutional Review Board of the University of Southern California (USC) and by the institutional review boards of all participating sites. All participants provided written informed consent. As this study used de-identified secondary data, no additional ethical approval was required.
From ADNIMERGE we first identified all individuals with a baseline diagnosis of AD and baseline cognitive, demographic and MRI-derived variables available (n = 411; Table 1). Model development and validation were then restricted to participants who also had 12-month Mini-Mental State Examination (MMSE) data available (n = 306); this analytic sample was used for all modelling steps and for Tables 2, 3 and 4; Figs. 1, 2 and 3.
The primary outcome was 12-month cognitive decline, defined a priori as a decrease of ≥ 3 points in MMSE from baseline to the 12-month visit. This threshold has been used in AD clinical studies to indicate a clinically meaningful decline and was consistent with the observed event rate in this ADNI sample (42.8%).
Candidate predictors were selected to reflect variables routinely collected in ADNI memory-clinic settings and those shown to be important in previous ADNI machine-learning
Demographics/ age, sex, years of education, APOE ε4 carrier (≥ 1 allele; derived from APOE4 allele count in ADNIMERGE);
Baseline cognition and clinical baseline MMSE, ADAS-Cog 13, ADAS-Cog 11, ADAS-Q4, CDR-Sum of Boxes (CDR-SB), and Functional Activities Questionnaire (FAQ); ADAS-Cog 11 and ADAS-Cog 13 are expected to be highly correlated; we retained both to capture overlapping but non-identical scoring ranges used in routine assessments, noting that tree-based models can accommodate correlated predictors, and we interpret feature importance with redundancy in mind.
MRI-derived structural hippocampal volume, ventricular volume, whole-brain volume, entorhinal cortex volume, fusiform volume, middle temporal gyrus volume, and intracranial volume (ICV) (see Supplementary Methods for details);
Other care-related or potentially modifiable factors that were repeatedly selected in the RF model were explored in sensitivity analyses (see Supplementary Methods).
All variables were inspected for missingness; missingness percentages for each candidate predictor in the analytic sample are summarized in Supplementary Table S4. Within each cross-validation training fold, continuous variables with missing values were imputed using the median and categorical variables were imputed using the most frequent category (single imputation). This pragmatic approach may introduce bias and underestimate uncertainty, particularly for MRI-derived variables with higher missingness (Supplementary Table S4); alternative strategies (e.g., multiple imputation) may yield different estimates and should be evaluated in future work. We did not formally test missingness mechanisms (e.g., MCAR/MAR/MNAR) or perform multiple imputation; given the modest sample size and the goal of a reproducible modelling pipeline, we used simple within-fold imputation and report missingness transparently. Continuous predictors were then standardised to zero mean and unit variance. Categorical predictors (e.g. sex, APOE4) were one-hot encoded. All preprocessing steps were performed within each cross-validation fold to avoid data leakage. MRI-derived volumetric measures were taken as provided in ADNIMERGE (centrally processed with standardized pipelines); we did not apply additional site/scanner harmonization methods (e.g., ComBat).
To make the incremental value of machine learning explicit, we compared two prespecified models, both trained and evaluated on the same analytic sample (n = 306):
Model 1 (baseline clinical model): logistic regression including age, sex, baseline MMSE and ADAS-Cog 13; presented in Table 4.Model 2 (machine-learning model): tuned RF using all candidate predictors (clinical variables and MRI-derived brain volumes). This model was used to generate the ROC, calibration and decision-curve plots (Figs. 1, 2 and 3) and is presented in Table 4. Key tuned random-forest hyperparameters are reported in Supplementary Table S9 (n_estimators = 200, max_features=sqrt, max_depth=None; full list in Table S9).Sensitivity because ventricular volume can be influenced by inter-individual head size, we repeated the main analyses replacing absolute ventricular volume and ICV with the ventricles/ICV ratio. Findings were consistent; details are provided in the Supplementary Methods and Supplementary Table S5.
We used five-fold stratified cross-validation. Given the moderate event rate (42.8%), we used outcome-stratified folds and did not apply additional imbalance techniques (e.g., SMOTE or class-weighting). In each fold, models were trained on 80% of the data and tested on the remaining 20%; all out-of-fold predictions were concatenated to obtain overall out-of-fold probabilities. Model performance was summarized (i) area under the ROC curve (AUC) for discrimination (Fig. 1); (ii) Brier score for overall accuracy of predicted probabilities (Table 4); (iii) calibration slope and intercept estimated from out-of-fold probabilities by fitting a logistic calibration model regressing the observed outcome on the logit of the out-of-fold predicted probabilities (Fig. 2; Table 4); and (iv) threshold-specific sensitivity, specificity, positive predictive value (PPV), negative predictive value (NPV) and Youden’s J at probability thresholds of 0.25, 0.40 and 0.50 (Table 3). Within each cross-validation training fold, we fitted an isotonic regression calibrator on the training predictions and applied it to the validation fold to obtain out-of-fold calibrated probabilities.
To assess potential clinical utility, we performed decision-curve analysis (DCA)^9,10^ on the out-of-fold calibrated probabilities of both models. Net benefit was calculated across probability thresholds from 0.20 to 0.80, a range relevant to risk-stratified follow-up in memory-clinic settings. Thresholds below 0.20 or above 0.80 were not emphasised because they imply very low or very high risk tolerance, where net benefit becomes unstable and decisions are less plausible in routine memory-clinic workflows.
Based on decision-curve analysis (DCA) and the distribution of predicted probabilities from the random forest model, we defined three risk strata low (< 0.25), intermediate (0.25–0.50) and high (≥ 0.50). For each stratum we calculated the number of patients, the observed 12-month decline rate and the mean MMSE change (Table 2). We further present an illustrative clinical application low-risk patients could be considered for routine annual assessment; intermediate-risk patients for closer monitoring (e.g., 3–6 months) and optimisation of modifiable factors; and high-risk patients for closer follow-up and further clinical assessment, subject to local resources and external validation. In the RF model, observed decline rates increased monotonically across 13.2% (9/68) in low-risk, 35.3% (42/119) in intermediate-risk, and 67.2% (80/119) in high-risk groups. Median predicted risks were 0.161, 0.370, and 0.644, respectively (IQRs 0.126–0.209; 0.305–0.429; 0.553–0.709).
All data management and modelling were performed in Python 3.13 (pandas 2.x, scikit-learn 1.4) and R 4.3 (for calibration and decision-curve plotting). Cross-validation and model tuning were scripted to ensure reproducibility. A large language model tool (ChatGPT, OpenAI) was used to assist with language editing only; all scientific content, analyses, and interpretations were performed and verified by the authors. Model interpretability was assessed using SHAP (SHapley Additive exPlanations). SHAP values were computed from a final random forest model refitted on the full analytic dataset using the tuned hyperparameters. These analyses were conducted post hoc and were not used in model training, feature selection, or performance evaluation.
A total of 411 ADNI participants met the AD diagnosis and had baseline cognitive, demographic and MRI-derived variables recorded (mean age 74.8 ± 7.9 years; baseline MMSE 23.15 ± 2.19). Of these, 306/411 (74.5%) also completed the 12-month MMSE assessment, and all analyses, tables (Tables 2, 3 and 4) and figures (Figs. 1, 2 and 3) were based on this analytic sample (n = 306). Additional baseline categorical characteristics are summarised in Supplementary Table S2.
At 12 months, 131/306 participants met the predefined endpoint of MMSE decrease ≥ 3 points, giving an event rate of 42.8%. This event rate was used in the threshold-based operating-characteristics analysis (Table 3) and in the decision-curve analysis.
Both prespecified models were trained and evaluated using five-fold stratified cross-validation on the same 306 participants. The baseline clinical model (age, sex, baseline MMSE, ADAS-Cog 13) achieved an AUC of 0.755 and a Brier score of 0.199.
The RF model achieved an AUC of 0.773 based on out-of-fold predictions from five-fold stratified cross-validation with a Brier score of 0.192. The absolute gain in discrimination over the baseline clinical model was modest, so decision-curve analysis is emphasized as the primary evidence for added clinical value. Calibration slopes remained below 1 for both models (Table 4), indicating some overfitting and underscoring the need for local recalibration when transported. Decision-curve analysis showed that the RF model provided higher net benefit than the “treat-all” and “treat-none” strategies across threshold probabilities of 0.20–0.80 (Fig. 3); ROC and calibration plots are shown in Figs. 1 and 2. A sensitivity analysis adjusting ventricular volume for intracranial volume (ventricles/ICV) produced consistent results (Supplementary Table S5).
To provide a clinically interpretable summary, we examined the out-of-fold predictions at the probability threshold that maximised Youden’s J (≈ 0.40). At this cut-off the confusion matrix was TP 103, FP 61, TN 114, FN 28, corresponding to sensitivity 78.6% and specificity 65.1%, in line with Table 3. This cut-off classified 164/306 participants (53.6%) as test-positive (predicted probability ≥ 0.40), capturing 103/131 decliners (78.6%) while yielding 61/175 false positives; PPV was 62.8% and NPV was 80.3%. At a threshold probability of 0.40 in DCA, the RF model achieved a net benefit of 0.204, equivalent to 20.4 net true-positive decisions per 100 patients versus treating none and a net reduction of 23.5 unnecessary interventions per 100 patients compared with treating all. The corresponding confusion matrix is provided in Supplementary Table S1.
Table 3 reports operating characteristics at prespecified thresholds of 0.25 and 0.50 and at the best Youden’s J (0.40). At 0.25, sensitivity was very high (93.1%) but specificity was low (33.7%), which is suitable for rule-out or early recall. At 0.50, specificity increased to 77.7% while sensitivity dropped to 61.1%, which is suitable for rule-in or intensified follow-up. The best Youden’s J was 0.438 at 0.40, with balanced accuracy 71.9% and overall accuracy 70.9%. Additional threshold-specific metrics are provided in Supplementary Table S3, and a clinician-oriented interpretation of key cutoffs is summarized in Supplementary Table S7.
When RF-predicted probabilities were grouped into the three prespecified strata, 68 (22.2%) patients were low risk (< 0.25), 119 (38.9%) were intermediate risk (0.25–0.50), and 119 (38.9%) were high risk (≥ 0.50). Observed 12-month cognitive-decline rates increased monotonically from 13.2% (low) to 35.3% (intermediate) and 67.2% (high); mean MMSE change also followed this gradient (≈ 0, − 1.6, − 4.4 points, respectively). This confirms that RF-generated probabilities can be translated into clinically interpretable tiers (Table 2). The distribution of participants across these strata is shown in Supplementary Figure S3. Observed decline rates by stratum are shown in Supplementary Figure S4.
Decision-curve analysis based on out-of-fold calibrated probabilities showed that the RF model provided net benefit across a wide range of threshold probabilities (approximately 0.20–0.80) compared with the “treat-all” and “treat-none” strategies (Fig. 3). In practical memory-clinic ranges (≈ 0.20–0.50), using the RF model could identify more true decliners without generating excessive false positives. Net benefit at representative decision thresholds is reported in Supplementary Table S8.
Fig. 1Receiver operating characteristic (ROC) curve of the random forest (RF) model for predicting 12-month cognitive decline in patients with Alzheimer’s disease (AD).
The model was developed with 5-fold cross-validation and evaluated on out-of-fold predictions among participants with complete 12-month MMSE data (n = 306) from the ADNI cohort. Cognitive decline was defined as an MMSE decrease ≥ 3 points. The diagonal line represents a non-informative classifier.
Fig. 2Calibration plot of the RF model for predicting 12-month cognitive decline in patients with AD (MMSE decrease ≥ 3 points).
Predicted probabilities were grouped into deciles, and the observed event rates within each decile were plotted against the corresponding predicted risks (n = 306). The dashed 45° line indicates perfect calibration, and the solid line shows the model-estimated calibration.
Fig. 3Decision-curve analysis (DCA) of the RF model for 12-month cognitive decline in patients with AD.
The RF model (solid line) was compared with the “treat-all” and “treat-none” strategies using out-of-fold calibrated probabilities. Within the clinically relevant threshold range of 0.20–0.80, the RF model provided higher net benefit, supporting its use for risk-stratified follow-up.
Table 1Baseline characteristics of participants with Alzheimer’s disease in the ADNI cohort (n = 411).Variable n MeanSDMedianMinMaxMissingAge, years41174.757.9475.355.190.90MMSE at baseline41123.152.1923.016.030.00ADAS-Cog 13 at baseline40229.968.029.3312.6754.679CDR-SB at baseline4114.431.694.51.010.00Hippocampal volume, mm³3385770.491025.375655.152991.09572.073Intracranial volume, mm³4021528337.54224488.961494575.01071900.03315210.09MMSE at 12 months30620.914.4622.04.029.010512-month MMSE change306-2.343.87-2.0-18.05.0105Note: Continuous variables are shown as mean (SD); for skewed distributions, median (min, max) is also presented. “Missing” indicates the number of participants without data for the corresponding variable. MMSE at 12 months and MMSE change were available for 306/411 participants and analyses in Tables 2, 3 and 4 were restricted to this subgroup.AD Alzheimer’s disease, ADAS-Cog 13 Alzheimer’s Disease Assessment Scale–Cognitive Subscale (13-item), CDR-SB Clinical Dementia Rating–Sum of Boxes, ICV intracranial volume, MMSE Mini-Mental State Examination.
Table 2Risk stratification of 12-month cognitive decline based on random-forest–predicted probabilities (ADNI participants with 12-month MMSE, n = 306).Risk stratum N Events, nEvent rate, %Mean MMSE change, pointsLow (< 0.25)68913.20.04Intermediate (0.25–0.50)1194235.3-1.62High (≥ 0.50)1198067.2-4.42Note: Cognitive decline was defined as a decrease of ≥ 3 points in MMSE from baseline to 12 months. Predicted probabilities were obtained as out-of-fold calibrated probabilities from five-fold cross-validation. Risk groups were prespecified as low (< 0.25), intermediate (0.25–0.50), and high (≥ 0.50). Event rate = events / N. Mean MMSE change is shown for descriptive purposes only.MMSE Mini-Mental State Examination.
Table 3Threshold-specific performance of the random forest (RF) model for 12-month cognitive decline (ADNI participants with 12-month MMSE, n = 306).Threshold typeThresholdPrevalence, %Sensitivity, %Specificity, %PPV, %NPV, %Accuracy, %F1, %Balanced accuracy, %Youden’s JBest Youden’s J0.4042.878.665.162.880.370.969.871.90.438Preset 0.250.2542.893.133.751.386.859.266.163.40.268Preset 0.50.5042.861.177.767.272.770.664.069.40.388Note: Cognitive decline was defined as an MMSE decrease of ≥ 3 points from baseline to 12 months. “Best Youden’s J” denotes the probability threshold that maximized (sensitivity + specificity − 1). PPV and NPV were influenced by the observed prevalence of cognitive decline in this sample (42.8%). Accuracy = (TP + TN) / total; F1 = 2 × (precision × recall) / (precision + recall); balanced accuracy = (sensitivity + specificity) / 2.FN false negative, FP false positive, MMSE Mini-Mental State Examination, NPV negative predictive value, PPV positive predictive value, RF random forest; TN true negative, TP true positive.
Table 4Comparative performance of baseline and machine-learning models for predicting 12-month cognitive decline (MMSE decrease ≥ 3 points) in ADNI participants (n = 306).ModelPredictors includedAUCBrier scoreCalibration slopeCalibration interceptNet benefit (0.20–0.80)Model 1 (baseline clinical)Age, sex, baseline MMSE, ADAS-Cog 130.7550.1990.1400.474Low, mainly at ≤ 0.30Model 2 (RF, main model)All candidate predictors, tuned random forest0.7730.1920.2370.511Highest across 0.20–0.80Note: Both models used the same analytic sample (ADNI participants with complete 12-month MMSE, n = 306) and five-fold cross-validation with out-of-fold predictions. Calibration slope and intercept were estimated from cross-validated predictions; a calibration slope < 1 suggests overfitting, while the intercept reflects calibration-in-the-large. The RF model included MRI-derived volumes (including ventricular volume and ICV), which were additionally explored in sensitivity analyses using ventricles/ICV (Supplementary Table S5). Bootstrap 95% confidence intervals for discrimination and calibration metrics are provided in Supplementary Table S6.ADNI Alzheimer’s Disease Neuroimaging Initiative, MMSE Mini-Mental State Examination, RF random forest.
In this ADNI-based prognostic study, we developed a tuned random-forest (RF) model to predict clinically meaningful 12-month cognitive decline (MMSE decrease ≥ 3 points) in patients with Alzheimer’s disease and benchmarked it against a transparent baseline clinical logistic-regression model. The RF model achieved moderate discrimination (out-of-fold AUC ~ 0.77) with acceptable Brier score and showed positive net benefit across clinically relevant thresholds on decision-curve analysis^11^. When translated into three prespecified risk strata, 12-month decline rates increased monotonically (~ 13% to ~ 35% to ~ 67%), suggesting potential utility for trial enrichment and risk stratification, pending external validation^12,13^. Independent external validation and local recalibration are required before clinical deployment.
Our findings align with recent ADNI-anchored and multimodal studies showing that combinations of routine cognitive measures and structural MRI can support short-term prognosis in AD and AD-spectrum cohorts, and that longitudinal modelling of ADNI participants enables more precise characterization of individual progression trajectories^12,14–16^. Although more complex deep/ensemble approaches that integrate PET, CSF, or validated plasma markers such as p-tau217 can further improve discrimination^17,18^, our results support the view that MRI+clinical pipelines remain clinically meaningful and far more deployable in settings where advanced biomarkers are not routinely available. The current ADNI programme and its prospective extensions continue to provide the methodological and data infrastructure for such ML-based prognostic studies^16,19^.
In line with contemporary recommendations on evaluating ML-based prediction models, we benchmarked complex models against parsimonious clinical baselines and reported discrimination, Brier score, calibration slope/intercept, and DCA^6,7,20^. The fact that the RF model maintained net benefit over a wide range of risk thresholds underscores that decision-analytic reporting—not accuracy alone—is essential when a model is intended to guide follow-up intensity, biomarker testing, or trial screening^11^. This approach is consistent with recent BMJ guidance on the development–validation continuum for clinical prediction models^20^. Calibration slopes substantially below 1 indicate potential overfitting and overly extreme risk predictions. Therefore, model outputs should be interpreted with caution, and recalibration is essential before any clinical application.
Model explainability using SHAP (Supplementary Figures S1–S2) highlighted baseline cognitive severity—especially ADAS-Cog components and global MMSE—together with medial-temporal and ventricular volumetrics as dominant risk drivers^21^. These attributions are biologically plausible given ADNI syntheses linking regional neurodegeneration, global atrophy and near-term clinical worsening^14^, and they are also consistent with emerging biomarker recent plasma p-tau217 work shows that blood-based markers can flag early amyloid/tau pathology in clinically unimpaired individuals^17^, while earlier studies demonstrated that p-tau217 discriminates AD from other neurodegenerative disorders^18^. Presenting patient-level feature-contribution plots can therefore bridge statistical output and clinician–patient communication, in line with broader work on explainable AI for medical imaging and neurodegenerative disease^22^. Consistent with clinical expectations, the most influential predictors included baseline ADAS-Cog 13, age, and MRI-derived volumetric measures (e.g., ICV, ventricles, hippocampal and temporal lobe volumes).
Translating continuous probabilities into low-, intermediate- and high-risk tiers produced clearly separated 12-month decline rates (for example, ≈ 13% → ≈35% → ≈67%), supporting two potential use-cases: (i) prognostic enrichment to raise event rates and reduce follow-up burden in disease-modifying or proof-of-concept trials; and (ii) risk-stratified follow-up to prioritize confirmatory biomarkers or closer monitoring for the highest-risk group^12,13^. This mirrors recent deep-learning–based patient-stratification work that seeks to optimize clinical dementia trials by oversampling those most likely to decline^13^.
Key strengths of this study the use of a harmonized multimodal cohort (ADNI); head-to-head comparison of ML algorithms under stratified cross-validation with probability recalibration; explicit reporting of decision-analytic results; and the integration of explainability to improve clinician trust. These elements are in line with current expectations for transparent, evaluable clinical prediction models and with the movement toward end-to-end clinical-AI pipelines^20,23^.
ADNI participants are typically research volunteers with relatively high education levels and limited ethnic diversity; this may restrict transportability to community, hospital or multiethnic cohorts, making independent, locally representative external validation and local recalibration essential^14,16,19^. ADNI is a multi-site study; residual scanner/site differences can affect volumetric predictors, and we did not apply additional harmonization. The MMSE-based decline definition may be influenced by test–retest variability and ceiling/floor effects; future work should evaluate robustness using alternative or complementary endpoints (e.g., ADAS-Cog or CDR-SB change), repeated assessments, and sensitivity analyses around the decline threshold. We used single imputation (median/mode) within training folds, which may introduce bias and underestimate uncertainty, particularly for MRI-derived predictors with non-trivial missingness (Supplementary Table S4). Our feature set intentionally emphasized routinely available clinical and MRI variables to enhance feasibility, which will likely cap discrimination compared with models that incorporate PET, CSF or plasma p-tau217^17,18^. We also did not include white matter hyperintensity burden or a broader set of subcortical volumes, which may contribute to cognitive decline and may interact with grey matter atrophy; these features should be explored in future studies when consistently available and during external validation. Moreover, ML models for AD progression can exhibit performance and calibration differences across subgroups defined by race/ethnicity, education or language, so subgroup-specific auditing, reporting and recalibration (e.g., intercept/slope adjustment, Platt scaling or isotonic regression) should be planned^24–26^. These issues become particularly salient when such models are translated to health-care systems that rely on culture-fair cognitive tools or cross-culturally adapted neuropsychological assessment^25,26^.
For real-world adoption, we (a) prespecifying calibration-monitoring and model-updating triggers to handle dataset shift; (b) embedding clinician-facing explanations (top features and directionality) in the interface; (c) mapping each risk tier to clear standard operating procedures (follow-up interval, confirmatory biomarker strategy, trial referral); and (d) conducting prospective or pragmatic impact evaluations within an end-to-end clinical-AI framework such as SALIENT to quantify effects on workflow, patient-centred outcomes and costs^22,23^.
With appropriate external validation and local recalibration, a 12-month decline predictor can target patients for prioritized biomarker testing or intensified monitoring, support shared decision-making using calibrated probabilities plus the dominant contributors for a given patient, and assist trial teams in enriching for likely decliners within typical trial windows^12,13,17^. Where MRI or advanced biomarkers are limited, simplified cognitive-centric workflows informed by the same modelling principles may still offer actionable stratification, provided local data are used for recalibration^16,19^.
Future work should (1) multi-region, multiethnic external validation in both clinical and community cohorts; (2) incremental improvement via validated blood biomarkers (for example, p-tau217) and PET where feasible; (3) integration of longitudinal digital phenotypes from wearable devices and other remote-monitoring technologies, as demonstrated in recent ML studies of MCI trials^27^; (4) pragmatic or stepped-wedge impact trials of model-guided care pathways; and (5) embedding continuous fairness monitoring and governance within implementation frameworks to ensure safe, equitable translation^23–26^.
A transparent ML workflow using routinely collected cognitive and structural MRI variables provided moderate discrimination and positive decision-analytic value for predicting 12-month MMSE decline in AD. Translating calibrated probabilities into three risk tiers offers a practical route for trial enrichment and risk-stratified care in settings without routine PET/CSF. Calibration required and remained imperfect in internal validation, and external/cross-cultural validation with fairness auditing and prospective impact evaluation are needed before routine clinical use.
Below is the link to the electronic supplementary material.
Supplementary Material 1