Authors: Walker S. McKinney (1Department of Behavioral Medicine and Clinical Psychology, Cincinnati Children’s Hospital Medical Center, 3333 Burnet Avenue, Cincinnati, OH 45229-3026, USA), Meredith Nelson (1Department of Behavioral Medicine and Clinical Psychology, Cincinnati Children’s Hospital Medical Center, 3333 Burnet Avenue, Cincinnati, OH 45229-3026, USA; 2Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH, USA), Rebecca C. Shaffer (1Department of Behavioral Medicine and Clinical Psychology, Cincinnati Children’s Hospital Medical Center, 3333 Burnet Avenue, Cincinnati, OH 45229-3026, USA; 2Department of Pediatrics, University of Cincinnati College of Medicine, Cincinnati, OH, USA), Kelli C. Dominick (3Division of Child and Adolescent Psychiatry, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH, USA; 4Department of Psychiatry and Behavioral Neuroscience, University of Cincinnati College of Medicine, Cincinnati, OH, USA), Craig A. Erickson (3Division of Child and Adolescent Psychiatry, Cincinnati Children’s Hospital Medical Center, Cincinnati, OH, USA; 4Department of Psychiatry and Behavioral Neuroscience, University of Cincinnati College of Medicine, Cincinnati, OH, USA), Lauren M. Schmitt (5Phelan-McDermid Syndrome Foundation, Osprey, FL, USA)
Categories: Article, Fragile X syndrome, Stanford–Binet, Intelligence, IQ, Cognition, Intellectual disability
Source: Journal of autism and developmental disorders
Authors: Walker S. McKinney, Meredith Nelson, Rebecca C. Shaffer, Kelli C. Dominick, Craig A. Erickson, Lauren M. Schmitt
Fragile X syndrome (FXS) is the most common inherited cause of intellectual disability and single-gene cause of autism. The Stanford-Binet, Fifth Edition (SB-5) is commonly used to assess IQ in FXS. It is not known if the SB-5 routing form’s abbreviated IQ (ABIQ) score accurately estimates full-scale IQ (FSIQ), limiting data-informed decision-making when choosing between an abbreviated or full SB-5 battery.
198 participants with FXS (143 males) aged 4 to 47 years of age completed the full SB-5. We calculated differences between abbreviated and full-scale estimates of IQ and assessed the extent to which the agreement between ABIQ and FSIQ varied as a function of age, routing subtest scatter, and FSIQ.
The abbreviated SB-5 battery over-estimated FSIQ in most school-age children (< 11 years), and under-estimated FSIQ in adolescents and adults. This under-estimate of FSIQ was larger when there was a greater discrepancy (scatter) between the two routing subtests that comprise ABIQ and in individuals with FSIQ < 68.
Clinicians and researchers should consider administering the full SB-5 battery to individuals with FXS when possible. If only an abbreviated estimate of IQ is available, ABIQ should be interpreted with caution based on our findings of over- or under-estimation occurring across development. Large discrepancies between verbal and nonverbal skills as well as greater severity of ID should both serve as cues to administer the full battery to avoid under-estimating cognitive skills that are otherwise only captured by FSIQ.
Fragile X syndrome (FXS) is the most common heritable form of intellectual disability and single-gene cause of autism spectrum disorder (ASD), occurring in approximately 1 in 7,000 males and 1 in 11,000 females (Hunter et al., 2014). FXS is caused by a trinucleotide repeat expansion (> 200 CGG repeats) in the Fragile X messenger ribonucleoprotein 1 (FMR1) gene on the X chromosome. Nearly all males with FXS have an intellectual disability, although the severity is highly variable (Schmitt et al., 2024). When present, intellectual disability tends to be milder in females with FXS, and many females with FXS have IQ scores in the borderline or average range due to the protective effect of a second unaffected X chromosome and random X-inactivation (Bartholomay et al., 2019; Kirchgessner et al., 1995; Schmitt et al., 2024).
Given the hallmark intellectual disability phenotype, the accurate assessment of IQ in FXS is integral to its clinical characterization. IQ scores, although just one criterion for having an intellectual disability, continue to guide educational and vocational supports and placements, inform clinical prognosis, and have implications for accessing services and other benefits (e.g., Supplemental Security Income) (Greenspan et al., 2015; Silverman et al., 2010). IQ scores also are used in research studies in FXS beyond clinical characterization, and often are correlated with primary measures and biomarkers like electrophysiological output (Pedapati et al., 2022) and fragile X messenger ribonucleoprotein (FMRP) and FMR1 RNA expression (Boggs et al., 2022; Schmitt et al., 2024; Straub et al., 2023) to determine the extent to which biomarkers co-vary with the hallmark phenotype.
One known challenge with assessing IQ in FXS is inadequate sensitivity to individual differences in IQ among individuals with moderate-to-severe intellectual disabilities (Hessl et al., 2009). This results in a floor effect, especially among adults with FXS due to age-associated “declines” in IQ driven by limited growth in skills relative to same-age peers. For example, as measured by the Stanford-Binet, Fifth Edition (SB-5) (Roid, 2003), which already has the benefit of an extended floor relative to other IQ instruments, 50% of the males and 7% of the females with FXS in the present study have a full-scale IQ of 40 (the floor of the SB-5). One well-established solution to this floor effect is the transformation of raw scores to z-scores based on deviation from the original normative sample (“deviation IQ”), which extends the floor of observed IQ scores. This method has been described for both the SB-5 and Weschler Scales (Hessl et al., 2009; Sansone et al., 2014), and is publicly available for use via the PRO-ED SB-5 scoring software. This extended range of IQ scores reduces floor effects and produces a near-normal distribution of deviation IQ scores in individuals with FXS (Schmitt et al., 2024). The SB-5 excels at meaningfully capturing cognitive differences in individuals with intellectual disability, and it is the only publicly available method for calculating deviation IQ (the use agreement for the Wechsler Scales was not renewed and thus the deviation IQ scoring method can only be calculated for the SB-5) (Sansone et al., 2014). This makes the SB-5 a valuable approach to measuring IQ in FXS.
The SB-5, like other IQ measures, allows for the calculation of an “abbreviated IQ” score, a combination of performance on the two Verbal Knowledge and Nonverbal Fluid Reasoning routing subtests (Roid, 2003). Administration of these two subtests provides a briefer estimate of IQ than comparable instruments like the Wechsler Abbreviated Scale of Intelligence, Second Edition (WASI-II) which is made up of four subtests (Wechsler, 2011). In our extensive experience, the routing subtests can be completed in 10–15 min in FXS, compared to the 60–90 min required for the full battery. Abbreviated IQ (ABIQ) has thus become a widespread method for brief cognitive screening in FXS across clinical and research settings (Abbeduto et al., 2021; Norris et al., 2022; Shaffer et al., 2020).
Despite its prevalence, the accuracy of SB-5 ABIQ in estimating FSIQ for individuals with FXS or intellectual disabilities more broadly is not known. The accuracy of ABIQ is better understood in autistic individuals, a population that significantly overlaps with FXS, but which has lower prevalence of co-occurring intellectual disability. For example, ABIQ over-estimates FSIQ by an average of 3.5 points in autistic preschoolers (FSIQ 40 to 117, M = 70, 42.5% with FSIQ ≥ 70) (Twomey et al., 2018). Stephenson and colleagues further demonstrated that ABIQ over-estimated FSIQ in autistic individuals and individuals with ADHD who have a high degree of scatter between the two routing subtests (FSIQ 40–133, M = 78) (Stephenson et al., 2023). These findings suggest ABIQ may over-estimate FSIQ in children with neurodevelopmental disabilities and that the magnitude of this over-estimation is greater when there is a large discrepancy between verbal and nonverbal skills. This is relevant to individuals with FXS, especially males, who often show greater verbal relative to nonverbal skills (Freund & Reiss, 1991; Huddleston et al., 2014). Knowledge of the extent of this error would allow for data-informed decision-making when choosing between the administration of an abbreviated or full SB-5 battery in FXS in clinical or research settings.
To support this data-informed decision-making, the present study aimed to determine the extent to which ABIQ under- or over-estimates FSIQ in individuals with FXS. We also aimed to assess whether this pattern varies by age, the degree of routing subtest scatter, and FSIQ (i.e., a proxy for ID severity). Based on previous findings in autistic youth (Twomey et al., 2018), we hypothesized that ABIQ would over-estimate FSIQ in children with FXS. Based on our clinical experience administering the SB-5 to adults with FXS, we also hypothesized that ABIQ would under-estimate FSIQ in adults with FXS and in those with more severe ID (i.e., lower FSIQ).
Participants included all male and female patients with FXS seen in the Cincinnati Fragile X Research and Treatment Center who were administered the Stanford-Binet, Fifth Edition (SB-5) full battery between 2015 and 2025. All participants had a confirmed diagnosis of FXS, defined as having the full FMR1 mutation (> 200 CGG repeats), confirmed via past testing results made available in a participant’s medical record or via Southern Blot and/or PCR conducted in collaboration with the Molecular Diagnostic Laboratory at Rush University. Only participants younger than 50 years of age were analyzed due to a small number of participants (N = 6) over the age of 50. A subset of participants (N = 61) provided clinical data across multiple 198 unique participants provided data across 281 visits (see Results for additional details). Full demographic details are reported in Table 1. Parents or legal guardians of participants younger than 18 years of age and legal guardians of adult participants with limited decision-making capacity stemming from their intellectual disability provided written informed consent. Adult participants otherwise provided written informed consent. All participants provided assent when possible given each participant’s expressive and receptive language abilities. All study procedures were approved by the Cincinnati Children’s Hospital Medical Center Institutional Review Board.
All participants completed the full version of the SB-5 (Roid, 2003), administered by either a licensed clinical psychologist, supervised post-doctoral clinical psychology fellow, or supervised clinical research coordinator. Deviation scores for the SB-5 were calculated using previously reported methods validated in FXS to minimize floor effects common in this population (Sansone et al., 2014). Subtest scatter was calculated as the absolute difference between scaled scores on the Verbal Knowledge and Nonverbal Fluid Reasoning subtests. In our Supplementary Results, to compare routing Verbal Knowledge performance with VIQ and Nonverbal Fluid Reasoning performance with NVIQ using the same scale, scaled scores for each routing subtest were transformed to standard scores. These are referred to as “abbreviated VIQ” and “abbreviated NVIQ”, respectively, for ease of reference.
Three separate mixed effects models were used to examine whether differences between full-scale and abbreviated IQ estimates (e.g., FSIQ minus ABIQ) varied by age, subtest scatter, or FSIQ. Linear and quadratic effects for age, subtest scatter, and FSIQ were examined and models with the strongest fit were implemented (see Results). When linear and quadratic effects provided models with similar fit, the more parsimonious linear effect was examined. Participant/subject was the random effect in all models to account for participants who were tested at multiple visits.
Given the unique cognitive profiles seen across males and females with FXS (Freund & Reiss, 1991; Huddleston et al., 2014), we also report whether the above associations were different between males and females (i.e., age × sex; subtest scatter × sex; FSIQ × sex) in our Supplementary Materials. Identical models in our Supplementary Materials also were used to examine the extent to which Verbal Knowledge and Nonverbal Fluid Reasoning routing subtests accurately estimated VIQ and NVIQ, respectively, and whether these estimates varied as a function of age and sex.
Changepoint detection was implemented to determine the age, degree of subtest scatter, and FSIQ score at which the mean difference between FSIQ and ABIQ estimates significantly changed. At Most One Change (AMOC; i.e., one changepoint) was allowed based on visual inspection of the linear/quadratic fit (see Figs. 1, 2 and 3) and to maximize interpretability and clinical utility.
All statistical analyses were conducted using R version 4.3.1 (2023). Mixed effects models used the lme4 R package (Bates et al., 2015). Changepoint detection used the changepoint R package (Killick et al., 2022). All data was visualized using the ggplot2 R package (Wickham, 2016).
Sixty-one (61) participants provided data across multiple timepoints representing 144 total repeated visits. The number of repeated visits ranged from 2 to 4 (M = 2.5 visits, SD = 0.6 visits). The average time between repeated visits ranged from 265 to 2856 days (M = 698.6 days, SD = 435.0 days). Change in ABIQ between subsequent visits ranged from 0 to 48.9 (M = 9.5, SD = 8.8). Change in FSIQ between subsequent visits ranged from 0 to 33.0 (M = 7.3, SD = 6.7).
ABIQ was within 1 SD (15 standard score points) of FSIQ for 77.2% of all timepoints (N = 217). We first examined whether the accuracy of ABIQ in estimating FSIQ varied as a function of age. A model including a quadratic age effect provided a better fit (AIC = 1992.4, BIC = 2010.6, log-likelihood = −991.18, marginal R^2^ = 0.263, conditional R^2^ = 0.693, adjusted ICC = 0.584) than a model including a linear age effect (AIC = 2026.8, BIC = 2041.4, log-likelihood = −1009.42, marginal R^2^ = 0.378, conditional R^2^ = 0.707, adjusted ICC = 0.529; χ^2^(1) = 36.474, p < .001). The degree to which ABIQ accurately estimated FSIQ changed non-linearly with age (F(1,221.71) = 38.998, p < .001; Fig. 1). Changepoint detection indicated that ABIQ over-estimated FSIQ in participants younger than 11.5 years of age (vertical dashed line in Fig. 1), after which ABIQ under-estimated FSIQ by an average of 9.2 standard score points (SD = 9.1, range = −15.5 to 35.4).
We also examined whether the accuracy of ABIQ in estimating FSIQ varied as a function of routing subtest scatter. A model including a linear subtest scatter effect provided a similar fit (AIC = 2042.7, BIC = 2057.3, log-likelihood = −1017.4, marginal R^2^ = 0.179, conditional R^2^ = 0.663, adjusted ICC = 0.589) as a model including a quadratic subtest scatter effect (AIC = 2044.7, BIC = 2062.9, log-likelihood = −1017.4, marginal R^2^ = 0.179, conditional R^2^ = 0.663, adjusted ICC = 0.589; χ^2^(1) < 0.001, p = .983). The more parsimonious model including a linear fit was selected. The degree to which ABIQ accurately estimated FSIQ changed linearly with subtest scatter (F(1,278.079) = 61.274, p < .001; Fig. 2). The greater absolute scatter there was between the verbal and nonverbal routing subtests, the more ABIQ tended to under-estimate FSIQ. Changepoint detection indicated that ABIQ began to consistently under-estimate FSIQ by an average of 12.5 standard score points (SD = 9.7, range = −21.7 to 27.9) after the difference between the routing subtests was greater than 5.5 scaled score points. Subsequent visual inspection of the raw scatter between subtests (i.e., non-absolute value) indicated that for most participants, this underestimate was occurring in participants whose verbal routing performance was greater than their nonverbal routing performance.
We also examined whether the accuracy of ABIQ in estimating FSIQ varied as a function of FSIQ (i.e., differences varying as a function of ID severity). A model including a linear FSIQ effect provided a similar fit (AIC = 2090.4, BIC = 2105.0, log-likelihood = −1041.2, marginal R^2^ = 0.027, conditional R^2^ = 0.674, adjusted ICC = 0.665) as a model including a quadratic FSIQ effect (AIC = 2090.8, BIC = 2108.9, log-likelihood = −1040.4, marginal R^2^ = 0.034, conditional R^2^ = 0.661, adjusted ICC = 0.650; χ^2^(1) = 1.658, p = .198). The more parsimonious model including a linear fit was selected. The degree to which ABIQ accurately estimated FSIQ changed linearly with FSIQ (F(1,241.082) = 6.521, p = .011; Fig. 3). ABIQ tended to under-estimate FSIQ more in participants with lower FSIQ. Changepoint detection indicated that this under-estimate was more consistent in those with FSIQ < 68. In participants with FSIQ < 68, this underestimate was an average of 6.3 standard score points (SD = 11.6, range = −23.1 to 35.4).
To inform clinical decision-making when administering the SB-5 to individuals with FXS (i.e., determining a full vs. abbreviated battery), we examined differences between FSIQ and ABIQ in nearly 200 participants with FXS. We also determined whether characteristics like age, scatter between routing subtests, and ID severity could inform an administrator’s decision to use ABIQ or administer the full SB-5 to calculate FSIQ. We have two primary findings from this analysis. First, ABIQ often over-estimates true abilities (FSIQ) before age 11 in individuals with FXS, after which ABIQ almost universally under-estimates FSIQ, suggesting over- or under-estimation varies across development, likely as skills emerge that are only captured by a full administration (Fig. 1). Second, the degree to which ABIQ under-estimates FSIQ worsens as the split between verbal and nonverbal skills grows and in individuals with more severe ID, providing simple indicators during routing subtest administration to determine if a full SB-5 battery is needed (Figs. 2 and 3). Based on these findings, we make three recommendations for clinicians and clinical researchers.
As the developers of the SB-5 acknowledge, ABIQ is only an estimate of FSIQ and errors in its estimation are expected (Roid, 2003). FSIQ itself is only an estimate of true cognitive abilities, although we treat it as “ground truth” for the purposes of discussion here. However, we demonstrate that errors in the estimation of FSIQ by ABIQ likely vary across development in individuals with FXS. Our finding that the SB-5 ABIQ score over-estimates FSIQ in school-age children with FXS is consistent with findings in autistic youth (Twomey et al., 2018). When conducting evaluations, we therefore recommend a more conservative a full administration may avoid over-estimating abilities in children and under-estimating abilities in adults.
The issue of age-related under-estimation and over-estimation is amplified in males relative to females (see Supplementary Materials). At the lower end of the performance distribution, small raw score differences translate into disproportionately large differences in standard scores. In males who have more significantly affected cognitive skills and who measure at the lower of the performance distribution (Schmitt et al., 2024), small differences in raw performance thus translate to large differences in estimated abilities. This effect is demonstrated in our finding that the under-estimation of skills by ABIQ is more drastic in individuals with lower FSIQ (Fig. 3).
Differences between the verbal and nonverbal routing subtest’s prediction of VIQ and NVIQ described in our Supplementary Materials clarify what is likely driving the under-estimation of skills by ABIQ in adolescents and adults with FXS. The accuracy of the Verbal Knowledge routing subtest in predicting VIQ is more consistent across development, whereas the Nonverbal Fluid Reasoning subtest begins to under-estimate NVIQ beginning in adolescence. This suggests there are nonverbal skills not captured by this lower ABIQ score in adolescents and adults that are otherwise captured by FSIQ. As Fluid Reasoning is already reflected in the ABIQ score, this suggests performance on the other domains (Nonverbal Knowledge, Quantitative Reasoning, Visual-Spatial Processing, or Working Memory) may be stronger.
There may be times when clinicians or researchers only have ABIQ available for analysis (e.g., a full administration is terminated early; retrospective studies). If presented with these issues, we encourage users to recognize the limitations of the ABIQ estimate and take appropriate precautions when interpreting the data. Within the context of a research study, this limitation should be explicitly reported. Ultimately, this underscores the importance of evidence-based psychological assessment that bases recommendations on the results of a comprehensive battery without relying on an IQ score as the sole data point (Wright et al., 2022). When other data points (e.g., adaptive behaviors, measures of co-occurring symptomatology) are available for clinical characterization, this other data may be weighed more heavily when interpreting research findings or making clinical recommendations.
Our finding that this ABIQ under-estimate is worse when routing subtest scatter is high and in individuals with more severe ID gives administrators a concrete decision point after completion of the routing subtests. If the difference between the nonverbal and verbal routing subtests is less than 5.5 scaled score points and if an individual appears to have an IQ > 68, ABIQ often adequately estimates FSIQ. However, under- or over-estimation of FSIQ by ABIQ may remain due to age-related concerns described above, and ABIQ should still be interpreted with caution.
Several limitations of our study inform directions for future research. First, many of our participants performed at the floor of the SB-5 (50% of males, 7% of females). This precluded our ability to conduct our analysis using the original SB-5 scoring method because there is a 7-point difference between the lowest possible FSIQ (40) and ABIQ scores (47) that would create an artificial “over-estimate” of 7 points in these participants. This would minimize individual differences in raw performance that are otherwise captured by the deviation IQ scoring method (Hessl et al., 2009; Sansone et al., 2014). We are interested in similar studies using the original scoring method in a sample with a greater number of participants with FXS performing above the SB-5 floor. Second, this study was conducted on a predominantly white and non-Hispanic sample, consistent with studies finding race- and ethnic-based disparities in FXS service access (Crawford et al., 2002; Kidd et al., 2017). The degree to which these findings generalize to a more representative sample is unknown. Last, IQ scores are only one broadband metric of cognitive abilities, and our paper frames as FSIQ as the truest available estimate of cognition for the sake of interpretability. It is imperative to recognize that IQ tests and scores are imprecise, reductionistic, and historically rooted in eugenics and other discriminatory practices (Au, 2013). Despite these issues, IQ testing remains a foundational, but imperfect, component of neurodevelopmental assessments and informs decision-making for families and providers.
Based on findings from a sample of nearly 200 participants with FXS, we demonstrate errors in the estimation of FSIQ by ABIQ when using the SB-5 in FXS. The SB-5 remains a high-quality instrument, but ABIQ estimates should be interpreted with appropriate precautions in FXS given these findings. We generally recommend that administrators aim to complete the full SB-5 when possible. When this is not possible, administrators should acknowledge potential sources of error in ABIQ estimates of intelligence. Our findings highlight the complexity of assessing uneven cognitive profiles in individuals with neurodevelopmental disabilities such as FXS and the care that must be taken when interpreting testing results.
Supplementary Information The online version contains supplementary material available at https://doi.org/10.1007/s10803-025-07062-w.