Authors: Laurin Mordhorst (Department of Neuroradiology, University of Luebeck, Luebeck, Germany; Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Luke J. Edwards (Department of Cognitive Neuroscience, Faculty of Psychology and Neuroscience, Maastricht University, Maastricht, The Netherlands; Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany), Maria Morozova (Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; Paul Flechsig Institute—Centre of Neuropathology and Brain Research, Leipzig, Germany), Mohammad Ashtarayeh (Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Tobias Streubel (Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Björn Fricke (Department of Neuroradiology, University of Luebeck, Luebeck, Germany; Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Francisco J. Fritz (Department of Neuroradiology, University of Luebeck, Luebeck, Germany; Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Henriette Rusch (Paul Flechsig Institute—Centre of Neuropathology and Brain Research, Leipzig, Germany), Carsten Jäger (Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; Paul Flechsig Institute—Centre of Neuropathology and Brain Research, Leipzig, Germany), Ludger Starke (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany), Thomas Gladytz (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany), Ehsan Tasbihi (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany; Charité—Universitätsmedizin Berlin, Berlin, Germany), Joao S. Periquito (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany), Andreas Pohlmann (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany), Herbert Mushumba (Institute of Legal Medicine, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Klaus Püschel (Institute of Legal Medicine, University Medical Center Hamburg-Eppendorf, Hamburg, Germany), Thoralf Niendorf (Berlin Ultrahigh Field Facility (B.U.F.F.), Max Delbrück Center for Molecular Medicine in the Helmholtz Association, Berlin, Germany; Charité—Universitätsmedizin Berlin, Berlin, Germany), Nikolaus Weiskopf (Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; Felix Bloch Institute for Solid State Physics, Faculty of Physics and Earth System Sciences, Leipzig University, Leipzig, Germany; Wellcome Centre for Human Neuroimaging, UCL Queen Square Institute of Neurology, University College London, London, UK), Markus Morawski (Department of Neurophysics, Max Planck Institute for Human Cognitive and Brain Sciences, Leipzig, Germany; Paul Flechsig Institute—Centre of Neuropathology and Brain Research, Leipzig, Germany), Siawoosh Mohammadi (Department of Neuroradiology, University of Luebeck, Luebeck, Germany; Institute of Systems Neuroscience, University Medical Center Hamburg-Eppendorf, Hamburg, Germany; Max Planck Research Group MRI Physics, Max Planck Institute for Human Development, Berlin, Germany)
Categories: Research Article, axon radius, microstructure, diffusion-weighted MRI, light microscopy, histology, validation
Source: Imaging Neuroscience
Doi: 10.1162/IMAG.a.1030
Authors: Laurin Mordhorst, Luke J. Edwards, Maria Morozova, Mohammad Ashtarayeh, Tobias Streubel, Björn Fricke, Francisco J. Fritz, Henriette Rusch, Carsten Jäger, Ludger Starke, Thomas Gladytz, Ehsan Tasbihi, Joao S. Periquito, Andreas Pohlmann, Herbert Mushumba, Klaus Püschel, Thoralf Niendorf, Nikolaus Weiskopf, Markus Morawski, Siawoosh Mohammadi
The axon radius holds promise as a clinical MRI biomarker for neurological disorders. However, in-vivo MRI estimation appears infeasible on clinical scanners and lacks experimental validation. Crucially, existing histology is only sparsely sampled, enabling primarily qualitative assessment. Here, we use large-scale human brain histology, sampling 46 million axons across 35 corpus callosum regions with MRI-like sizes. By demonstrating a significant spatial correlation with histology on an advanced research scanner, we provide quantitative proof that MRI radius estimates reflect underlying microstructure—a critical milestone. The next milestone—translation to clinical scanners—appears feasible with now-available high-gradient systems according to simulations, but would require substantial SNR gains. Yet, we also identify a sensitivity bottleneck in current modeling that may offer a complementary path to improved sensitivity through future modeling advances. Overall, we provide promising evidence for the validity of MRI-based axon radius estimation and identify challenges that must be solved for clinical adoption.
Axons are critical to neural communication, with their radii influencing communication speed (Waxman, 1980). Axon radii vary spatially along and across white matter fiber bundles (Caminiti et al., 2009; Tomasi et al., 2012; Veraart et al., 2021) but also change temporally over the lifespan. While typical changes occur during development and aging, others can indicate neurodevelopmental disorders (Stassart et al., 2018; Wegiel et al., 2018) or neurodegenerative diseases (Evangelou et al., 2001; Judson et al., 2017), positioning the axon radius as a potential clinical biomarker.
This biomarker might be measurable via water diffusion-weighted magnetic resonance imaging (dMRI). In the dMRI-based signal models relevant here, axons are represented as cylindrical structures, with diffusion perpendicular to the cylinder axes reflecting the axon radius (Alexander et al., 2010; Assaf et al., 2008; Sepehrband, Alexander, Kurniawan, et al., 2016; Veraart et al., 2020). Since axon radii are micrometer-sized—much smaller than millimeter-scale dMRI voxels—the dMRI signal reflects a combined contribution from the individual axons within a voxel. This combined contribution of all axons within a voxel has been proposed to be captured in a scalar metric, the effective axon radius (reff ) (Burcaw et al., 2015; Veraart et al., 2020), which is heavily influenced by the largest axons, representing the tail of the axon radius distribution.
Estimating reff in real-world dMRI experiments, and axon radius metrics more broadly, has proven challenging throughout the past decades, mainly due to dMRI’s weak sensitivity to small axon radii (Huang et al., 2015; Nilsson et al., 2017; Sepehrband, Alexander, Kurniawan, et al., 2016; Veraart et al., 2020). In-vivo, the gradient amplitudes required to exploit this sensitivity have only recently become accessible with the advent of a handful of specialized research scanners reaching up to 300 mT/m (Fan et al., 2020; Pizzolato et al., 2023; Veldmann et al., 2024; Veraart et al., 2020, 2021). While current clinical systems are limited to gradient amplitudes up to 80 mT/m , next-generation scanners reaching up to 200 mT/m may help bridge this gap and open new possibilities for clinical translation. In parallel with these hardware advances, the understanding of dMRI signal contributions has evolved. One line of work revisits the core intra-axonal signal model, using simulations of at most a few hundred axons (M. Andersson et al., 2020, 2022; H.-H. Lee, Papaioannou, et al., 2020; H.-H. Lee et al., 2019, 2024; Winther et al., 2024) to explore how complex axonal morphology alters the dMRI signal relative to the simplified cylinder assumption underlying reff . Another line of research focuses on confounding signal contributions from outside the axon, such as axonal surface relaxation effects (Barakovic et al., 2023), additional signal compartments (Alexander et al., 2010; Palombo et al., 2020; Pizzolato et al., 2023; Stanisz et al., 1997; Veraart et al., 2019, 2020), orientation dispersion (Drobnjak et al., 2016; Nilsson et al., 2012), and Rician noise bias (Gudbjartsson & Patz, 1995), with various strategies proposed to address these confounding signal contributions via experimental design (Veraart et al., 2020, 2021), modeling (Barakovic et al., 2023; Jespersen et al., 2013; Kaden et al., 2016; Mollink et al., 2017; Pizzolato et al., 2023), and processing (Varadarajan & Haldar, 2015).
Despite these challenges, spatial trends in reff have emerged with some consistency, particularly in the corpus callosum (Fan et al., 2019; Horowitz, Barazany, Tavor, Bernstein, et al., 2015; Pizzolato et al., 2023; Veraart et al., 2021). In rats, a low-high-low profile has been reported (Barazany et al., 2009), aligning well with histological findings (Barazany et al., 2009; Veraart et al., 2020). However, similar patterns in humans (Aboitiz et al., 1992; Barakovic et al., 2021, 2023; Caminiti et al., 2009; Horowitz, Barazany, Tavor, Bernstein, et al., 2015; Huang et al., 2015; Pizzolato et al., 2023; Veraart et al., 2021) and nonhuman primates (Caminiti et al., 2009; Lamantia & Rakic, 1990) appear less consistent. Critically, these dMRI-histology comparisons have remained largely qualitative rather than quantitative—mainly due to limitations in current histological datasets, particularly for humans. These comparisons rely on 2D histology, which attempts to sample axon radius distributions orthogonal to the fiber direction in highly aligned white matter such as the corpus callosum (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014; Veraart et al., 2020). However, in humans, these datasets typically include only a small number of regions of interest (ROIs) (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014), precluding meaningful spatial correlation analysis. Moreover, the ROIs often cover only a fraction of a dMRI voxel, which—due to the tail-weighting of reff —makes estimates prone to strong statistical fluctuations (Barakovic et al., 2023; Veraart et al., 2020). In rats, broader spatial sampling has been attempted, including 20 ex-vivo dMRI voxel-sized ROIs in the corpus callosum (Veraart et al., 2020). Yet, the variation in reff across these ROIs is minimal, and the analysis largely focuses on distribution-level agreement rather than ROI-wise correspondence, limiting conclusions about spatial sensitivity and potentially masking systematic error, for example, through compensating contributions of confounding factors.
Recently, 3D histology has emerged as a complementary approach (M. Andersson et al., 2020; H.-H. Lee et al., 2019; Shapson-Coe et al., 2024; Tian et al., 2025), though only few studying human tissue (Shapson-Coe et al., 2024; Tian et al., 2025). These data reveal complex axonal morphology, including undulations and radius fluctuations, inaccessible to 2D histology. However, current 3D datasets cover at most three ROIs in animals (M. Andersson et al., 2020; H.-H. Lee et al., 2019)—or even just a single ROI in humans (Shapson-Coe et al., 2024; Tian et al., 2025)—and have largely been used to simulate morphological effects on dMRI signal formation, rather than for direct experimental validation.
In this study, we aim to provide the first quantitative proof of sensitivity of dMRI-based reff measurements to underlying tissue microstructure. To this end, we assess quantitative spatial correlations between reff from experimental in-vivo and ex-vivo dMRI against corresponding histological values, enabled by a densely sampled light microscopy dataset comprising 35 in-vivo dMRI voxel-sized ROIs across two human corpora callosa. First, we establish our light microscopy-based histology dataset as a benchmark for experimental reff validation by demonstrating gains in accuracy and precision over existing datasets (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014). We then show that reff estimates from in-vivo dMRI significantly correlate with histological values, despite a newly identified model-inherent sensitivity reduction in the dMRI-based reff. Finally, we show that next-generation clinical scanners—now available on the market—may bring in-vivo reff mapping within the realm of clinical adoption, although substantial improvements in SNR remain necessary.
Here, we provide a brief introduction to the dMRI-visible effective axon radius reff and the dMRI signal model used for its estimation. For a more detailed description, see Appendix A.1. The dMRI-visible effective axon radius
reff=〈r6〉〈r2〉4.(1)
is a scalar, tail-weighted statistic of the axon radius distribution within a dMRI voxel (Burcaw et al., 2015; Veraart et al., 2020). reff can be estimated in a regime of strong diffusion weighting (b), suggested to be around b ≥ 6 ms/µm2 for in-vivo dMRI (Veraart et al., 2019, 2020) and b≥20 ms/µm2 for ex-vivo dMRI (Veraart et al., 2020). In this regime, one can approximate the powder-averaged dMRI signal as
S∘(b)≈βb⋅e−bDa⊥+fim,(2)
where β is a signal scaling factor, Da⊥ is the intra-axonal perpendicular diffusivity, and fim is the signal of the immobile water compartment (Alexander et al., 2010; Stanisz et al., 1997). reff is directly linked to Da⊥ through
reff=487δ(Δ−δ3)Da⊥D04,(3)
where δ is the diffusion gradient duration, Δ is the diffusion gradient separation, and D0 is the diffusivity of the axoplasm (Veraart et al., 2020). Using Equations (2) and (3) and an estimate for D0, one can jointly estimate reff and β, for example, via non-linear fitting (Veraart & Novikov, 2019).
We used two human corpus callosum tissue samples, CC-01 and CC-02. See Table 1 for sample information.
We immersion-fixed whole brains in 3%. paraformaldehyde and 1% glutaraldehyde in phosphate-buffered saline (pH 7.4). We then extracted the corpora callosa and bisected them along the mid-sagittal plane, yielding hemispheric sections. We prepared one hemispheric section from each donor for histology and preserved the other hemisphere for ex-vivo dMRI. In the histology hemisphere, we cut a slice of the tissue sample orthogonal to the mid-sagittal plane and extracted 35 ROIs in total, each roughly including the cross-sectional area of in-vivo dMRI voxels used in our study (see Fig. 1). The ROI segments were then contrasted with osmium tetroxide and uranyl acetate, dehydrated in graded acetones, and embedded in Durcupan resin. For imaging with light microscopy, we cut semi-thin sections (≈500 nm thickness) parallel to the mid-sagittal plane on a Reichert Ultractut II. The sections were mounted on Thermo Scientific SuperFrost Plus glass slides, stained with 1% toluidine blue, air dried, and coverslipped with Sigma-Aldrich Entellan toluene.

We acquired light microscopy images (one per ROI) using a Zeiss AxioScan Z1 (objective: 40× , numerical 0.95, 0.1112 µm /pixel; resolution 292 nm ). An example image is shown in Figure 1b.
We segmented axons using a deep-learning-based method (Mordhorst et al., 2022) (see Fig. 1c) and derived empirical axon radius distributions by calculating the radius of circles with equivalent areas for each axon. For comparison with in-vivo dMRI, we compensated for tissue shrinkage by scaling each axon radius (scaled r′=1.3r ), where the factor 1.3 was estimated as the mean of previously reported values (Aboitiz et al., 1992; Tang et al., 1997). Finally, we computed reff values from the empirical axon radius distributions using Equation (1).
For ex-vivo dMRI, we used the hemisphere of CC-01 that was not processed for histology (see Fig. 2e for the bisection into hemispheres). Sample information is provided in Table 1.

We cut the tissue sample into five segments along the anterior-posterior axis (see Fig. 2b) using a Reichert Ultracut II, and embedded the segments in 1.5% agarose in phosphate buffer saline in a custom-made container.
We acquired magnitude dMRI data using a Bruker Biospin 9.4T scanner with a single-channel transceiver volume coil and a gradient insert coil with a maximum gradient amplitude of 1500 mT/m at the Berlin Ultrahigh Field Facility in Berlin, Germany. We followed a protocol similar to that suggested by Veraart et al. (2020) for ex-vivo reff mapping in rats. Briefly, we applied diffusion-weighting for 65 gradient directions per b, with directions based on the approach by Dubois et al. (2006). We used a segmented EPI sequence with four segments and the following fixed δ=7 ms , Δ=20.1 ms , echo time TE=34.7 ms , repetition time TR=[15000, 25000] ms (segment-dependent), and isotropic voxel edge length of 0.35 mm . For different tissue segments, the field-of-view was adjusted between 22 × 28 × 9 mm3 and 25 × 30 ×
10.5 mm3. We varied b between 2.5 and 100 ms/µm2, and gradient amplitude (g) between 200 and 1278 mT/m , as detailed in Table 2. To enhance SNR, we averaged repeated measurements prior to image reconstruction for higher b-values, as shown in the Repetitions column of Table 2.
We corrected for Gibbs ringing artifacts (Kellner et al., 2016; Tournier et al., 2019). To account for signal drift across b-shells, we normalized images within each b-shell to an S(b=0) image acquired at the start of acquisition for that shell.
Per b, we estimated the noise level σ^ using Marchenko-Pastur principal component analysis (Cordero-Grande et al., 2019; Tournier et al., 2019; Veraart et al., 2016) prior to preprocessing. Then, we estimated S∘(b) as the zeroth-order spherical harmonic using an estimator of the even order spherical harmonic coefficients up to the sixth order. Specifically, we determined the spherical harmonics basis functions (Coelho et al., 2022, 2023; Novikov et al., 2018; Reisert et al., 2017) and estimated the coefficients using a Rician maximum likelihood estimator (Varadarajan & Haldar, 2015), which relied on the b-dependent σ^ maps.
For b≤10 ms/µm2, we estimated the main fiber direction μ→ using NODDI (Zhang et al., 2012a, 2012b), using the σ^ map for b=2.5 ms/µm2.
For b≥20 ms/µm2, we estimated reff from S∘(b) using Equations (2) and (3) via non-linear fitting (Veraart & Novikov, 2019), assuming D0 =0.35 µm2/ms (West et al., 2018). We estimated fim (see Equation (2)) from strongly decayed directional signals. To this end, we selected signals from the highest b-shell with high alignment between g→ and μ→ (angle ≤20° ), fitted a Rician distribution, and approximated fim as its expected value.
We recruited five healthy adult subjects (age: 31 ±3 years, representing mean ± standard deviation; 2 male, 3 female).
We acquired magnitude dMRI data using a 32-channel receive coil and 300 mT/m gradient coils on a Siemens Connectom 3T scanner at the Max Planck Institute for Human Cognitive and Brain Sciences in Leipzig, Germany. We followed the dMRI protocol described by Veraart et al. (2021). Briefly, we used a single-shot multi-band echo-planar imaging (EPI) sequence with blipped-CAIPI (multi-band 2) and in-plane GRAPPA acceleration (acceleration 2). We applied diffusion-weighting with the following fixed δ=15 ms , Δ =30 ms , TE=66 ms , TR=3500 ms , matrix size of 88×
88 with 54 slices, and isotropic voxel edge length of 2.5 mm . We varied b={0.5, 1, 2.5, 6, 30.45} ms/µm2 for {30, 30, 30, 60, 120} gradient directions isotropically distributed on a sphere (Jones et al., 1999) and used variable gradient amplitude g={36, 51, 80, 124, 279} mT/m .
For geometric susceptibility correction, we acquired 23 non-diffusion-weighted images with the same and 10 images with reverse phase encoding. Additionally, we acquired T1-weighted MP-RAGE images (Brant-Zawadzki et al., 1992).
We corrected for Gibbs ringing artifacts (Kellner et al., 2016; Tournier et al., 2019), eddy current and motion artifacts (J. L. R. Andersson & Sotiropoulos, 2016; J. L. R. Andersson et al., 2016; Tournier et al., 2019), and gradient non-linearity distortions (Janke et al., 2004; Jovicich et al., 2006).
For b ≤2.5 ms/µm2, we estimated the noise level σ^ using Marchenko-Pastur principal component analysis (Cordero-Grande et al., 2019; Tournier et al., 2019; Veraart et al., 2016) prior to preprocessing. After preprocessing, we estimated the apparent diffusion tensor (Basser et al., 1994a, 1994b; Tournier et al., 2019) and mapped fractional anisotropy (FA) (Basser et al., 1994a; Tournier et al., 2019).
For b≥6 ms/µm2, we estimated S∘(b) as the zeroth-order spherical harmonic using an estimator of the even order spherical harmonic coefficients up to the sixth order. Specifically, we determined the spherical harmonics basis functions (Coelho et al., 2022, 2023; Novikov et al., 2018; Reisert et al., 2017) and estimated the coefficients using a Rician maximum likelihood estimator (Varadarajan & Haldar, 2015), which relied on the σ^ maps. Finally, we estimated reff using Equations (3) and (A14), assuming D0=2.07 µm2/ms (Veraart et al., 2018) and fim=0 (Tax et al., 2020).
We conducted simulations to replicate dMRI signal generation and reff -estimation under ex-vivo and in-vivo conditions. The simulations are described in greater detail in Appendix A.2. In brief, we generated dMRI signals for each diffusion gradient direction by computing volume-weighted average signals (Packer & Rees, 1972) over our empirical axon radius distributions. For in-vivo simulations, we used axon radius distributions adjusted for tissue shrinkage as described in the histology section. For signal simulation, we modeled three intra-axonal, extra-axonal, and immobile water compartment with T2-weighted volume fractions fa, fe and fim . We simulated intra-axonal signal using the matrix method (Callaghan, 1997) to capture effects beyond the Gaussian phase approximation (Van Gelderen et al., 1994). We assumed fully decayed extra-axonal signal (but fe >0 ) and used fixed fim . To estimate reff from simulated signals, we followed the procedure for experimental ex-vivo and in-vivo dMRI data, assuming known fim . We considered both an idealized scenario (SNR=∞ ), as well as an “experiment-like” scenario mimicking experimental Rician noise conditions (in-vivo: SNR = 32; ex-vivo: b-dependent SNR ranging from 17 to 51), for which we repeated simulations 1000 times.
For each histology ROI, we empirically determined the sampling distribution of reff for different subsample sizes between 102 and 105 axons, reflecting smaller ROI sizes typical for existing histology data (Aboitiz et al., 1992; Caminiti et al., 2009; Liewald et al., 2014). Per subsample size, we assessed accuracy using the normalized mean bias error
NMBE=∑m=1M(r^eff,m−reff)reff(4)
and precision using the coefficient of variation
CV=std({r^eff,m|m∈M})reff,(5)
where M=1000 , r^eff,m is a subsample estimate, reff is the reference value computed from the full empirical axon radius distribution, and std({r^eff,m|m∈M}) denotes the standard deviation across all subsample estimates. This subsampling analysis assumes zero spatial autocorrelation of axon radii.
across modalities
patterns
To compare the spatial patterns of reff in the mid-sagittal section of the corpus callosum across modalities, we mapped reff values from all modalities onto the mid-sagittal slice in Oxford-MultiModal-1 (OMM-1) atlas space (Arthofer et al., 2024). Specifically, we proceeded as
patterns
To quantitatively compare reff from dMRI experiments to histological values, we determined corresponding reff values in the respective native spaces.
Error metrics Between histological values (reff ) and fitted/estimated values (r^eff ) from dMRI experiments as described in Sections 3.2.4 and 3.3.4, we computed a linear regression slope to assess the scaling behavior of r^eff . Additionally, we determined the fitting success rate
S=1N∑i=1NI(r^eff,i≥0.1 µm),(6)
as the proportion of biologically feasible estimates (r^eff≥0.1 µm , Waxman et al. (1995)), where N=35 is the number of ROIs and I(⋅) is the indicator function, which equals 1 if the condition inside is true and 0 otherwise. ROIs that did not meet voxel-selection criteria (corpus callosum masking or insufficient subject support for group-averages) were also counted as unsuccessful.
To quantify absolute agreement, we computed the normalized root-mean-square
NRMSE=∑i=1N(r^eff,i−reff,i)2∑i=1Nreff,i,(7)
using r^eff,i=0 µm for unsuccessfully fitted values.
To quantify the ability to capture linear relationships, we computed Pearson’s correlation
R=∑i=1N(r^eff,i−〈r^eff〉)(reff,i−〈reff〉)∑i=1N(r^eff,i−〈r^eff〉)2∑i=1N(reff,i−〈reff〉)2,(8)
where 〈reff〉 and 〈r^eff〉 denote the mean histological and estimated reff values across ROIs. To assess statistical significance, we performed a Monte Carlo permutation test under the null hypothesis that r^eff and reff are uncorrelated (R=0 ). We computed the associated p-value
p=1K∑i=1KI(|R′i|≥|R |),(9)
where R′i were computed from shuffled reff and fixed r^eff to approximate the null distribution, using K=106 permutations.
For dMRI simulations with M=1000 repetitions, we pooled over all M×N values to compute the linear regression, S, R and NRMSE . Accordingly, we computed p over M×K iterations with K=1000 so that M×K=106.
The assessment in Section 3.6 revealed a model-inherent bias of dMRI-based reff . To investigate the origins of this bias, we assessed the influence of signal approximations involved in deriving reff (see Appendix A.1). These approximations, applied in sequence,
We evaluated the accuracy of these approximations at the level of the powder-averaged signal S∘(b) (see Equation (2))—the signal to which reff is fitted. While the Taylor approximation and exponential approximation directly model S∘(b) , the matrix method, the GPA and the WPA model signals per gradient direction S(b,g→) as per Equation (A1). For the latter methods, we first simulated S(b,g→) and then estimated S∘(b) using a Gaussian maximum likelihood estimator.
To study reff mapping in anatomies beyond our human corpus callosum dataset, we extended simulations to mimic other axon populations by scaling the radius distributions from our data. Specifically, we scaled axon radii to extrapolate axon radius distributions of the rat corpus callosum (scaling 0.5 (Veraart et al., 2020)) and the human corticospinal tract (scaling 1.15 (Veraart et al., 2021)). For in-vivo simulations, these scaling factors were applied alongside the tissue shrinkage compensation factor (1.3). For example, we used the scaling factor 1.15⋅1.3≈1.5 for in-vivo corticospinal tract simulations.
mapping
We optimized in-vivo dMRI protocols for a range of maximum gradient strengths (gmax=[40,600] mT/m ), covering the capabilities of existing clinical and research 3T scanners. The optimization aimed to maximize R between histological reff and simulated reff for protocol candidates, where the simulations were conducted analogously to those for our experimental protocols described. In contrast to these simulations, we assumed Gaussian rather than Rician noise, as Rician noise more strongly obscures correlation (see Supplementary Section S2) but can, in principle, be mitigated through advanced preprocessing techniques (Eichner et al., 2015; Fan et al., 2020; Manzano Patron et al., 2024).
To streamline the parameter search, we formulated the optimization problem using the constraints of our experimental in-vivo dMRI protocol. Specifically, we focused on two-shell protocols with fixed diffusion timing parameters (δ, Δ) and fixed minimum b (bmin=6 ms/µm2) to suppress extra-axonal signal, while allowing the minimum g (gmin ) to vary. Thus, we modeled the optimization problem as
θ*|gmax=argmaxθR(θ)|gmax, θ={δ,Δ,gmin}.(10)
The search grid for θ and additional parameters are detailed in Table 3. Additionally, we enforced Δ ≥δ+4 ms .
We modeled the effect of the protocol-dependent TE on T2-weighted intra-axonal water fraction (fa) and SNR. The echo time of a protocol candidates was estimated as
TE(θ)=δ+Δ+C,(11)
where the constant C=21 ms , derived from our experimental protocol, accounts for additional contributions, such as the RF pulse and readout gradients. Assuming fim=0 (Tax et al., 2020), we computed fa as
fa(TE)=f0⋅e−TE/T2,af0⋅e−TE/T2,a+(1−f0)⋅e−TE/T2,e,(12)
where f0=0.41 is the non-T2-weighted intra-axonal water fraction, T2,a=82 ms is the intra-axonal T2-value, and T2,e=44 ms is the extra-axonal T2-value, as reported by Veraart et al. (2018). Using SNRref=32 and TE,ref=66 ms , derived from our experimental protocol as reference values, we extrapolated
SNR(TE)=SNRref⋅f0⋅e−TE/T2,a+(1−f0)⋅e−TE/T2,ef0⋅e−TE,ref/T2,a+(1−f0)⋅e−TE,ref/T2,e.(13)
To evaluate a potential elevation of the baseline SNR level through technical or acquisition improvements, we repeated the protocol optimization analyses for SNRref increased by 75% and 150% , yielding SNR values of 56 and 80 for our experimental protocol.
We analyzed 35 light microscopy ROIs from two human corpus callosum samples to establish an in-vivo dMRI-scale histological reference for reff . Figure 3 summarizes how our dataset improves spatial sampling, as well as the precision and accuracy of reff estimation compared to existing data.

Figure 3a and b presents a quantitative comparison of our dataset with existing 2D histology of the human corpus callosum. Our dataset improves spatial sampling both by including a greater number of ROIs and by increasing ROI size, translating into three orders of magnitude more axons per ROI.
through enhanced tail sampling
Figure 3c and d illustrates that light microscopy ROI sizes enable smoother sampling of the tail of the axon radius distribution than ROI sizes used by Aboitiz et al. (1992), which would result in occasional spikes in the tail and deviations in reff . This effect of ROI size on reff is further explored in Figure 3e, which shows sampling distributions of reff computed from repeated sampling at different ROI sizes. Smaller ROIs underestimate reff but increase the likelihood of overestimated outliers (see example in Fig. 3d), indicating lower accuracy and precision. The low accuracy and precision of smaller ROIs is quantified across all ROIs in Figure 3f and g. As ROI size increases, both accuracy (bias) and precision (coefficient of variation) improve, with accuracy improving more rapidly. For ROI sizes of existing histology data (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014), the expected bias would be 4 to 12% , whereas the expected coefficient of variation would be 14 to 21% .
across modalities
To validate reff experimentally against our histological reference, we acquired in-vivo and ex-vivo dMRI data. In-vivo, we scanned five healthy human subjects using a high-gradient scanner (300 mT/m ), whereas we conducted ex-vivo dMRI scans on tissue sample CC-01 on a preclinical scanner (9.4 T ). As a bridge between histology and dMRI, we conducted dMRI simulations under both idealized (SNR = ∞) and experiment-like conditions with added Rician noise (in-vivo: SNR = 32; ex-vivo: SNR ranging from 17 to 51, depending on the diffusion-weighting b). For a fair comparison between in-vivo dMRI experiments/simulations and histology, we scaled axon radii from histological distributions by 1.3 to account for tissue shrinkage (Aboitiz et al., 1992; Tang et al., 1997). Figure 4 compares spatial reff patterns across modalities, whereas Figure 5 provides a quantitative comparison.

Figure 4a and b show the spatial patterns of reff in histology across the mid-sagittal section of the corpus callosum. Both samples exhibit similar inter-region trends, with an alternating low-high pattern across the anterior midbody, midbody, posterior midbody, and splenium (see also Supplementary Section S3). However, there is strong intra-region variability within the splenium, inconsistent across tissue samples. In other regions, intra-region variability cannot be conclusively assessed due to sparser sampling.
pattern at reduced dynamic range
Figure 4c–j compare spatial reff patterns across the corpus callosum between histology, dMRI experiments and simulations, both for the ex-vivo (see Fig. 4c–f) and in-vivo scenario (see Fig. 4g–j). Ex-vivo dMRI-based reff underestimate histological values (see Fig. 4d), aligning with simulations (see Fig. 4e, f). While in-vivo simulations also predict an underestimation of reff (see Fig. 4i, j), experimental reff overestimate histological values (see Fig. 4h), indicating effects not captured by simulations. Overall, both ex-vivo and in-vivo reff patterns exhibit a reduced dynamic range compared to histology, suggesting low sensitivity to microstructure (see Fig. 4d–f, h–j). This low sensitivity complicates the capture of spatial patterns ex-vivo (see Fig. 4d–f), whereas the group-average pattern of in-vivo dMRI experiments (see Fig. 4h) shows some resemblance to histology, hinting at a similar alternating low-high pattern across anterior midbody, midbody, posterior midbody, and splenium. However, the high values in the genu, relative to other regions, do not align with histological patterns. Additionally, partial volume effects may influence the in-vivo dMRI pattern, as suggested by extreme values near border regions, an effect that is likely exacerbated by the relatively large voxel size of our in-vivo acquisition.
The resemblance of the group-average spatial reff pattern from in-vivo dMRI experiments with histology is reflected in a significant correlation (see Fig. 5a). However, this analysis exhibited some variability due to the non-deterministic nature of our in-vivo dMRI processing (Fig. 5a shows a representative iteration; over 10 iterations we R=0.414±0.03 , all p<0.05 ; see Supplementary Section S4). The significant correlation of in-vivo dMRI-based reff with histological values was not predicted by our simulations (see Fig. 5b), which, however, assume a single subject rather than a group-average. Supplementary Section S5 provides a more comparable scenario to simulations by showing per-subject correlations, revealing no significant correlation with histology for most individual subjects as predicted by simulations. Yet, in terms of absolute values, there is an offset of about 0.5 µm between reff from in-vivo dMRI experiments and simulations (see Fig. 5a, b). For ex-vivo dMRI experiments (see Fig. 5a), there is no significant correlation with histology. This is likely due to reduced precision compared to simulations (see Fig. 5b) and the need to estimate an additional parameter, fim , which can confound reff estimation (see Supplementary Sections S6 and S7).

The idealized dMRI simulations (see Fig. 5c) reveal a primary cause of the reduced dynamic range of dMRI-based reff : a proportional bias at larger reff , which we refer to as “model-inherent bias”. This bias affects absolute agreement, as measured by the normalized root mean square error (NRMSE), by shifting values below the unity line. Additionally, it reduces R by limiting the dynamic range at the upper end of reff values, thereby obscuring correlations under noisy conditions (see Fig. 5a, b). Notably, the model-inherent bias is stronger ex-vivo than in-vivo (see slopes in Fig. 5c).
additionally reduces the dynamic range
In-vivo experiment-like simulations (see Fig. 5b) show a mild, noise-induced overestimation of smaller reff values, affecting sensitivity at the lower end of reff values. This reduced sensitivity to small reff hints at the practical resolution limit, below which reff values may no longer be reliably distinguished from noise (Nilsson et al., 2017).
The simulations in Figure 5c reveal a model-inherent bias of dMRI-based reff . Here, we investigate the origins of this bias by assessing the signal approximations involved in deriving reff (see Equations (A5), (A6), (A9), and (A10)). Figure 6 shows the powder-averaged signal S∘(b) simulated for both in-vivo and ex-vivo experimental MRI protocols, comparing different levels of signal approximation. The matrix method (Callaghan, 1997) is the reference method for our simulations, whereas remaining methods introduce successive simplifications, as arranged from left to right in the legend, ultimately leading to the signal model used for reff fitting.

Both in-vivo and ex-vivo, the Taylor approximation introduces by far the largest deviations between successive approximations, driving the model-inherent bias. The wide-pulse approximation (WPA) also introduces slight deviations, whereas the Gaussian phase approximation (GPA) aligns almost perfectly with our reference method. While relative differences between approximations are preserved across experimental conditions, deviations are generally stronger for the ex-vivo protocol.
and gmax
The observed scaling of deviations with reff and gmax for the Taylor approximation and the WPA can be understood by examining their underlying dependencies. For the Taylor approximation, accuracy improves when the necessary but not sufficient condition κreff4≪1 is satisfied. While the violation of this condition with increasing reff is evident, its violation with increasing g is implicit in the quadratic dependency g2∼κ=748g2γ2δD0. Similarly, the deviations caused by the WPA (δ>>r2/D0) are related to both reff and gmax . Higher reff are linked to the prevalence of larger r, making the assumption less valid, while higher gmax typically allow for shorter δ, further undermining the assumption. In ex-vivo dMRI, the reduced D0 further amplifies the deviations due to both Taylor approximation and WPA.
Interestingly, S∘(b) appears to decay almost linearly with reff at high b. This behavior suggests an intrinsic property of the signal model, warranting further exploration.
To investigate how model-inherent bias may affect reff measurements outside the human corpus callosum, we extrapolated axon radius distributions from our human corpus callosum dataset to the human corticospinal tract and the rat corpus callosum. We repeated idealized simulations (SNR = ∞) as in Figure 5c, applying additional scaling factors on top of tissue shrinkage correction (only in-vivo; scaling 1.3) to adjust for the target anatomies (rat corpus 0.5 (Veraart et al., 2020)); human corticospinal 1.15 (Veraart et al., 2021)). Figure 7a shows the resulting axon radius distributions and corresponding reff values. Figure 7b and c presents ROI-wise comparisons between simulated and histological reff values for these populations, evaluated under both our ex-vivo and in-vivo protocols.

Across both in-vivo and ex-vivo conditions (see Fig. 7b, c), the broader reff range observed across axon populations reveals a nonlinear scaling of the model-inherent bias. The trend indicates a tendency toward saturation at higher reff values—indicating not just a fixed bias, but a loss of sensitivity as reff increases.
The trends observed previously—stronger sensitivity reduction under ex-vivo conditions—persist across axon populations. In-vivo (Fig. 7c), this implies that brain regions with larger axons, such as the corticospinal tract, may be particularly affected, potentially imposing anatomical constraints on reff mapping. Ex-vivo (see Fig. 7b), sufficient sensitivity appears retained for small-axon populations such as the rat corpus callosum, provided lower reff values remain distinguishable from noise under experimental conditions (Nilsson et al., 2017). However, for human white matter, sensitivity is minimal, raising concerns about whether the remaining sensitivity can be meaningfully exploited in real-world ex-vivo dMRI acquisitions.
mapping
We optimized in-vivo dMRI protocols for next-generation 3T clinical scanners with maximum gradient amplitude (gmax ) up to 200 mT/m . To this end, we conducted a grid search for optimal protocol parameters and evaluated protocol candidates by simulating their reff estimates for our corpus callosum dataset and maximizing the correlation (R) with histological reff , assuming a single dMRI subject. We accounted for SNR variations due to protocol parameter choices, but also considered increased baseline SNR levels, independent of protocol parameters, to explore potential gains achievable through technical or acquisition improvements. In contrast to our dMRI experiments, we assumed Gaussian- rather than Rician-distributed signals, which can be achieved with advanced preprocessing techniques (Eichner et al., 2015; Fan et al., 2019; Manzano Patron et al., 2024). Figure 8 summarizes the protocol optimization results.

Figure 8a and b show R and NRMSE as a function of gmax , contextualizing the achievable performance of next-generation clinical scanners in comparison to state-of-the-art 3T clinical scanners, state-of-the-art research scanners as used in our dMRI experiments, and next-generation research scanners, assuming 90% of nominal gmax values. For any baseline SNR level, R converges to a maximum value at a certain minimum gmax , where NRMSE is also optimal or close to optimal. While state-of-the-art clinical scanners consistently perform well below optimal R and NRMSE values, next-generation clinical scanners achieve R values much closer to those of research scanners and can reach near-optimal NRMSE at higher SNR baseline levels. The protocol parameters and further analyses for all optimal protocols referenced in Figure 8a and b are presented in Supplementary Section S8.
Figure 8c–e show simulated reff for optimal next-generation clinical scanner protocol at each baseline SNR level (corresponding to colored markers in Fig. 8a, b). At SNR≈27 , reflecting the expected SNR of the protocol candidate under our experimental conditions, next-generation clinical scanners would not reveal a significant correlation for a single subject with our histology data (R=0.29 , p=0.13 , see Fig. 8c). However, our simulations suggest that a significant correlation could be revealed at SNR≈48 (R=0.49 , p=3.7e−3 , see Fig. 8d) and a stronger correlation at SNR≈68 (R=0.63 , p<0.05 , see Fig. 8e).
While R remains stable after reaching the optimum at some gmax (see Fig. 8a), NRMSE increases thereafter (see Fig. 8b). We attribute this loss of absolute agreement to the increasing influence of the model-inherent bias, which drives large reff further away from the unity line with increasing gmax (see Supplementary Section S8). Hence, within the employed model’s constraints, model-inherent bias is a relevant factor in protocol design for scanners with very high gmax , such as next-generation research scanners. To fully exploit the potential of these scanners, improved modeling is required.
We addressed the longstanding challenge of quantitatively validating axon radius measurements from diffusion-weighted MRI (dMRI)—a challenge made increasingly pressing by recent advances in acquisition hardware and signal modeling. While previous studies hint at qualitative patterns common with histology, they fall short of quantitatively demonstrating that experimental dMRI captures spatial variation in axon radii—relying instead on simulations, sparsely sampled histology, or histological datasets with insufficient variation to reveal spatial trends. Here, we assessed quantitative spatial correlations between the dMRI-visible effective axon radius (reff ) and densely sampled histology from two human corpora callosa. While ex-vivo dMRI showed no significant correlation with histological values, a significant group-level correlation in-vivo provides the first quantitative evidence that dMRI-based axon radius estimates reflect underlying tissue microstructure in the human brain. While we demonstrate this correlation with in-vivo dMRI data acquired on an advanced research scanner, histology-grounded simulations suggest that emerging high-gradient clinical scanners—now available—may bring in-vivo reff mapping within reach of clinical adoption, although substantial improvements in SNR remain necessary. In particular, clinical adoption may be facilitated by addressing a newly identified sensitivity limitation in current models—caused by how axon populations are collapsed into a single scalar metric—through future modeling advances.
From a mechanistic perspective, our findings experimentally demonstrate that a specific axon morphology marker, reff , leaves a detectable signature in the in-vivo dMRI signal in the human brain. This sensitivity to reff survives despite unmodeled competing intra-axonal signal contributions. In particular, recent 3D histology (M. Andersson et al., 2020; H.-H. Lee et al., 2019; Shapson-Coe et al., 2024; Tian et al., 2025) has challenged the core assumption of perfectly cylindrical axons underlying reff , suggesting substantial impact of complex axonal morphology on the dMRI signal in Monte Carlo simulations (M. Andersson et al., 2020, 2022; H.-H. Lee, Papaioannou, et al., 2020; H.-H. Lee et al., 2019, 2024; Winther et al., 2024). In addition, our study reveals another fundamental a model-inherent proportional bias that reduces sensitivity to reff , potentially challenging application in brain regions with very large axons, such as the corticospinal tract (see Figs. 5c and 7). This bias arises from the reduction of the full axon radius distribution to a single scalar, reff , and scales with its magnitude (see Fig. 6). As such, we expect it to persist even if reff were not only computed from radius distributions across axons, as done here, but also incorporated along-axon radius variations, as recently suggested (H.-H. Lee, Jespersen, et al., 2020).
The sensitivity of in-vivo dMRI to reff persists not only in the face of intra-axonal modeling limitations, but also against a broader array of confounding signal contributions, including unmodeled compartment signals (Alexander et al., 2010; Palombo et al., 2020; Pizzolato et al., 2023; Stanisz et al., 1997; Veraart et al., 2019, 2020), relaxation effects (Barakovic et al., 2023), orientation dispersion (M. Andersson et al., 2022; Drobnjak et al., 2016; Nilsson et al., 2012), partial volume effects (Alexander et al., 2001; Vos et al., 2011), and Rician noise bias (Gudbjartsson & Patz, 1995). In addition, in-vivo dMRI-histology comparisons are further affected by tissue deformation and shrinkage (Aboitiz et al., 1992; Dyrby et al., 2018; Tang et al., 1997; Yendiki et al., 2022), inter-individual differences, inter-cohort differences, and scan-rescan variability, although the latter has been reported to be low for the in-vivo protocol we adopted (Veldmann et al., 2024; Veraart et al., 2021).
This myriad of confounding factors suggests that achieving specificity to reff is challenging and indeed, our in-vivo data hints at such effects. While there appears to be only a slight overestimation of reff compared to histological values, our simulations indicate that the model-inherent bias alone should introduce substantial underestimation. This suggests the presence of an additional overestimation effect of similar magnitude that counteracts the model-inherent bias—highlighting the potential for ambiguities in absolute value range comparisons, such as those conducted in rats (Veraart et al., 2020). The better agreement in absolute values between simulations and ex-vivo dMRI experiments—despite a lack of correlation—raises the possibility that Rician noise bias, more pronounced at the lower SNR of in-vivo dMRI, may drive part of the observed shift. Yet, in light of the complex interplay of additive and compensating confounding effects, attributing such discrepancies to specific sources remains difficult. Hence, given the previously unproven sensitivity of dMRI to reff , and more broadly to axonal morphology in the human brain, the next logical step was to provide quantitative experimental evidence for sensitivity.
The spatial variation we validate largely reflects a coarse low-to-high pattern across the corpus callosum—spanning the anterior midbody, midbody, posterior midbody, and splenium—as suggested by visual inspection. This pattern aligns most closely with previous findings in humans (Horowitz, Barazany, Tavor, Bernstein, et al., 2015) and nonhuman primates (Caminiti et al., 2009; Lamantia & Rakic, 1990). Differences to regional trends reported in other human studies (Aboitiz et al., 1992; Barakovic et al., 2021, 2023; Caminiti et al., 2009; Huang et al., 2015; Pizzolato et al., 2023; Veraart et al., 2021) may stem from variations in anatomical definitions, ROI placement, acquisition protocols, and analysis methods, but could also partially reflect inter-individual variability in axon morphology. Indeed, the described pattern is visible only at the group-level, whereas individual subjects and donors show notable variability (see Fig. 4a, b for inter-subject differences and Supplementary Section S5 for a quantification of inter-donor differences). A further contributor is the age mismatch between in-vivo subjects averaged 31 years, whereas histology donors averaged 61 years. Axon radius is known to increase with age in human white matter (Aboitiz et al., 1996; Fan et al., 2019), but our simulations suggest a small impact on correlations based on an age-related change estimated from a dMRI study (Fan et al., 2019) (see Supplementary Section S9). While part of the variability across subjects and donors may be due to imperfect alignment across subjects and modalities, these findings raise the question of whether finer-grained common spatial patterns exist at all. In light of this variability and the limited number of histology donors in our study, confirming sensitivity of in-vivo dMRI to reff on independent datasets remains among the most immediate priorities.
Ex-vivo dMRI is often used as an intermediate step toward in-vivo validation, as it excludes inter-individual differences and allows higher resolution than in-vivo dMRI; however, our results suggest it introduces its own challenges. Unlike in-vivo, we found no correlation ex-vivo. This may be partly due to the smaller number of ROIs available (15 ex-vivo vs. 35 in-vivo), but more fundamentally our simulations show that model-inherent bias strongly reduces sensitivity to reff under typical ex-vivo conditions, as it scales with gradient amplitude and decreased D0. While ex-vivo dMRI measurements may retain sufficient sensitivity in species with predominantly small axons, such as rats, the sensitivity in human white matter appears prohibitively low. Consequently, validation studies performed in rats (Barazany et al., 2009; Veraart et al., 2020) are limited not only by anatomical differences from humans (Leenen et al., 1982), but also by fundamentally different sensitivity constraints. For the specific protocol we assessed, the sensitivity to reff leaves little headroom to detect correlations if additional unaccounted for effects are present. One strong candidate for such an effect is the presence of an immobile water compartment in ex-vivo tissue, which introduces an additional parameter estimation step—along with its associated uncertainty. Our simulations suggest this uncertainty has stronger impact ex-vivo than in-vivo (see Supplementary Sections S6 and S7). Recent methods propose addressing this confound by modeling the spherical variance rather than the spherical mean of the dMRI signal (Pizzolato et al., 2023; Veraart, Raven et al., 2023), an approach that could be evaluated in future studies.
validation
As a secondary outcome, the correspondence of reff between in-vivo dMRI and our 2D histology suggests that local axon radius distributions from 2D cross-sections are somewhat representative of the full 3D voxel environment forming the dMRI signal. This interpretation aligns with findings from recent 3D histology (M. Andersson et al., 2020), which indicate that axon radius distributions—and by extension reff —may remain relatively stable along fiber bundles, though this has so far only been shown for a limited number (∼50 ) of large axons. Moreover, our approach captures these axon radius distributions more comprehensively than existing 2D histology, as it samples entire cross-sections of in-vivo dMRI voxels rather than small fractions (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014), yielding more accurate (up to 12% ) and precise (up to 21% ) reff estimates. These improvements in accuracy and precision are based on subsampling our large-scale axon radius distributions under the assumption of zero spatial autocorrelation. Deviations from this assumption would alter the variance of reff estimates, increasing it if large axons cluster and decreasing it if they are evenly dispersed. In conclusion, our 2D histology dataset represents a substantial advance over existing human 2D datasets (Aboitiz et al., 1992; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014), providing unprecedented spatial sampling with access to 46 million axons across 35 ROIs. While current 3D histology efforts (M. Andersson et al., 2020; H.-H. Lee et al., 2019; Shapson-Coe et al., 2024; Tian et al., 2025) uniquely enable dMRI simulations (M. Andersson et al., 2020, 2022; H.-H. Lee, Papaioannou, et al., 2020; H.-H. Lee et al., 2019, 2024; Tian et al., 2025; Winther et al., 2024) based on a few hundred complex axonal reconstructions, our large-scale 2D histology approach offers a complementary and scalable route for practical validation of real-world dMRI measurements.
Axon radius measurements have so far been restricted to advanced research scanners with gradient amplitudes of up to 300 mT/m , as used in our study. In contrast, standard clinical scanners typically operate at gradient strengths up to 80 mT/m . However, emerging high-gradient clinical systems—now available on the market—narrow this gap with amplitudes up to 200mT/m . In histology-informed protocol optimization simulations, we show that these scanners could approach the performance of research systems for reff mapping, provided future advances can improve SNR. Yet, these advances would have to be to support a significant subject-level correlation (R=0.49 , p=3.7×10−3 ) with our histology, simulations indicate that a ∼75% SNR increase would be required, reaching SNR≈48 . These requirements may be even higher when unmodeled signal components reduce correlation strength or when Gaussian signal distributions cannot be enforced through advanced preprocessing techniques (Eichner et al., 2015; Fan et al., 2020; Manzano Patron et al., 2024). Overall, these SNR requirements pose a substantial barrier and indicate that near-term clinical translation remains challenging. This is especially true for robust application at the individual-subject level, where our supplementary analyses showed considerable variability in correlations (see Supplementary Section S5). Nonetheless, compared to the prohibitive SNR demands of current clinical systems (Huang et al., 2015; Nilsson et al., 2017; Sepehrband, Alexander, Kurniawan, et al., 2016; Veraart et al., 2020), next-generation clinical scanners with higher gradient amplitudes offer a realistic avenue for translation, provided that further methodological and hardware advances can deliver the necessary SNR gains.
As for candidates to increase SNR, enforcing Gaussian-distributed signals (Eichner et al., 2015; Fan et al., 2020; Manzano Patron et al., 2024) could enable effective SNR gains through denoising, which we did not apply in our experiments or simulations, in order to preserve Rician-distributed signals for Rician maximum likelihood fitting (Varadarajan & Haldar, 2015). Importantly, the identified model-inherent bias highlights a key target for future modeling advances that could help relax SNR demands. Recent developments also promise future SNR gains, for example, a 30% SNR gain was recently demonstrated using advanced acquisition techniques for the same in-vivo protocol applied in our study (Veldmann et al., 2024). Looking further ahead, improvements may arise from ongoing innovations in coil design and gradient hardware (Feinberg et al., 2023; Niendorf et al., 2015). Encouragingly, our simulations in Supplementary Section S10 suggest that meaningful clinical applications could become feasible if such SNR gains are achieved—illustrated by successful discrimination between individuals with autism spectrum disorder and healthy controls within realistic group sizes.
Expanding beyond clinical protocols, we provide broader insights into the design of in-vivo reff measurements, exploring gradient amplitudes ranging from state-of-the-art clinical scanners with 80 mT/m to next-generation research systems with 500 mT/m . Previous studies have primarily evaluated scanner performance in terms of achievable precision as a function of gradient amplitude and noise, typically assuming idealized models (Huang et al., 2015; Nilsson et al., 2017; Sepehrband, Alexander, Kurniawan, et al., 2016; Veraart et al., 2021). In contrast, our data-driven approach uses histology-informed simulations, capturing both sensitivity reduction and loss of accuracy due to model-inherent bias. This bias increases with both gradient amplitude and reff , degrading absolute agreement as measured by normalized root-mean-square error (NRMSE). By contrast, higher gradient amplitudes tend to improve correlation (R) through improved precision, aligning with previous findings (Huang et al., 2015; Nilsson et al., 2017; Sepehrband, Alexander, Kurniawan, et al., 2016; Veraart et al., 2021). However, these gains in R become marginal for gradient amplitudes beyond 300 mT/m in regimes of high SNR and large reff , at the expense of reduced accuracy. This becomes particularly pronounced for next-generation research scanners with gradient amplitudes of 500 mT/m , suggesting that further improvements in modeling are required to fully harness the potential of these scanners.
We focused exclusively on reff , a scalar metric with a direct analytic relation to the axon radius distribution, albeit one that is dominated by its large-axon tail. A less tail-weighted scalar metric than reff can be defined in the narrow-pulse limit (Burcaw et al., 2015; Sepehrband, Alexander, Kurniawan, et al., 2016), although practical measurements are technically demanding (Nilsson et al., 2017). Other strategies attempt to recover full distributions through parametric fits (Assaf et al., 2008), yet the validity of the assumed distributional forms remains uncertain (Sepehrband, Alexander, Clark, et al., 2016). Even if such assumptions were correct, the additional degrees of freedom in the fit and the disproportionate sensitivity of the dMRI signal to larger axons (Burcaw et al., 2015; Neuman, 1974; Van Gelderen et al., 1994) constrain robust inference about the body of the distribution. With regard to clinical translation, the sensitivity of reff to large axons may be advantageous, as large axons are preferentially affected in several neurological disorders (Judson et al., 2017; Stassart et al., 2018; Wegiel et al., 2018; Zikopoulos & Barbas, 2010). Nonetheless, the usefulness of reff as a biomarker in pathological tissue remains to be better understood, as alterations may reflect distinct effects such as caliber variations, beadings, or undulations (M. Andersson et al., 2020, 2022; H.-H. Lee, Papaioannou, et al., 2020; H.-H. Lee et al., 2019, 2024).
We applied spatial smoothing to in-vivo dMRI-based reff maps to mitigate alignment inaccuracies and reduce apparent noise in comparisons with histology (see Supplementary Section S1). This noise is partly a consequence of our interpolation-efficient pipeline, which applies all spatial transformations in a single resampling step, preserving not only fine-grained spatial detail but also noise. While smoothing reduces sensitivity to subtle spatial variation, substantial inter-individual differences already raise the question of whether such common structure exists at all.
For reff estimation, we assume a fixed literature value for the axoplasmic diffusivity (D0), which could vary across brain regions and individuals (Veraart et al., 2018). Our simulations in Supplementary Section S11 suggest that realistic variation of D0 has only a mild impact on dMRI-histology correlations. Beyond D0, we also fixed additional parameters across all voxels in simulations, such as, for example, axonal volume fraction, which cannot capture voxel-wise heterogeneity and may calibrate outcomes to specific tissue regimes.
A further consideration is how individual radii are derived from 2D histology cross-sections. Various approximations exist, most commonly the radius of a circle with equivalent area (used in this study) or the minor axis of a fitted ellipse (Abdollahzadeh et al., 2021; Aboitiz et al., 1992; M. Andersson et al., 2020; Barakovic et al., 2023; Caminiti et al., 2009; Liewald et al., 2014; West et al., 2016). In Supplementary Section S12, we show that reff based on these two approximations are closely aligned, and their choice has little impact on experimental validation results. In contrast, ellipse major axes (Veraart et al., 2020) produced inflated reff values and no significant correlation with group-level in-vivo dMRI, likely reflecting artifacts from non-orthogonal sectioning rather than true axonal morphology (see example in Mordhorst et al. (2022)).
For protocol optimization, we assumed that Gaussian-distributed signals can be achieved through advanced preprocessing techniques (Eichner et al., 2015; Fan et al., 2020; Manzano Patron et al., 2024). However, such methods are not yet widely adopted in practice, and deviations from the Gaussian assumption could further increase the SNR demands of the proposed next-generation clinical scanner protocols. See Supplementary Section S2 for a comparison of Rician and Gaussian noise in our experimental protocols.
We used relatively large voxels in our in-vivo dMRI acquisition compared to our ex-vivo acquisition. This is a common trade-off to account for the high SNR demands of in-vivo dMRI-based axon radius mapping at strong diffusion-weighting (Fan et al., 2019, 2020; Veldmann et al., 2024; Veraart et al., 2020, 2021).
To account for tissue shrinkage for comparison with in-vivo dMRI, we applied a uniform scaling of axon radius distributions. While this may oversimplify biological reality (Horowitz, Barazany, Tavor, Yovel, et al., 2015) and does not capture potential changes in fiber orientation (Yendiki et al., 2022) or non-linear shrinkage of the extra-axonal space (Dyrby et al., 2018), our correlation analysis inherently accommodates such systematic effects.
Refining the signal model appears to be an immediate avenue for improvement—either by incorporating higher-order terms in Equation (A9) or by exploring the apparent linear decay of S∘(b) with reff (see Fig. 6).
Finally, our open-access histology dataset can serve as a benchmark for studies assessing dMRI-based axon radius estimation beyond the particular metric evaluated (reff ). Even beyond the realm of dMRI, it may support efforts to better characterize the tail of the axon radius distribution using parametric descriptions (Sepehrband, Alexander, Clark, et al., 2016).
We provide the first quantitative evidence that in-vivo MRI-visible axon radius estimates reflect underlying microstructure in the human brain. This quantitative proof of sensitivity to axon morphology marks a critical milestone. At the same time, we outline limitations of our validation design, underscoring the need for independent replication. To support such efforts, we release an open-access histological dataset comprising 46 million axons across 35 ROIs for experimental validation. Our findings further suggest that the newest generation of clinical scanners may support clinical adoption with future technical and modeling advances. Overall, our results motivate continued model development and exploration of potential applications of reff as a neuroimaging biomarker.