Authors: Matthew Mattoni (1Temple University, Department of Psychology and Neuroscience, 1801 N Broad St., Philadelphia, PA, USA), Aaron J. Fisher (2University of California-Berkeley, Department of Psychology, 2121 Berkeley Way, Berkeley, CA, USA), Kathleen M. Gates (3University of North Carolina at Chapel Hill, Department of Psychology and Neuroscience, 235 E. Cameron Avenue, Chapel Hill, NC, USA), Jason Chein (1Temple University, Department of Psychology and Neuroscience, 1801 N Broad St., Philadelphia, PA, USA), Thomas M. Olino (1Temple University, Department of Psychology and Neuroscience, 1801 N Broad St., Philadelphia, PA, USA)
Categories: Article, Ergodicity, Idiographic, Heterogeneity, Stationarity, Precision Imaging
Source: Neuroscience and biobehavioral reviews
Authors: Matthew Mattoni, Aaron J. Fisher, Kathleen M. Gates, Jason Chein, Thomas M. Olino
Much of cognitive neuroscience research is focused on group-averages and interindividual brain-behavior associations. However, many theories core to the goal of cognitive neuroscience, such as hypothesized neural mechanisms for a behavior, are inherently based on intraindividual processes. To accommodate this mismatch between study design and theory, research frequently relies on an implicit assumption that group-level, between-person inferences extend to individual-level, within-person processes. The assumption of group-to-individual generalizability, formally referred to as ergodicity, requires that a process be both homogenous within a population and stationary within individuals over time. Our goal in this review is to assess this assumption and provide an accessible introduction to idiographic science (study of the individual) for the cognitive neuroscientist, ultimately laying a foundation for increased focus on the study of intraindividual processes. We first review the history of idiographic science in psychology to connect this longstanding literature with recent individual-level research goals in cognitive neuroscience. We then consider two requirements of group-to-individual generalizability, pattern homogeneity and stationarity, and suggest that most processes in cognitive neuroscience do not meet these assumptions. Consequently, interindividual findings are inappropriate for the intraindividual inferences that many theories are based on. To address this challenge, we suggest precision imaging as an ideal path forward for intraindividual study and present a research framework for complementary interindividual and intraindividual study.
Many theories in cognitive neuroscience, especially in applied clinical, developmental, or social research questions, reflect intraindividual processes. For example, hypotheses about neural mechanisms of psychopathology, brain and behavioral development, and the effects of environmental stressors all reflect processes that occur within individuals. Despite this intraindividual theoretical focus, most study designs examine interindividual variance. Moreover, many studies that test interindividual processes will interpret results with intraindividual speculation or conclusions (e.g., suggesting mechanistic explanations in the discussion). Interpreting interindividual results to support an intraindividual theory relies on an implicit assumption that group-averaged or between-person inferences extend to individual-level, within-person processes. For example, at the between-person level, higher anterior default mode connectivity is associated with depression (Mulders et al., 2015). This may lead to the intuition that when an individual experiences momentary increases in connectivity across these regions, they will experience heightened depressive symptoms. However, this assumption of group-to-individual generalizability, formally referred to as ergodicity (Molenaar, 2004), may not be supported. The (lack of) support for the assumption of group-to-individual generalizability has critical consequences for cognitive neuroscience. If interindividual findings do not generalize to intraindividual processes, many studies may have a mismatch between their intended goal and conclusions with their study design and statistical operationalization, leaving gaps in our understanding of individual-level functioning.
There is increasing attention to how interindividual differences in brain structure and function may limit the ability of group-average brain maps to reflect individuals (Gordon, Laumann, Gilmore, et al., 2017; Horn et al., 2008; Miller et al., 2009; Nakuci et al., 2023; Seitzman et al., 2019). These studies and commentaries have highlighted this neural heterogeneity as a major limitation of group-level analyses and emphasize the need for an individual-level focus. More recently, researchers have further considered the subsequent inability of interindividual study to inform intraindividual inferences, which is especially important for clinical research (Gell et al., 2024; Kraus et al., 2023). This culminated in the Hunter et al. (2024) conclusion that most processes in cognitive neuroscience are unlikely to be ergodic, or demonstrate literal group-to-individual replaceability.
As a result of non-ergodicity, within-person study designs are necessary to achieve the field’s goal of understanding intraindividual phenomena such as brain-behavior mechanisms. This review targets several goals to further aid this research direction. First, the idiographic science literature, largely coming from the field of psychology, has built a substantial conceptual and methodological framework for individual-level study. This provides an excellent resource for parallel goals in cognitive neuroscience, and recent reviews have applied this knowledge to outline important study design considerations for individual-level study (Gell et al., 2024; Kraus et al., 2023). However, there have been few direct descriptions of idiographic research concepts in a cognitive neuroscience context, with existing work largely focusing on the mathematical foundation of ergodicity (Hunter et al., 2024). Here, we provide an accessible introduction to the origins and key concepts of idiographic science. By outlining key theory, our intention is to broaden recognition of why studying the individual is important and why between-person studies are likely insufficient for this goal.
Second, most prior work has focused on limitations of group averages in representing individual persons for issues such as precise functional localization (Dubois & Adolphs, 2016; Foulkes & Blakemore, 2018; Roalf et al., 2024). While this is an important topic on its own, the subsequent problem of applying interindividual results for intraindividual inferences was largely unaddressed in cognitive neuroscience until the discussion of ergodicity by Hunter et al. (2024). As discussed below, ergodicity is inherently an all-or-nothing property, and likely an unachievable property for most human processes (Voelkle et al., 2018). Instead of focusing on ergodicity, we contend that focusing on identifying patterns of processes that may exhibit group-to-individual generalizability is more appropriate. We operationalize patterns as the structure of associations (e.g., bivariate correlation, network or factor model) and direction (Liu et al., 2023). This broader conceptualization of group-to-individual generalizability enables a research direction of empirically testing which, if any, group-level processes may be appropriate for individual-level inferences. Toward this goal, we review a broad literature of brain function and brain-behavior association studies to consider the plausibility of group-to-individual generalizability assumptions.
Finally, we apply the idiographic knowledge foundation to outline a research direction framework for the distinct goals of individual-level study and group-level study, and how they can complement each other. Key to this direction is the growth of precision imaging to better characterize individual-level brain functioning.
Toward these goals, we begin with a conceptual review of the idiographic science literature in psychology to introduce key concepts and their context. We then focus on the implications of group-to-individual generalizability in cognitive neuroscience with an emphasis on brain-behavior relationships. To examine if group-to-individual generalizability is likely to be supported in cognitive neuroscience, we review empirical literature on the two necessary assumptions for group-to-individual between-person homogeneity and within-person stationarity. Through this review, we suggest that for many processes in cognitive neuroscience, group-averages do not reflect individuals and cross-sectional observations do not reflect the dynamic nature of individuals. While problems by themselves, these challenges also indicate that group-to-individual generalizability is unlikely to be supported for many processes in cognitive neuroscience. To address these challenges, we conclude by connecting the goals of idiographic science to the rise of precision imaging, highlighting it as an ideal path forward for study of the individual. We further emphasize the distinct uses of individual-level and group-level research, and how they can complement each other. Overall, this article is intended to bridge idiographic science and cognitive neuroscience to promote attention to individual-level study.
Behavioral sciences have focused on averages for centuries. Most research attempts to identify the average value of a measure in a group, associations between measures in a group, or differences between groups. This approach is the foundation of nearly all empirical research in psychology and cognitive neuroscience. Group-level statistics are popular as they presume random error in single observations and enable programmatic research that builds on findings generalized from prior studies. However, the dominance of this group-level focus has a potential limitation in that a science of averages obscures the individuality of people (Allport, 1937; Cattell, 1952; Fisher et al., 2018; Hamaker, 2012; Lamiell, 1981; Molenaar, 2004; Molenaar & Campbell, 2009). In one example, Allport (1937) discussed the inability of researchers to identify a universal affective value of different colors. Instead, the affective value of green, for example, differed based on context (e.g., traffic light vs. laboratory) and the individual’s personal experiences (e.g., mood, memories, interests). Although intuitive, this outcome was demonstrative of an overarching issue in focusing on the there is no universal meaning of green. Or as Allport put it, “the generalized human mind is entirely mythical” (Allport, 1937, p.5). Allport’s criticism was that generalizing group-level findings to individuals has questionable validity because individuals vary from each other and from themselves across time and contexts. Averaging across individuals, he argued, provides a broad result that does not reflect the dynamic nature of any individual. Instead, he suggested that individuals should be studied intensively over different occasions, contexts, and other factors.
The differing scientific philosophies of studying groups and individuals were coined nomothetic and idiographic, respectively, by Windelband (1894). A nomothetic (Greek root word nomos, meaning law) approach to science is concerned with establishing fundamental truths that are applicable to the population. The cross-sectional studies that dominate psychology and cognitive neuroscience fall under the nomothetic umbrella, and nomothetic statistics are a function of sample means and the deviations of individuals from those means. Alternatively, an idiographic (idios, meaning own or personal) approach to science is focused on intensive study of single individuals across time and contexts. Person-specific research methods, such as single case designs, intensive use of ecological momentary assessment (EMA), and precision imaging are idiographic approaches. In the idiographic case, statistics are a function of intraindividual means and the deviations of time points from those means within individuals. Cattell’s Data Box (Figure 1) is the classic visual depiction of the nomothetic-idiographic distinction (Cattell, 1988; Cattell et al., 1947). It illustrates that the possible data space is a combination of measures, discrete time points, and individuals. Cattell’s Data Box highlights how studying a measure across individuals, with time being held constant (nomothetic), is distinct from studying an individual across time (idiographic). While the nomothetic approach aims for the important goal of understanding human nature, the idiographic approach suggests that a sole focus on interindividual study cannot accomplish this goal. Noting that individuals are dynamic systems and that interindividual and intraindividual variation are independent (Figure 1), idiographic science instead argues that to understand individuals, they must be studied intensively across time, context, and other dimensions. The idiographic approach still seeks generalizable knowledge, but through a different process of identifying intraindividual processes and examining their generalizability in other individuals.
The nomothetic vs. idiographic distinction is ultimately a contrast of deviations of persons from sample means (interindividual variation) and deviations of instances from an individual’s mean (intraindividual variation). This difference is not itself problematic; interindividual and intraindividual studies address different research questions (Cattell, 1952; Epstein, 1983). Using treatment as an example, nomothetic science is concerned with treatment outcomes – are there neural measures that predict differences in treatment response? In contrast, idiographic science examines treatment mechanisms – are there neural processes that change across treatment as a pathway for symptom reduction? Importantly, these different inferences can be complementary. Treatment requires both between-person knowledge, such as risk factors and treatment efficacy rates, as well as within-person knowledge, such as treatment mechanisms and precipitating factors for a given patient. Moreover, many theories, especially when not precisely specified, contain elements of both intraindividual and interindividual variation (Borsboom & Haslbeck, 2024).
As highlighted by Fisher et al. (2018), problems arise, however, when results obtained from interindividual study are interpreted for intraindividual inferences, or vice versa. This misattribution is known as the ecological fallacy (Robinson, 2009; Schwartz, 1994) and is a common inferential error, especially in neuroscience (Cragg et al., 2019). Many theories in psychology (Hamaker, 2012) and cognitive neuroscience focus on intraindividual processes, even when the vast majority of studies examine interindividual variation. While it is enticing to assume that a between-person association found at a group-level would extend to the processes in individuals within that same group, this assumption is only sometimes supported (Molenaar, 2004). A classic example (Hamaker, 2012) is that typing speed and accuracy are positively correlated between individuals (individuals are more or less skilled at typing). However, the inverse is true within individuals (speed vs. accuracy tradeoff). The two processes reflect distinct levels of variance. Differences in the direction of association at between- vs. within-person levels of analysis is an instance of Simpson’s Paradox (Blyth, 1972; Kievit et al., 2013; Wagner, 1982), a phenomenon where relationships in subgroups of data (e.g., observations within individuals) are masked in an aggregate effect (e.g., observations across individuals; Figure 2).
The potential exchangeability of between-person and within-person processes is frequently discussed in terms of ergodicity. Molenaar (2004) introduced the concept of ergodicity to psychology from mathematics (Birkhoff, 1931). The ergodic theorem holds that, given sufficient time and the absence of external forces, dynamic systems will achieve all possible states. Consequently, the long-term time-average of the system will equal the average possible state at any point in time. One example of an ergodic system comes from in a finite container where gas particles move freely and without external force, the average position of a single particle across time (i.e., time-average) is equal to the average position of all particles at any point in time (i.e., group-average).
To bring ergodicity to psychology, Molenaar (2004) contextualized the congruence of time and group averages with the congruence of within-person and between-person variation. He demonstrated that, for a process to be ergodic, it must be equivalent between individuals in the population (homogeneity criterion), and the mean and variance of each variable must be stable over time (stationarity criterion). These properties will be described in more detail and in the context of cognitive neuroscience later. Briefly, the homogeneity criterion is necessary to assume ergodicity because if individual-level models are heterogenous, a group-averaged model will reflect few, if any, of the individuals (Figure 3A). Similarly, stationarity is necessary because if an individual’s structure or pattern of associations systematically changes across time, a time-averaged model will not reflect that individual’s functioning at different points in time (Figure 3B). If either of these assumptions are violated, the process cannot be considered ergodic, because between-person processes cannot be assumed to equal within-person processes (Fisher et al., 2018; Molenaar, 2004). Consistent with the thinking of Allport (1937), Molenaar (2004) suggested that the vast majority of psychological and biological processes are likely nonergodic, citing substantial genetic variability between individuals and development within individuals as easily observable evidence. Moreover, he presented a proof that, even with relatively lenient conditions (every individual shares the same factor structure, parameters are distributed normally, alpha = .05) at least 5% of individuals will significantly differ from the group estimate. In other words, most of our statistical methods include a distributional assumption that inherently prevents ergodicity, an all-or-nothing property.
Generally, ergodicity is an umbrella term for the portability of all statistical observations and phenomena from the group to the individual level of analysis. It represents an empirical question about the variances, covariances, models, networks, and systems inherent to human subjects research (Fisher, Medaglia, & Jeronimus, 2018). True ergodicity is likely an untenable goal for human processes. We do not expect monozygotic twins, let alone larger populations, to behave identically and unchanging over time. With ergodicity likely untenable for human processes, research following Molenaar (2004) has suggested that the equivalence of interindividual and intraindividual processes should be viewed along a continuum, rather than as a binary classification of systems as ergodic or nonergodic (Liu et al., 2023; Voelkle et al., 2014, 2018).
Along this continuum, a more achievable task is testing which, if any, aspects of group-level inferences generalize to individual-level processes. That is, are there any interindividual processes that can be used for intraindividual inferences for most, if not all, individuals, even if the criteria for ergodicity are not met? Toward this goal, we contend that a focus on the congruence of patterns, operationalized as the structure (e.g., linear relationship, network, or factor model) and direction of the relationship is more appropriate (Liu et al., 2023). Generalized inferences would be further strengthened by also specifying bounds for the strength of parameter estimates (e.g., within a standard deviation). Using this framework, group-to-individual generalizability can be empirically tested for different processes. For example, group-to-individual generalizability of the relationship between striatal reward responsiveness and behavioral risk-taking could be tested by examining if the bivariate association is linear, positive, and within specified effect size bounds when examined between-individuals (e.g., N = 300, t = 1) and within-individuals (e.g., N = 30, t = 25). Critically, as interindividual and intraindividual variance are distinct, group-to-individual generalizability must be tested, rather than assumed, for any given process (Fisher et al., 2018). This leaves a challenging task for the field, but also indicates exciting potential for exploring intraindividual variation that has long been understudied (e.g., Lynch et al., 2024).
Without empirical demonstration, group-to-individual generalizability is an unsupported assumption (Molenaar, 2004), and one must be cautious to not misinterpret findings (i.e., ecological fallacy). However, interindividual findings hold substantial value. This value is primarily in their intended interindividual research questions. For example, interindividual study is needed to identify potential neural markers of disorders or predictors of treatment outcomes—inherently between-persons classifications. Additionally, interindividual findings can be useful for constraining hypotheses for intraindividual study (Borsboom & Haslbeck, 2024). A salient example (Borsboom & Haslbeck, 2024) is evidence that individuals who smoke cigarettes are at greater risk for cancer. This interindividual finding identifies cigarette smoke as a relevant exposure that scientists interested in biological causes or mechanisms of cancer can investigate in intraindividual studies. Importantly, this interindividual observation did not provide intraindividual inferences. Rather, it identified a target (i.e., smoke) that can be tested in intraindividual studies for intraindividual observations. As an example in cognitive neuroscience, striatal reward responsiveness is frequently found to be blunted in individuals with depression (Forbes & Dahl, 2012), an interindividual finding. This may lead to an intraindividual hypothesis that blunted striatal functioning is mechanistically related to depression within individuals. However, this hypothesis must then be empirically tested, not assumed. As these are two distinct levels of association, the generated hypothesis has no greater likelihood of being true than any other; the interindividual finding simply serves as a hypothesis-constraining observation. Finally, we also note the potential complementary use of interindividual and intraindividual study through the investigation of cross-level moderators (e.g., identifying neural markers of individuals that exhibit a particular intraindividual process).
Group-to-individual generalizability is a critical assumption for cognitive neuroscience research. Like previous suggestions in psychology (Curran & Bauer, 2011; Gates et al., 2023; Hamaker, 2012), we contend most theories in cognitive neuroscience reflect within-person processes. This is especially true in the context of brain-behavior associations in affective, clinical, developmental, social, and more applied subfields. Research questions are frequently based on a hypothesized neural mechanism for cognitive processes or behaviors, which are inherently within-person processes. For example, theories of biological mechanisms of psychopathology (e.g., decreased reward response in depression), neural development (e.g., changes in structural connections across age), and neural effects of environmental factors (e.g., effects of stress on brain structure) all reflect intraindividual processes. Despite this theoretical focus, these research questions are frequently examined using interindividual designs.
As reviewed, without group-to-individual generalizability, interindividual studies cannot provide valid intraindividual inferences. Thus, with a reliance on interindividual study designs, we may know much less about neural and cognitive processes than we hope (Fisher et al., 2018). It is therefore critical to understand the assumptions for group-to-individual generalizability, and test which processes these assumptions hold for. For instance, Molenaar’s proof demonstrated that frequentist statistical parameter averages hold a distributional assumption that prevents ergodicity. By extension, functional connectivity matrices for any two individuals will almost certainly differ, suggesting that the group-averaged matrix may not reflect individuals. However, it could be tested if the group-level patterns of associations (e.g., nodes present in the default mode network, correlation directions and strength bounds) are consistent with patterns identified in individuals. Group-to-individual generalizability is also crucial to examine for brain-behavior associations. While this has been infrequently studied, Lynch et al. (2024) recently provided an excellent example with findings that, between-individuals, there is a reliable association between the salience network and depression, but within-individuals, there was no significant relationship between expansion of the salience network and depression (i.e., group-to-individual generalizability was not supported). Below, we review studies relevant to the homogeneity and stationarity assumptions of group-to-individual generalizability in cognitive neuroscience, with focuses on fMRI measures and brain-behavior relationships. With robust findings of heterogeneity and non-stationarity, we suggest that group-to-individual generalizability is an unlikely assumption for most processes. Without group-to-individual generalizability, there is a common mismatch between theoretical goals and study design, and an increased focus on idiographic approaches is necessary for individual-level research questions.
The first requirement of ergodicity is that the processes are homogenous between individuals. In addition to preventing ergodic assumptions, heterogeneity also risks that group-averages may reflect few, if any, of the contemporaneous individual-level models in the same group. As most human processes will not meet the ergodicity requirements, researchers are increasingly viewing the homogeneity assumption along a continuum.
Liu et al. (2023) recently outlined four levels of assumptions of homogeneity required for current analytic approaches. The first assumption, strict homogeneity, assumes that both the patterns of relationships (e.g., direction of relationship, network or factor structure) between variables and the parameter values (e.g., correlation strength or factor loading) are identical across individuals. Analytic approaches where data are concatenated to arrive at one “group” results would be an example of an approach that assumes strict homogeneity among individuals. The second, pattern homogeneity, assumes that only the pattern of relationships is constant across individuals, with parameter values being allowed to vary. Approaches built from multilevel modeling assume pattern homogeneity, since all individuals have the same models yet different estimates. Individuals are assumed to have qualitatively the same processes. Third, weak homogeneity, describes analytic approaches that assume that some patterns of processes are similar across individuals, yet the estimates and other structures of relations likely vary. Approaches that assume weak homogeneity include multitask machine learning methods (Caruana, 1997; Fisher et al., 2022) that utilize some information from across individuals to detect any consistent or similar aspects of the pattern of relations. From there, individual-level relations are added as needed. Note that this differs from multilevel modeling which assumes that participants have identical model patterns of relations (i.e., pattern homogeneity). A similar approach is Group Iterative Multiple Model Estimation (GIMME; Gates & Molenaar, 2012), which assumes weak homogeneity in that some relations exist across individuals while others may be unique to some individuals. Approaches that assume weak homogeneity allow for individual-level estimation of effect sizes for all relations. Finally, no homogeneity assumes that both relationship patterns and parameters freely vary between individuals.
Formal ergodicity requires strict homogeneity across the population. However, this assumption is unlikely to be met for most human processes. For group-to-individual generalizability in cognitive neuroscience, pattern homogeneity is likely a more appropriate assumption to test. For individuals to demonstrate pattern homogeneity, individuals must share the same (or largely similar) structure of relationships, but homogeneity requirements for parameter estimates are loosened. With pattern homogeneity, the aggregate model will reflect the structure of individual’s models. For example, in the left panel of Figure 3A, all individuals share the same network structure but different weights, and the aggregate network structure consequently reflects the structure of each individual. In contrast, the right panel shows pattern heterogeneity of a network structure, and the aggregate model fails to represent any single individual. Following this framework, researchers could examine the homogeneity of patterns of task-related BOLD activation maps, regions belonging to a large-scale connectivity network, or regions with decreased cortical thickness in one behaviorally-defined group vs. another. In each of these examples, the key consideration is the structure of relationships, with more stringent tests also considering bounds of the effect size. Beyond being a requirement for group-to-individual generalizability, examining pattern heterogeneity is necessary to understand if group-level models reflect individual-level models (e.g., do individuals share the same set of connections in a large-scale network). In the case of heterogeneity, individual-level inferences may be inaccurate and comparisons of parameter estimates may be inappropriate if a model does not match the individual. On the other hand, research can examine predictors or outcomes of different structural patterns
One key context where heterogeneity is frequently examined is the localization of brain function in response to a task (Horn et al., 2008; Miller et al., 2002). Task-based studies typically estimate the group-average voxel-wise BOLD response to a particular stimulus to identify where activation occurs (Fedorenko, 2021). However, interindividual differences in functional neuroanatomy threaten the ability for an average location to reflect individuals. While the field largely relies on spatial normalization to account for morphological differences between individuals (Friston et al., 1995), normalization does not account for the anatomical and functional heterogeneity frequently present within a given brain area. Adjacent regions within the same macroscopic anatomical region may have different, or even opposing functional responses (Fedorenko & Blank, 2020). Consequently, voxel-wise averages risk aggregating functionally heterogenous regions across participants, leading to inaccurate group estimates (Fedorenko, 2021). Evidence of heterogeneity in BOLD response patterns is widespread across numerous fMRI tasks, including language (Fedorenko & Blank, 2020; Heun et al., 2000), visual processing (Swisher et al., 2007), face recognition (Kanwisher et al., 1997), memory retrieval (Feredoes & Postle, 2007; Miller et al., 2009), motor functions (Gordon, Laumann, Gilmore, et al., 2017), and decision making (Nakuci et al., 2023). Several of these studies explicitly found the group-level activation map did not reflect individual-level maps (Feredoes & Postle, 2007; Gordon, Laumann, Gilmore, et al., 2017; Miller et al., 2002, 2009; Nakuci et al., 2023; Seghier et al., 2007). To address challenges related to pattern heterogeneity of brain function, the field has developed analytical and study-design methods that attempt to parse functional signal from distinct neuroanatomy, such as hyperalignment (Feilong et al., 2018; Guntupalli et al., 2016; Haxby et al., 2011) and localizer tasks (Fedorenko, 2021; Kanwisher et al., 1997).
Heterogeneity has also received substantial attention in the context of functional connectivity, where high dimensionality increases the number of ways in which individuals can differ. Gordon et al. (2021) delineated three manners of functional connectivity variation, particularly in the context of large-scale brain networks. These include (1) spatial variation, where functional nodes differ in anatomical location or size, (2) topological variation, where functional networks are comprised of different nodes, and (3) connectivity strength variation, where the strength of correlation between nodes differs. Spatial and topological variation can be considered forms of pattern heterogeneity, and strength variation can be considered parameter heterogeneity. Without spatial and topological homogeneity, group-level connectivity networks are at risk of not reflecting relationship patterns of individuals. Moreover, spatial and topological variation are especially important as they may drive strength variation and be more relevant for behavioral associations (Bijsterbosch et al., 2018).
Studies have consistently found that large-scale brain networks and network topography systematically differed between individuals in location and features, and that the group model obscured these individualized patterns (Braga et al., 2020; Feilong et al., 2018; Gordon, Laumann, Adeyemo, et al., 2017; Gordon, Laumann, Gilmore, et al., 2017; Greene et al., 2020; Seitzman et al., 2019). Gordon et al. (2021), for example, found that the precuneus node of the parietal memory network is located on the medial surface of the brain for some individuals, and wraps around to the lateral surface for others. Our work has also found substantial between-person differences in effective connectivity network models during a reward task, and showed that the group-level directional connectivity network poorly reflected most individuals (Mattoni et al., 2023). Similar findings of heterogeneity have been found in studies of structural connectivity, where Gu et al. (2021) and Sun et al. (2022) found idiosyncratic patterns of structural connections and that these differences were associated with functional connectivity heterogeneity. Group-to-individual generalizability thus appears to be unsupported in most studies of brain network patterns. Overall, while group-to-individual generalizability must be tested in each process individually, prior work raises doubts for many areas.
Heterogeneity between individuals may also threaten group-to-individual generalizability of brain-behavior associations. Variation in brain-behavior associations can take two forms, equifinality and multifinality (Cicchetti & Rogosch, 1996). Equifinality (also referred to as degeneracy; Price & Friston, 2002) suggests that distinct biological processes can result in the same cognitive or behavioral outcome (many-to-one). Equifinality of brain structure and function has been demonstrated in cognitive processes, such as emotion (Doyle et al. 2022), as well as clinical disorders such as depression (Drysdale et al. 2016), attention-deficit/hyperactivity disorder (ADHD; Costa Dias et al. 2015), and autism spectrum disorder (Shan et al. 2022). Equifinality presents substantial challenges to identifying specific brain-behavior associations (Bernard, 2023; Feczko et al., 2019), since if individuals differ in the neural correlate for a given phenotype, the average brain-behavior association may not reflect the biological processes underlying phenotypes in individuals (i.e., Simpson’s Paradox). See Feczko et al. (2019) for an extended discussion of equifinality in clinical neuroimaging research.
In contrast to equifinality, multifinality is a phenomenon where the same process can result in different behavioral outcomes across individuals (one-to-many). Multifinality is a substantial concern for brain-behavior prediction, as there are numerous possible phenotypes that may be associated with a given neurobiological measure. For example, decreased hippocampus volume between-individuals relative to healthy controls is associated with higher risk for depression (Videbech & Ravnkilde, 2004), anxiety (van Tol et al., 2010), ADHD (Hoogman et al., 2017), and other phenotypes (though, this example could also be explained by higher-order clinical dimensions; Zald & Lahey, 2017). Given a single individual with smaller than average hippocampus volume, we cannot confidently predict the specific phenotype of the individual, lessening the value of group-level findings. Together, equifinality and multifinality present substantial challenges for individual-level prediction (Westlin et al., 2023), and likely contribute to overall low prediction accuracy and association strength between fMRI measures and most cognitive or behavioral measures (Marek et al., 2022; Schulz et al., 2024).
In sum, there is widespread evidence of robust individual differences in brain structure, function, and brain-behavior associations. This pattern heterogeneity limits the ability of group-level models to explain or account for individual-level models. Moreover, evidence of heterogeneity suggests that many processes of brain function and brain-behavior associations do not meet the pattern homogeneity assumption for group-to-individual generalizability. The lack of group-to-individual generalizability has critical consequences, such as limiting efficacy of group-based neuroimaging for clinical translation (Fisher et al., 2018; Kraus et al., 2023). This potential limitation is especially important as between-person heterogeneity appears to be greater in cortical association areas of the brain (Finn et al., 2015; Mueller et al., 2013; Seitzman et al., 2019) that are frequent targets for study of behavioral and clinical prediction.
Despite these limitations, group-based research is ideal for certain research questions. Studies of individual differences, group classification, and risk identification are inherently interested in interindividual variance. For example, early life adversity is associated with individual differences in cortico-limbic functional connectivity and psychopathology (Brieant et al., 2021), suggesting an early risk factor affecting later brain functioning and mental health. On the other hand, for research questions interested in intraindividual processes, such as how stress affects the brain, within-person research designs are necessary.
Homogeneity also is important to consider with data-driven machine learning approaches, which are becoming increasingly popular. Machine learning approaches in cognitive neuroscience largely include supervised methods that classify individuals into categories such as diagnoses or predict behavioral outcomes (e.g., Bondi et al., 2023; Marek et al., 2022),or unsupervised approaches that separate individuals into more homogenous subgroups (e.g., Drysdale et al., 2017; Gates et al., 2017). Two observations are important to consider in the context of these machine learning goals. First, group-to-individual generalizability is solely based on the source of information, not analytical method. The issues discussed in this paper apply to any statistical method. Second, predicting individual differences in classifications or outcomes and identifying subgroups of individuals both rely on interindividual information. Thus, machine learning approaches face the same challenges of group-to-individual generalizability as frequently relied upon statistical relations that exist at the between-individual level may not apply within individuals examined across time. Moreover, this challenge may be exacerbated in classification methods that include a large number of features or interactions, as highly complex process are unlikely to be homogenous across individuals. In the context of subgrouping approaches, one promising direction is clustering individuals with the explicit goal of creating ergodic subgroups (Golino et al., 2022). However, while it is plausible to identify homogenous processes in smaller sets of individuals, it must be stressed that even here the between-person results do not necessarily inform the within-person process. For example, if individuals in a subgroup characterized by higher functional connectivity in a cognitive control network have greater anhedonia (Tozzi et al., 2024), this does not necessarily mean that any individual’s functional connectivity changes across time with increases or decreases in anhedonia. Like with traditional statistical methods, to use machine learning to make intraindividual inferences, applications to intraindividual data is likely necessary (e.g., Leenaerts et al., 2024).
The second assumption for group-to-individual generalizability involves stationarity. Stationarity can be tested with increasing first-order stationarity, where means are time-invariant; second-order (“weak”) stationarity, where means, variances, and covariances are time-invariant; and strict stationarity, where the joint cumulative distribution is time-invariant (Gates et al., 2023). Typically, when researchers test for stationarity, they examine weak stationarity. Molenaar (2004) demonstrated that second-order stationarity (referred to as stationarity from here on for brevity) is required for ergodicity. With the stationarity assumption met, observing an individual at one instance would not produce a substantively different conclusion than if they were observed at any other instance. In Cattell’s Data Box, single observations from one person could be replaced with another from a different point in time (same row), or the average across time (Figure 3B, left panel). Without stationarity, the parameters governing the process change across time and cross-sectional studies would fail to reflect the dynamic nature of the individual, or an observation at a different point in time (Figure 3B, right panel). Moreover, if individuals change across time, a cross-sectional correlation does not reflect the true between-person effect nor a “middle-ground” of within-person and between-person effects (Hamaker, 2023). Thus, non-stationarity limits both group-to-individual generalizability and the ability to make clear between-person inferences.
Whereas the homogeneity assumption for group-to-individual generalizability requires congruent patterns across individuals, stationarity requires congruent patterns across time. As Molenaar (2004) noted, since neurobiological and psychological processes change substantially across human development, they are unlikely to be stationarity across longer time spans, such as years. Childhood and adolescence in particular are characterized by dramatic changes, including widespread changes in gray matter (Bethlehem et al., 2022; Houston et al., 2014; Mills et al., 2016), maturation of white matter (Lebel & Deoni, 2018), and a shift to more distributed functional networks (Fair et al., 2009; Power et al., 2010). Development continues through adulthood, with findings of robust neurobiological changes across the entire lifespan (Giorgio et al., 2010; Mills et al., 2016; Petrican et al., 2017). Developmental non-stationarity limits the ability to generalize inferences obtained at one age to other ages and creates the risk that cross-sectional studies that average across different ages may obscure effects present at specific ages (i.e., Simpson’s Paradox). For example, Shaw et al. (2006) found that the relationship between cortical thickness and general intelligence was negative in early childhood, but positive in later childhood. Furthermore, when they averaged across all ages, they did not find any significant relationship between cortical thickness and general intelligence. Similar changes in the relationship between cortical thickness and cognitive performance have also been found at different stages of adulthood (de Chastelaine et al., 2019). There are several other examples of brain-behavior relationships changing over time, including findings that age moderates the associations between resting state connectivity and personality traits (Simon et al., 2020) and intelligence (Lindbergh et al., 2019). Overall, these examples as well as a general recognition that the brain, behavior, and environments all change across development suggest that stationarity is a dubious assumption for much of cognitive neuroscience at longer time scales.
At shorter time scales (i.e., days, weeks, or months), evidence for stationarity is more mixed. While measures of brain morphology may be expected to be stationary over short time periods, there are diurnal changes in brain volume thickness that must be accounted for (Murata et al., 2024; Nakamura et al., 2015). Large-scale brain network organization appears to be more stable over shorter time scales, though this is variable across different networks. In one example, network stability was measured in the MyConnectome Project, which examined brain functioning of a single individual for over 100 assessments across 500 days (Poldrack et al., 2015). While there was substantial stability of network functioning between assessments, Laumann et al. (2015) found higher intersession variability of network connectivity in visual, somatomotor, and dorsal attention networks relative to other regions. Notably, a substantial amount of both interindividual and intraindividual variance is attributable to state-like factors, mainly arousal (Bijsterbosch et al., 2017; Glasser et al., 2018; Laumann et al., 2017; Liu et al., 2018; Orban et al., 2020; Schneider et al., 2016; Tagliazucchi & Laufs, 2014). Additional variance may be explained by other factors such as task state, thought patterns, and affect (Gonzalez-Castillo et al., 2015, 2019, 2024; Gratton et al., 2018; Vatansever et al., 2020; Wang et al., 2018), though potentially to a lesser extent (Gratton et al., 2018; Laumann et al., 2024).
Other insights into brain network stationarity may come from retest reliability studies. However, it should be noted that retest reliability does not distinguish whether differences are due to changes in signal (i.e., non-stationarity) or measurement noise, and thus must be interpreted with caution. Meta-analytic work has suggested overall poor retest reliability for both functional connectivity (Noble et al., 2019) and task-based activation (Elliott et al., 2020). However, as retest reliability for functional connectivity can be dramatically improved with longer scanning durations (Laumann et al., 2015). Similarly, retest reliability of task-based activation may be improved by increasing the number of task events and analytical methods that account for cross-trial variability (Chen et al., 2021; Freund et al., 2024). Overall, while more intensive longitudinal research is needed, early results suggest that some aspects of brain functioning, such as network organization, may be stable over short periods. However, the evidence is inconclusive and variable across different processes. Investigators must be careful to consider the specific process and timescale at hand and avoid assumptions of stationarity without prior empirical demonstration.
In sum, stationarity at longer at timespans (i.e., years) is a tenuous assumption due to robust neurodevelopment across the lifespan. At shorter timespans, some processes likely exhibit stationarity while others may not, but more work is needed. Moreover, even for stationary processes, the homogeneity assumption is also required to assume group-to-individual generalizability. For non-stationary processes, single observations may not reflect observations at other points in time or overarching dynamic patterns and assumptions for group-to-individual generalizability are not met (Figure 3B). To more precisely examine within-person processes, longitudinal studies are necessary. Longitudinal study designs also improve interindividual inferences, as they permit disaggregation of between-person and within-person effects and allow for research questions such as differences in relationships across stages of development.
The previous sections highlight that few processes in cognitive neuroscience will exhibit both between-person homogeneity and within-person stationarity, undermining efforts to generalize from groups-to-individuals. The predominant focus on between-person designs is thus inappropriate for understanding within-person functioning for most human processes. Without group-to-individual generalizability, an increased focus on the individual in both theory and study design is necessary. Fortunately, focused study of individuals is not a novel concept for cognitive neuroscience, and the field is well suited for this goal. Foundational discoveries that are the basis of introductory textbooks largely come from N-of-1 studies, including the theme of brain localization from Phineas Gage (Harlow, 1999), the hippocampus’s role in memory from H.M. (Corkin, 1984), and the role of Broca’s Area in language (Berker et al., 1986). While these examples resulted from unique circumstances with extraordinary effect sizes, they provided lasting impacts in our understanding of the brain and demonstrate the high potential of intensive study of individuals. Recently, the study of individuals has returned to cognitive neuroscience in the form of precision imaging (Gordon, Laumann, Gilmore, et al., 2017; Gratton et al., 2022; Laumann et al., 2015; Michon et al., 2022; Naselaris et al., 2021), also referred to as intensive or deep imaging (Kupers et al., 2024).
Precision imaging inverts the typical allocation of resources that maximizes sample size (Figure 1, y-axis) to instead focus on the intensive scanning of single individuals (Figure 1, z-axis). Consistent with an idiographic approach to science, its overarching goal is to fully characterize the neural functioning of single individuals and avoid inaccuracies introduced by group averaging and standardized anatomical atlases (Fedorenko, 2021; Seitzman et al., 2019). Precision imaging studies scan individuals for longer periods of time, typically on multiple occasions, providing hours of data. While most studies focus on resting state scans, others implement task-based designs as well, better enabling comparisons of context-dependent brain functioning (Flournoy et al., 2024; Gordon, Laumann, Gilmore, et al., 2017; Gratton et al., 2018). Classic examples of precision imaging studies include the aforementioned MyConnectome project (Laumann et al., 2015; Poldrack, 2021; Poldrack et al., 2015) and the Midnight Scan Club (Gordon, Laumann, Gilmore, et al., 2017). These studies produced important insights into brain measurement and functioning, including the amount of data necessary for reliable functional connectivity estimation, robustness of individual differences in brain network topography, and overlap between task activation and intrinsic resting state networks. Since these demonstrations of the feasibility and value of precision imaging, several reviews have highlighted it as an important path forward for cognitive neuroscience (Fair & Yeo, 2020; Gell et al., 2024; Gratton et al., 2020a; Kraus et al., 2023; Laumann et al., 2023; Michon et al., 2022). See Gratton & Braga (2021) for an overview of a special issue that reviews precision imaging, including its history, uses, technicalities, implications, and applied examples.
Briefly, precision imaging is an ideal approach to studying the individual for several reasons. First, longer scanning times and multi-echo imaging reduce noise in functional signals to provide more reliable characterization of individual-level brain functioning (Gordon, Laumann, Adeyemo, et al., 2017; Lynch et al., 2020). Second, precision imaging addresses anatomical heterogeneity by localizing brain activation or networks using BOLD patterns that are specific to each individual. For instance, functional localizer tasks can increase precision of individual-level BOLD responses (e.g., Fedorenko et al., 2010), and precision functional mapping can better characterize individual-level network functioning (Gordon, Laumann, Gilmore, et al., 2017). Individualized functional localization can have critical implications. For example, Cash et al. (2021) found that transcranial magnetic stimulation (TMS) treatment for depression was more effective for individuals when stimulation was applied closer to their functionally defined dorsolateral prefrontal cortex, rather than a location defined from a group-average. Due to the benefits of individual-level precision, several commentaries have emphasized precision imaging for clinical translation (Gratton et al., 2020b; Kraus et al., 2023; Laumann et al., 2023).
Third, repeated sampling permits longitudinal assessment of dynamic changes that are core to an idiographic approach to science. For example, the previously mentioned study by Lynch et al. (2024) found that expansion of the salience network was a reliable between-person correlate of depression across multiple precision imaging and consortia datasets. However, in an individual scanned 62 times across 1.5 years, salience network expansion was not associated with depressive symptoms. Instead, functional connectivity between the nucleus accumbens and anterior cingulate cortex predicted depressive symptoms, suggesting that depression and functional connectivity patterns have distinct interindividual and intraindividual relationships. In another example, Flournoy et al. (2024) scanned 30 adolescents for 10 assessments across a year during an emotional processing task. They found low retest reliability for BOLD responses to aversive stimuli across the brain. Conversely, they found strong internal consistency within sessions and that monthly variation in mood, sleep, and stress were associated with within-person changes in BOLD responses. These results suggest that the study had low validity for between-person comparisons but was well suited for within-person study.
Additionally, while intensive sampling approaches have mostly been used in the study of intrinsic functional connectivity, they also provide opportunity for task-based designs. With intentional experimental design, numerous and varied tasks can more richly examine different aspects, contexts, and dynamic processes of the targeted process (Kupers et al., 2024). For example, Allen et al. (2022) showed participants tens of thousands of natural scene images, enabling more robust examination of the visual system within individuals. Moreover, trial-level analysis (Chen et al., 2021) of task-based studies can provide intensive time series of neural data and behavioral responses. These time series can enable intraindividual study and be more practical than numerous repeated scans for certain research questions. For example, trial-by-trial modeling has been used for study of many processes such as reinforcement learning, working memory, and cognitive control (Cho et al., 2022; Colas et al., 2022; Pessoa et al., 2002). More recently, Mistry et al. (2024) also compared within-person and between-person relationships using trial-level analysis to find that the association between inhibitory control-related BOLD activation in cognitive control regions and reaction time was positive at the within-person level, but negative at the between-person level. Like most studies, however, the within-person association was aggregated across individuals, leaving uncertainty how well this pattern reflects different individuals. Idiographic study of trial-by-trial brain-behavior associations and the heterogeneity of these patterns across individuals is a promising and understudied direction for task-based studies.
Finally, repeated sampling better allows for mechanistic examination using experimental manipulation to obtain large effect sizes. Functional connectivity, for instance, appears to respond quickly and drastically to experimental manipulation. Newbold et al. (2020) found that casting the dominant arm changes the connectivity of somatomotor regions, including within-person changes in connectivity correlation coefficients ranging from .22–.86 within a single day across participants. Similarly, Siegel et al. (2024) observed very large effects in a precision imaging study on the within-person effects of psilocybin on network synchrony.
With these key benefits, precision imaging has been presented as an exciting path forward in cognitive neuroscience (Fair & Yeo, 2020; Gratton et al., 2022; Michon et al., 2022). Gratton et al. (2022) specifically identified precision imaging and population-based (e.g., consortia) studies as key directions for detecting reliable brain-behavior associations, as they increase power by improving signal-to-noise ratio and increasing sample sizes, respectively. Gratton et al. (2022) emphasized that precision imaging and population studies are also useful for different research questions and goals, which we note mirrors the idiographic and nomothetic distinction. Precision imaging enables deep characterization of the individual over time and context, providing an ideal study design for intraindividual processes such as cognitive mechanisms, development, or within-person brain-behavior associations. Alternatively, population studies provide the sample size to power more reliable detection of between-person associations, providing insights such as risk-factor identification that can inform public policies.
Gell et al. (2024) recently expanded on Gratton et al. (2022) by further discussing how individualized and consortia studies can reciprocally complement each other with mechanistic study and generalizability tests. This parallels early suggestions by Allport (1937) of the complementary use idiographic and nomothetic approaches. Here, we conclude with a heuristic (Figure 4) that builds on these previous frameworks with an additional focus on intraindividual vs. interindividual inferences, and the potential (if not likely) differences between results at these distinct levels of analysis. The first step is for investigators to explicitly determine if their research question involves intraindividual or interindividual processes (or both) and choose the appropriate study design. As group-to-individual generalizability is an unsupported assumption for most processes, it is critical to match inferential goals with study designs.
If primarily interested in an interindividual process (e.g., neural predictor of clinical diagnosis) a population-based study design that maximizes generalizability and power to detect small effects is most appropriate (Marek et al., 2022). Moreover, interindividual findings can then constrain hypotheses for investigators interested in additional intraindividual study. For example, with several interindividual findings of amygdala reactivity to threat predicting anxiety treatment response (Klumpp & Fitzgerald, 2018), investigators may be interested in if treatment affects amygdala reactivity, and how this in turn reduces anxiety symptoms. However, with orthogonal variance, this intraindividual study could produce parallel, opposite, or null effects relative to interindividual study (e.g., Lynch et al., 2024). Group-to-individual generalizability must be tested explicitly for any specific process, rather than assumed.
With the unlikeliness of group-to-individual generalizability assumptions, it is inefficient to begin with interindividual study when the primary goal is testing an intraindividual hypothesis. Instead, investigators should first turn to individualized, intensive longitudinal designs such as precision imaging. After completing an intraindividual study, we agree with Gell et al. (2024) that more population-based studies can then be useful to assess generalizability. However, we must be careful with what we specifically mean by assessing generalizability. That is*, if an intraindividual process was examined*, generalizability is not simply a test of replicability in a larger, more representative sample. Rather, generalizability tests may assess (1) group-to-individual generalizability or (2) the generalizability of the intraindividual processes (i.e., homogeneity assumption). Following an intraindividual study, population-based studies are appropriate for directly testing group-to-individual generalizability, or in this case, individual-to-group generalizability. On the other hand, resource constraints result in population-based studies rarely, if ever, having sufficient timepoints to test the homogeneity of intraindividual processes across individuals. Smaller-sampled cohort studies are likely better equipped to answer this question, as well the important follow-up in the likely case of identifying interindividual moderators of intraindividual processes (i.e., for who does this intraindividual process apply).
In sum, there is no “best” study design. Rather, study designs must be matched to the inferential goals of the investigator. Precision imaging offers an excellent framework for intraindividual study, and population-based studies are ideal for many interindividual studies. Moreover, purposeful combination of intraindividual and interindividual study provide complementary insights to an overarching research program, as well as explicit tests of group-to-individual generalizability. See Kraus et al. (2023) and Gell et al. (2024) for additional discussions of practical concerns with different study design frameworks, such as observation time intervals and resource tradeoffs.
For the past century, behavioral sciences have been grappling with the contrasting nomothetic and idiographic philosophies of science—a question of whether to focus resources on the study of the group for generalizable truths or on the intensive study of individuals to better characterize functioning across time and context. This debate culminated in Molenaar’s (2004) idiographic manifesto, which introduced the concept of ergodicity to the psychological sciences and concluded that between-person studies are unlikely to capture the nature of within-person processes that most theories are interested in. More recently, history is repeating itself in cognitive neuroscience, where two paths of population-based studies and individualized precision imaging studies are each being propelled as paths forward for distinct reasons. We connected reviews of idiographic science to robust findings of heterogeneity between individuals and non-stationarity within individuals for both patterns of brain function and brain-behavior associations. With these complex interindividual and intraindividual differences, we suggest that, like in psychology, group-to-individual generalizability is an unlikely assumption for much of cognitive neuroscience. We therefore conclude that cross-sectional, between-person studies are unlikely to provide appropriate insights into the within-person functioning that is critical to our theories.
To address this dilemma, we unite nomothetic and idiographic goals in a heuristic of complementary paths forward for advancing cognitive neuroscience research. We emphasize that neither a nomothetic, population-based approach or idiographic, precision approach is superior. Rather, study designs must be directly tied to the research question at hand. Our review highlights, however, that the dominance of cross-sectional study designs is inappropriate for testing within-person processes, resulting in a knowledge gap of individual-level functioning. To fully understand a construct, whether a basic science question such as neural functioning underlying a cognitive process or an applied question such as mechanisms of psychopathology, we must obtain inferences across both interindividual and intraindividual levels. Figure 4 provides a framework for researchers to select a study design that matches their theoretical question and then leverages the benefits of alternate study designs to complement and extend findings. Overall, our hope is that this review is useful for researchers in being more precise regarding the intraindividual or interindividual focus of their research question and matching study designs to their inferential goal.