Authors: Anirban Basu
Categories: Social and Interdisciplinary Sciences, Health and Medicine, Research Methods, Research Article, SciAdv r-articles
Source: Science Advances
To answer whether patients’ race belongs in clinical prediction algorithms, two types of prediction models are (i) diagnostic, which describes a patient’s clinical characteristics, and (ii) prognostic, which forecasts a clinical risk or treatment effect that a patient is likely to experience in the future. The ex ante equality of opportunity framework is used, where specific health outcomes, which are prediction targets, evolve dynamically due to the effects of legacy levels of outcomes, circumstances, and current individual efforts. In practical settings, this study shows that failure to include race corrections will propagate systemic inequities and discrimination in any diagnostic model and specific prognostic models that inform decisions by invoking an ex ante compensation principle. In contrast, including race in prognostic models that inform resource allocations following an ex ante reward principle can compromise the equality of opportunities for patients from different races. Simulation results demonstrate these arguments.
Health equity is viewed as a necessary aim for achieving high-value health care (1). It has also been pegged as a core aim of U.S. federal, local, and private health policy initiatives (2, 3). Despite these initiatives, systemic disparities, especially across racial and ethnic groups, continue to exist (3, 4). One concern regarding health equity is clinical algorithms’ role in exacerbating such disparities (5–7). Specifically, the President’s Big Data Working Group in 2014 (8) raised concern about “algorithmic discrimination” as an unintentional outcome of how big data technologies and structures are used that can potentially encode discrimination in automated or humanistic decisions. The recent empirical literature has shown that algorithmic discrimination is present in various sectors outside of health, where algorithms are used to guide decision-making. These areas include judicial practices (9–11), hiring and promotion decisions (12), and customs and border protection (13).
In the clinical and health economics literature, concerns about algorithmic discrimination are also evident (14, 15). Although many attributes of a healthcare application can lead to algorithmic biases, to what extent can the inclusion or exclusion of race or ethnicity variables as predictors or features in developing clinical algorithms induce algorithmic discrimination has not been rigorously addressed. Many clinical algorithms used in practice carry race (often as Black/non-Black classification) as one of the predictors. However, a growing voice within the clinical community signals that information on patients’ race should not be used as a feature within predictive algorithms employed in clinical practice (16). The general argument is that since race is not a biological feature and often the true underpinnings of how race can affect clinical outcomes are unknown, including race in such algorithms may perpetuate the inequities across patients of different racial groups. However, much of the technical literature on algorithmic bias does not answer whether race should be included or excluded in algorithms. Instead, it focuses on cases where the underlying data-generating processes could lead to discrimination even when the race variable is not directly included in the algorithms (14, 15). Recently, Manski (17) has argued that dropping race in such an algorithm may lead to suboptimal clinical decision-making outcomes, which could exacerbate disparities in the long run. A more recent working paper (18) “Race is not a particularly special Any covariate with predictive power (i.e., one that changes conditional probabilities of illness) should be used to optimize clinical decisions.”
A motivating example of this debate comes from the clinical algorithm literature on predicting glomerular filtration rate (GFR), which measures how well one’s kidneys function. A low GFR indicates poorer kidney health. To accurately measure GFR, one needs a series of blood and urine measurements. These measurements represent the gold standard for establishing the true GFR but are costly and time-consuming. Instead, researchers in the 1990s, and lately in 2009 (Chronic Kidney Disease Epidemiology Collaboration) (19), developed an estimated GFR (eGFR) based on an algorithm that was trained with gold-standard measurements and includes features such as age, biological sex, race (Black/non-Black), and serum creatinine. The variable Black carries a positive coefficient in this algorithm, implying that a Black patient would be assigned a higher eGFR than a non-Black patient of the same age, sex, and serum creatinine level.
Since eGFR is a central criterion used to allocate kidneys for transplantation, at the margin, some Black patients would be deemed to have healthier kidneys than observably identical non-Black patients and less likely to be assigned a kidney. This issue has led to national calls to remove race from these algorithms. In response, some institutions omit the Black race in the computation of eGFR, thus assigning the value for non-Black persons to Black persons. On the basis of these suggestions and practices, a recent task force on reassessing the inclusion of race in estimating GFR recommended that all equations be refit without the race variable, including or excluding additional biomarkers (20). In late 2021, researchers published work addressing these recommendations and found that excluding race in the original equations could lead to larger biases within each race group and larger differential biases between the groups, leading to systematic differences in care between race groups, especially at higher GFR. They also found that these algorithms performed better when an additional biomarker, cystatin C, was used for prediction. However, differential bias across races continued to be more when using predictions from algorithms that exclude race versus not. Although it remains to be seen which algorithm gets adopted in clinical practice, this example helps to set up a discussion to resolve the ambiguity of when race may or may not be included in such algorithms.
Here, this study asks whether race/ethnicity variables should be used in developing clinical prediction algorithms and, if so, under what conditions. This study unequivocally assumes that race is a social and not a biological construct. The vast majority of genetic variation associated with race-specific traits exists within racial groups and not between them (21). This distinction matters because certain arguments against using race in developing clinical prediction models invoke this notion of race not being a biological feature. Furthermore, this line of questioning has also invoked principles of individual-level inequality, i.e., when two individuals, each from a different race group but otherwise observationally identical, are offered a different clinical decision based on an algorithmic prediction (22). However, this study assumes that a central goal of algorithms is to promote population-level equity rather than individual-level equality. Under population-level equity, identical risk scores (predicted) across the population of different races represent the same underlying “true” health and, therefore, should be offered similar resources. Failure to do so is formally defined as a discrimination metric (23, 24) that this study adopts and discusses in detail. Under this premise, the arguments made in the clinical literature about dropping race as a predictor conflate various issues that may be useful to disentangle and answer the question in a nuanced fashion. These issues include (i) understanding the purpose of predictions and the fairness criterion invoked to address bias, (ii) disentangling the issue of using a mismeasured proxy for “true” health outcome in developing an algorithm from the use of race as a predictor in those algorithms, and (iii) understanding the role that systemic racism may play in inducing biological associations between biomarkers and health risks, even when no true biological associations exist. This paper studies these issues through a dynamic framework for a health production function and the lens of equality of opportunity (E.O.) to identify situations where race should or should not be included in algorithms.
A central theme that this study invokes for discussing the principles of discrimination is that of the E.O. (25, 26), a social ideal that combines concern with freedom and equality, and this social ideal provides a vision of how one ought to live together (27, 28). Specifically, as its name suggests, it has two principles, equality and opportunity. Agent(s) have an opportunity when they have a chance to attain a specified goal(s) without the hindrance of some obstacle(s). The E.O. framework has been broadly used in political science (29, 30), economics (25, 31), and law (32) to assess the equity footing built into specific policies. The U.S. Supreme Court has continuously relied on E.O. principles for several landmark rulings (33–36). This is the first time anyone will apply this framework to hold machine learning and other artificial intelligence algorithms to the same standard of equity and, in the process, answer whether race should be included in these algorithms.
Specifically, this study uses Roemer’s approach in suggesting that individual outcomes are a consequence of an individual’s circumstances, effort, and the resources/treatment that the individual receives (25). Efforts represent all possible actions individuals take, such as health care seeking, adherence to treatment, healthy behaviors, investments in education, hours spent in labor markets, etc. Preferences, free will, and perceived benefits and costs drive individual actions. Efforts are believed to be explicit decisions individuals make, often shaped by their circumstances but can also be through preferences and free will. Circumstances represent immutable conditions that individuals face that are beyond their direct control. They include biological circumstances (e.g., age and biological sex) and nonbiological circumstances (e.g., race, parental income, an environment with systemic racism). Outcomes represent any outcome for individuals (e.g., income, education, blood pressure, etc.). This study will focus on biological outcomes as the primary outcome of interest, including various biological features such as heart health, kidney health, and blood nutrients, and various clinical outcomes such as pain and symptoms. One can easily extend this framework to any other type of outcome.
The E.O. principle (25) invokes a normative framework that creates a level playing field across individuals, ensuring that allocating resources is such that each individual would have an equal opportunity to generate the same distribution of outcomes. Note that this framework differs from a purely utilitarian framework (18). The utilitarian framework aims to maximize an (positive) outcome in the population, irrespective of who can do so, suggesting allocating resources to those with the most opportunities to generate specific outcomes or utilities. One may view such a utilitarian approach as problematic under any notion of fairness. There is no reason why clinical decision-making should not include some notions of fairness beyond just a utilitarian framework. Under the E.O. rubric and following Roemer’s original approach, which focused on a treatment allocation decision, inequality of outcomes is unethical if it arises due to differences in circumstances. Such unethicality can be remedied using an ex ante approach that prescribes compensating individuals with disadvantaged circumstances, giving them the same footing/opportunity to generate outcomes. Similarly, inequality in outcomes arising from differential effort across individuals within each level of circumstances is not a moral bad (i.e., such inequality is acceptable) (25). That is, society should hold individuals responsible for their accountable effort (likely driven by preferences and free will) once the influence of circumstances on observed efforts has been removed (37). Another ex ante approach prescribes that two individuals with the same circumstance characteristics should be rewarded differentially to preserve differences in their expected outcomes (38). Both the concepts of compensation and reward also have an ex post version available (37, 38). However, this study only uses ex ante versions since, aligned with the goals of predictions, E.O. is evaluated during decision-making before the realization of efforts or the preferred outcomes become known.
The distinction between morally acceptable and unacceptable inequality highlighted by the E.O. framework is perhaps the most important contribution of egalitarian philosophical thought during the last 40 years. Much of the literature on E.O. relies on theoretical arguments and refines the content of this egalitarian normative framework (37). However, there are tremendous challenges in the empirical implementation of such a framework. The critical setback lies in the limitation of data that fails to capture the full range of so-called circumstances that can affect a preferred outcome. Such data limitations create uncertainty about how to invoke the compensation principle to achieve E.O. Much of the empirical work to date has looked at a limited set of circumstances, invoking principles of unconfoundness from unobserved confounders (or circumstances) (39). This limitation of data and its consequences for invoking the normative tenets of E.O. will help us answer the question of whether to include race in algorithms.
This study tries to answer the question of how an analyst can invoke principles of E.O. to decide whether race, a specific type of circumstance and a proxy for many other such circumstances, should be used as a predictor in clinical algorithms. The dynamics of outcomes as a function of circumstances and efforts. Such a dynamic model is conceptualized in Fig. 1, where the central theme is the development of a prediction algorithm that can inform the allocation of resources/treatment at time t. The key features of this framework
Fig. 1. The dynamics of circumstances and effort in the context of race and equality of opportunity.Circumstances either remain fixed or evolve outside the individual’s control. Circumstances can also directly affect outcomes in the next period. However, circumstances like race would only directly affect outcomes if there were differences in accountable efforts independent of all other circumstances. Outcomes are distinct from circumstances. They can be influenced by circumstances and through one’s own efforts independent of circumstances from last period. Past circumstances and outcomes can also influence past efforts that lead to current outcomes. For example, a health outcome in this period is directly affected by all health outcomes, resources received, and efforts exerted by individuals to improve health in the last period. However, circumstances such as systemic racism and other social determinants of health can prevent individuals from accessing health care, timely follow-up in care, or paying for care, affecting their biological outcome in the next period.
· Circumstances (biological, e.g., age and sex, and nonbiological, e.g., race, systemic racism, and parental characteristics) either remain fixed or evolve outside the individual’s control. Circumstances can also directly affect outcomes in the next period (e.g., parental relationship, a circumstance directly affecting the mental health outcomes of children). Circumstances like race would only directly affect outcomes if there are differences in accountable efforts (see below) independent of all other circumstances.
· Outcomes are distinct from circumstances. They can be influenced by circumstances and also through one’s own efforts independent of circumstances. Outcomes at time t have legacy effects from the same outcomes in time t − 1. They are also affected by efforts incurred by individuals and directed policies/resources/treatments in t − 1. Past circumstances and outcomes, in turn, can influence past efforts that lead to current outcomes. For example, an individual’s income this year is a direct function of the income last year. Also, today’s income is affected by last year’s efforts to invest in human capital or job-seeking behavior and the availability of job training services (i.e., resource allocations). Specific neighborhood characteristics (circumstances) and poor health (past outcome) impede investments in human capital or other efforts. Similarly, a health outcome in this period is directly affected by all health outcomes, resources received, and efforts exerted by individuals to improve health in the last period. However, circumstances such as systemic racism and other social determinants of health can prevent individuals from accessing health care, timely follow-up in care, or paying for care, affecting their biological outcome in the next period.
Typically, in clinical decision-making, one biological outcome is identified as a fair allocation score to allocate resources/treatment to all individuals. For example, determining whether one would benefit from a kidney transplant would be straightforward if one truly knew an individual’s kidney function/filtration rate level. Following such allocation would be fair for all individuals. However, a clinician may not readily observe such biological outcomes, especially those that occur in the future, and rely on predictions from an algorithm for decision-making. Two scenarios arise.
In the first scenario, the prediction target is a biological outcome in t. These include the presence of gallstones among patients with flank pain, urinary tract infections, detection of breast cancer using radiologic images, and the quality of kidney filtration. Here, algorithms are used to predict the current level of a biological outcome (already realized at time t, in Fig. 1, relative to the decision-making time) that remains unobserved to the decision-maker. These algorithms are often classified as diagnostic algorithms (40). These predictions guide the resource allocation decision, which aims to reduce inequalities of a future outcome (e.g., life expectancy) across circumstances. A classic example of scenario #1 is when the average treatment effects in an randomized clinical trials (RCT) are heterogeneous across groups defined by a baseline biological factor. Sometimes measuring this baseline factor remains difficult and costly in clinical practice. In such a case, a predicted score of this factor is used to allocate treatment, hoping that such allocation would reduce rather than enhance inequalities across patients with different circumstances in a preferred outcome in the future.
To formalize these concepts, this study follows Davillas and Jones (2020) (39) to write a general health production function for a target outcome (Bj,t) at t as a function of circumstances (Ct−1), efforts (Et−1), and resources offered (St−1) at time t − 1. Efforts in t − 1 are directly influenced by concurrent circumstances, both biological and nonbiological (Fig. 1). Boldface notation is used to reflect a vector of variables
where v and u represent the unobserved random variation in Bt
(41, 42). Here, the vector of v represents random variation in efforts that are independent of B~t~ and Ct, and u represents random variation in the Bt that is independent of Bt−*1, and E, Ct−*1*t−*1.
In scenario #1, suppose the decision-maker wants to allocate resources St based on a health outcome j at t, Bj,t, so that inequality in a future outcome k (k =,≠ j) at t + 1*, B~k,t*+1,~ across circumstances is reduced. Let there already be reduced-form evidence (i.e., circumstances and efforts affecting Bk,t+1 are included in ut+1) that links Bj,t to Bk,t+1. This is also the decision-making model at t
When Bj,t remains unobserved to the decision-maker at the point of decision-making at t, it must rely on a predicted value of this health outcome from an algorithm developed by an analyst. For example, consider deciding whether a patient should get a kidney transplant. Say there is already established data saying that kidney transplant at time t will improve future survival at t + 1 if GFR(t) < K (i.e., some threshold value). This is Eq. 2. This information is already available to the decision-maker, perhaps through past RCTs. However, at the point of decision-making at t, the decision-maker does not observe GFR(t), and so they rely on an algorithm to predict eGRF at t. The structural model in Eq. 1 is the basis for this algorithm.
To develop this algorithm, suppose the analyst observes data on Bj,t along with all possible circumstances, past efforts, and past resources used. In such cases, let the algorithmic prediction of B~j,t ~be denoted by B^j,t. Note that B^j,t contains no information about efforts that will be incurred at t, since Et does not belong in the structural Eq. 1. So when B^j,t is plugged into Eq. 2 to inform the decision of St, it will appropriately capture the variation in circumstances that will influence the production of the future outcome Bk,t+1 through Bj,t.
Therefore, using B^j,t in the decision-making model of Eq. 2 can invoke the ex ante compensation principle where more resources are provided to certain groups (i.e., certain racial groups or other circumstances) with compromised B^j,t, such that people with different circumstances have the equal opportunity to produce a similar distribution of Bk,t+1.
In practice, the analyst trying to develop a prediction algorithm for Bj,t does not observe most past efforts or resource uses and many circumstances. Instead, the analysts typically use a vector of other outcomes at t, B−j,t, as predictors. Following Eq. 1, this vector of outcomes serves as an (imperfect) proxy for all circumstances, past efforts, and resources. Prediction quality can be substantially enhanced if additional information on circumstances is directly used in the model. For example, let there be at least one circumstance besides race that remains unobserved. As long as the unobserved circumstance is correlated with the race indicator, the estimated coefficient on race in the algorithm should reflect the contribution of unobserved circumstances on Bj,t. Consequently, variations in B^j,t would better reflect variations across different circumstances if variables like race are included in the prediction model to better allow for applying the E.O. ex ante compensation principle.
In the second scenario, the prediction target is a biological outcome in the future (t + 1). These algorithms are classified as prognostic algorithms (40). These include the risk of heart failure in a year, inpatient mortality in patients with acute heart failure, long-term complications with certain surgeries, and cancer survival in general. These algorithms are used to predict future levels of a biological outcome (time t + 1 in Fig. 1) relative to the time of decision-making that the algorithms attempt to inform. Like scenario #1, these predictions guide the resource allocation decision at time t. However, unlike scenario #1, one of the goals is to invoke E.O. principles with respect to the predicted outcome from the algorithm itself. Under scenario #2, two cases In the first case (scenario #2a), the goal is the same as in scenario #1. The decision informed by the predictions is to ask whether to offer differential compensation to individuals with different circumstances. In the second case (scenario #2b), the decision is about rewarding individuals with expected high levels of the preferred outcome, controlling for circumstances. The preface for such decisions under scenario #2b is to reward accountable efforts (in line with the ex-ante reward approach) or to reward returns to certain biomarkers levels at baseline due to biological rationale, again controlling for circumstances.
An example of scenario #2 is developing a prediction algorithm for future survival with a kidney transplant. If such predictions are used to inform and overcome the burden of disease, i.e., allocate resources toward those who are predicted to have low survival in the future, it is scenario #2a. If the predictions are used to reward individuals who benefit more from certain resources due to independent contributions of their biology or accountable efforts, then it falls under scenario #2b.
For the two scenarios described above, following Davillas and Jones (2020) (39), a general health production function can be written as
So, Eq. 3 is the structural underpinning of Eq. 2. A linear reduced form for the production function can be represented as
Here, the coefficients α1and α2 represent the total effects of the legacy outcomes and circumstances at t, respectively. Each total effect consists of the direct (main) effect of the covariate and the indirect effect of that covariate through efforts that are correlated with it.
Before considering an algorithm that predicts Bj,t+1, let us consider a scenario where a decision-maker can have Oracle vision and observe E(Bj,t+1) (or causal average effects of St on Bj,t+1) at each level of combinations of Ct and Bt, at t. Could she institute an E.O. principle based on such an Oracle vision? Unlike scenario #1, variations in E(Bj,t+1) across Bt or Ct would include the influences of efforts (accountable and circumstance-based) at t. This creates challenges for precisely invoking the ex ante E.O. compensation principle across non-race circumstances (e.g., income levels) or baseline/legacy outcomes. Decision-makers will face ambiguity about how much of the variation in outcomes is driven by accountable efforts versus circumstance-driven efforts.
Nevertheless, one can invoke an additional egalitarian principle saying that circumstance-driven efforts represent the major share of all efforts in generating health outcomes. For example, consider efforts on drinking habits that would affect future kidney health and survival. One may argue that independent of all circumstances, intrinsic preferences for drinking (representing accountable efforts) would vary little across different groups of people. Therefore, an exhaustive list of circumstances can explain the substantial observed variation in drinking habits. Under such an assumption, variation in E(Bj,t+1) across any combination of non-race circumstances and baseline biological factors and outcomes are primarily independent of accountable efforts and can be used to implement the ex ante E.O. compensation principle (i.e., scenario #2a). For example, decision-makers can use the variation in the future survival of individuals with kidney disease across combinations of baseline biological statuses and non-race circumstances to allocate resources (e.g., transplants) toward those expected to have low survival. Since race is assumed to have no direct effect on outcomes except through efforts and other circumstances, when all other non-race circumstances levels and baseline outcomes levels are held constant, the conditional variance of E(Bj,t+1) across levels of race variable would only reflect the impact of accountable efforts. Therefore, such a variation of E(Bj,t+1) across levels of the race variable should not be used to implement an ex ante E.O. compensation principle.
An example of scenario #2b is when one tries to develop an algorithm predicting treatment effect heterogeneity within a randomized clinical trial. One often tries to estimate conditional average treatment effects to identify levels of baseline outcomes where returns to treatment are highest and reward those individuals with treatment. Lower returns to the treatment could directly result from baseline outcomes and circumstances that individuals face (e.g., exposure to parental smoking) that may directly interfere with some biological mechanisms generating outcomes. Lower returns could also be shaped by future efforts such as medication adherence, healthy lifestyle, etc., some of which may, in turn, be shaped by circumstances such as systemic racism and affordability. If oracle-level information is also available for the average causal effects of St on Bj,t+1, i.e., ∂Bj,t+1/St, at each level of combinations of Ct and Bt, one can also institute an ex ante E.O. reward principle based on average causal effects across baseline outcomes. For example, a rationale to reward individuals with GFR < K with transplants because the effect of transplant on survival is larger than those with GFR ≥ K, would meet the ex ante E.O. reward principle only when all other circumstances are held constant. Then, differences in transplant effects in the two groups must be driven by the causal biological mechanisms or accountable efforts. Similarly, one can also invoke the ex ante E.O. reward principle based on the race variable. A decision-maker can reward individuals from race A with kidney transplant at a higher rate over those from race B if survival gains with a transplant are higher among the former, as long all other potential circumstances and baseline factors have been controlled. A decision-maker, however, cannot invoke the ex ante E.O. reward principle for any non-race circumstances. The differential returns across that nonrace circumstance could be generated by that specific nonrace circumstance-driven differential efforts, even when all other circumstances are controlled for.
Suppose the analyst intends to predict Bj,t+1, observes data on it along with baseline outcomes and all possible circumstances at t, and uses them for prediction. In such cases, let the algorithmic prediction of Bj,t+1 be denoted by B^j,t+1. Note that the properties of B^j,t+1 would reflect those described above under the Oracle vision for Bj,t+1. Under such an ideal algorithm development situation, algorithmic predictions can be used to implement an ex ante E.O.–based compensatory principle for all non-race and baseline outcomes levels, but not race. This implies that for scenario #2a type decisions, race need not be included in algorithm development as long as all other circumstances are included.
In contrast, algorithmic predictions for returns to St under such an ideal situation can be used to design ex ante E.O. reward principle for any baseline outcomes and race but not any other non-race circumstances. However, all circumstances, race or non-race, should belong in the prediction model.
In the more practical empirical scenario, suppose the analyst does not observe all circumstances in the data. Here, estimates of α2 for an observed circumstance would capture not only the main effects of that circumstance and the indirect effects of efforts due to that circumstance but also the direct and indirect effects of unobserved circumstances correlated with the observed circumstance. For example, the coefficient on race would capture the effects of the differences in accountable efforts and the direct and indirect effects of unobserved circumstances correlated with race. Moreover, estimates of α1 on a legacy outcome would also pick up some influences due to unobserved circumstances that are correlated with the legacy outcome. Similar to the ideal situation above, decision-makers cannot cleanly enforce an ex-ante E.O. compensation principle based on predictions B^j,t+1, since some of this variation would remain to be driven by accountable efforts. However, under the assumption that the variations in circumstances-based efforts represent a larger influence on future outcomes compared to variations in accountable efforts, the major share in the variation of B^j,t+1 would still reflect variations due to observed and unobserved circumstances. Therefore, a decision-maker can invoke a second-best ex ante compensation principle under scenario #2a. However, in contrast to the ideal situation above, race should belong in prediction algorithms for scenario #2a, since it can help capture the variations in the preferred outcome due to unobserved circumstances to better implement an ex ante E.O. compensation principle.
For example, an algorithm that predicts the future survival of individuals with kidney disease as a function of baseline biological status and race can still be used to allocate resources (e.g., transplants) toward those expected to have low survival. In this case, the allocation of resources could be unequal across races (i.e., invoking the compensation principle), reflecting the unequal circumstances (observed or unobserved) faced by these groups so that each, irrespective of their own’s race, has an equal opportunity for survival.
For scenario #2b, using ∂B^j,t+1/St, generated with partially observed circumstances, to enforce E.O. is tricky. Unlike the ideal situation, the estimated returns of St across legacy outcomes or race will be contaminated with the influence of correlated unobserved circumstances and associated efforts. There is no perfect way to implement an ex ante E.O. reward principle in this situation. However, let us consider a situation where the variation in Bj,t+1 due to any individual efforts (accountable or circumstance based) can be removed (i.e., kept constant at some fixed level e¯), then the resulting hypothetical outcome Bj,t+1∗could represent a fairer allocation score, based on which an acceptable reward principle could be instituted.
Even though not perfect, as the confounding effects of unobserved circumstances would still be embedded in these hypothetical predictions, variations in ∂B^j,t+1∗/St across levels of the baseline/legacy outcomes would have an attenuated influence of circumstances-driven effort channels compared to variations in ∂B^j,t+1/St. This rationale implies that although the algorithms in practice would be trained on Bj,t+1, one should be open to accepting biased predictions as long as B^j,t+1 aligns with hypothetical Bj,t+1∗. One should try algorithmic development strategies that minimize the influences of efforts in future outcome production for these predictions to instill ex ante reward–based E.O. on legacy outcomes. Whether inclusion or exclusion of observed circumstances such as race achieves such fair predictions remains to be an open question. The inclusion of race can likely minimize the distance between B^j,t+1 and Bj,t+1, but it could increase the distance between B^j,t+1 and Bj,t+1∗ and, consequently, is discouraged in such situations.
Consider when an algorithm is developed to represent the treatment-effect heterogeneity of kidney transplants on survival using data from a randomized controlled trial. Although randomization may help equate the distribution of circumstances across treatment groups, it does not equate the distribution of circumstances across race or the levels of baseline outcomes. In most situations, one cannot control all possible circumstances for developing this algorithm. Consequently, using a predicting algorithm that includes the race variable but no other circumstances to reward individuals by allocating treatment differentially across races or, say, baseline levels of GFR would be discriminatory. The endogeneity of the race variable will create bias in its own coefficients and the coefficients on other baseline outcome variables so that the variation in predicted effects of transplant would carry the influences of the unobserved circumstances-driven efforts.
Under practical situations with unobserved circumstances, invoking the E.O. approach in scenarios #1 and #2a corresponds to a utilitarian approach to allocating resources where race should be considered in developing algorithms and whose predictions are intended to implement an ex ante E.O. compensation principle. In contrast, the E.O. approach deviates from a pure utilitarian framework in scenario #2b because the allocation decision involves a reward principle. Including race in algorithms meant to inform such purpose could be problematic.
This study considers dynamic biological outcomes over three time periods (t = 0, 1, 2) in a population of individuals that consist of two race groups, R Є {a, b}, representing one of the durable circumstances. Race a (R = 1) consists of about 40% of the population. Starting conditions are t = 0 with an initial distribution of two correlated biological outcomes that are independent of race.
Over time, the evolution of these outcomes occurs as a function of legacy biological outcomes and efforts (E), both of which positively influence the next period’s biological outcomes.
B1t and B2t, t = 1, 2, were each standardized to have a mean of zero and a SD of one.
To keep it simple, only the influence of circumstances on outcomes through past efforts was allowed. For example, efforts at any time may depend on current biological outcomes, many social determinants of health, and systemic racism. Here, those circumstances, which are different by race, are not made explicit. However, these factors, and hence the race indicator representing the underserved population, negatively influence the efforts exerted. In addition, this study allows for the interaction of race with biological outcomes in exerting effort. The negative influence of the social determinants gets amplified at lower (negative) values of biological outcomes.
This study considers two distinct linear prediction models that aim to inform a resource allocation decision at time t = 1. The first one, corresponding to scenario #1, aims to predict B21. The other, corresponding to scenario #2, aims to predict B22. Specifically, for each of these prediction models, a model with and without race in it is explored
In models with race, the race variable was entered with and without an interaction with the other predictor.
The base case simulations will develop the algorithms using {B11, B21, B22, R} data. This study directly observes the true levels of biological outcomes while developing the algorithms. Efforts always remain unobserved. Predictions will be based on {B11, R}.
Here, B~11 is not available while developing algorithms but rather, a proxy B11 is available, which is a noisy measure of B11 and noise level is heteroscedastic over race.
An example of such a situation is patient weight. Survey data on which algorithms are developed contain the self-reported weight, a noisy proxy for true weight available in clinic offices that are used for predictions. In this case, simulations will develop the algorithms using {B~11,B21,B22,R} data. However, predictions will be based on {B11, R}.
The third case considers the situation where the algorithm is developed on proxy outcomes that are noisy measures of B21 and B22.The noise levels are heteroscedastic over race.
This is a scenario highlighted by Obermeyer et al. (14), where the authors showed that a widely used algorithm developed using healthcare cost as a proxy for health exhibited significant (statistical and substantive) racial bias. This gap between the proxy B~2j used to develop an algorithm and the actual target B2j that one wishes to inform is referred to as the label choice bias and is responsible for algorithmic biases in many clinical applications (43, 44).
This last case considers the situation where for each of the algorithms developed under each of cases 1 to 3 that have the race variable in it, predictions used for decision-making were made using a misclassified race variable. A classic example of this situation is when the algorithms are developed using self-reported race (considered to be the gold standard here; see Discussion) but deployed in clinical practice where clinicians rely on observed race to make predictions. In this case, misclassified race (R~) is given as R~ = 0 in random 15% where R = 1. See Discussion for the implications of the simulation results in the context of the empirical evidence on such misclassification.
This study starts with a recently proposed fairness metric in the economics literature by Arnold et al. (23) and Rambachan et al. (24). They consider an allocation problem for each individual, where “qualification” to receive a specific level of a resource Si for individual i is based on some unobserved (latent) variable Yi^^, denoted as the “qualifying variable.” In the context of this study, one can think of this qualifying variable as the “true” biological outcomes of an individual, and treatment is efficacious below a certain threshold of health. Without loss of generality, let Yi be absolutely continuous with respect to the Lebesgue measure and indicate the “true” qualification for the allocation level of that resource for all individuals.
In an ideal world, the absence of discrimination implies that one can allocate resources fairly to those who qualify for them, and this allocation rule is the same in each group. Suppose the allocation rule is, for some reason, different between the two groups. In that case, the difference in the mean allocation between the two groups, conditional on the latent variable Y^*^, should be viewed as the extent of “discrimination” between these groups. Thus, total discrimination between the two groups is given by (23)
When the E.O. lens is put onto this metric, it becomes clear that future efforts should not influence any such qualification variable. This is because a qualification variable that incorporates future individual efforts also incorporates all the systemic inequalities that will generate those efforts. Therefore, equal allocation of resources at the current time based on such a qualification variable will still be discriminatory since one group may have a more difficult path to reach certain levels of the qualification variable. Consequently, for scenario #2b, even though B22 is being predicted, which will be realized in the future, the appropriate qualification variable is given as
B22∗ is standardized to have a mean zero and a SD of one. Compared to B22, the appropriate qualification score does not include the current efforts that would be exerted after resource allocation is made at time t = 1.
On the contrary, for scenario #1, where the target of predictions is B21, B21 can directly serve as the qualification score, B21∗(=B21). This score is already realized and would not depend on any additional efforts that are exerted after resource allocation has been made. However, it does include past efforts, which would be reflected in current circumstances.
Scenario #1 is identical to scenario #2a, where the prediction target is B22, but the resource allocation decision is preventive and, therefore, should allow past (in this case, at time t) efforts to shape the fair allocation score. Here, scenario #2b is not studied separately, as the simulation results from scenario #1 should directly apply to scenario #2b.
Last, to evaluate algorithm performance, it is important to focus on algorithmic discrimination by separating aspects of human discrimination from Eq. 11. An algorithm produces predictions B^21 and B^22. Resource allocation decisions are based on these predictions. While discrimination can be represented as the difference in the allocation of treatment or resources for two groups of individuals with the same predicted score from an algorithm, such difference may not entirely represent algorithmic discrimination. A decision-maker may still allocate treatment or resources differentially across certain groups even when directly observing the qualification variable levels (and not through a prediction model). For example, researchers had pointed out that the allocation of kidneys for transplant was differential across races well before any algorithms were used to predict kidney health. These differences represent aspects of human discrimination that arise from prejudices, preferences, and statistical discrimination in decision-making (45, 46). Instead, an appropriate measure of algorithmic discrimination should look at each possible predicted risk score and then assess whether the (expected) qualification score among individuals assigned that risk score is the same for each relevant group. One can average these deviations over the same marginal distribution of risk scores for all relevant groups to determine a population-level discrimination metric. Such a metric is equivalent to “calibration targets” in the computer science literature (10, 47). Therefore, for the purposes of this study, the main criterion for evaluating algorithmic fairness is given by the discrimination metric
There is a growing recognition that the discrimination metric is the superior metric to consider when assessing the fairness property of an algorithm (44, 48). While the discrimination metric is the most natural metric to evaluate algorithmic fairness, other metrics are also available. A closely related metric, which was used to assess the eGFR predictions, is differential bias. For any group, bias is calculated by averaging the deviations (as in discrimination) over the group-specific marginal distribution of risk scores.
In essence, it is the mean residuals for each group. Therefore, differential bias is the difference in mean residuals between two groups. The disadvantage of differential bias over discrimination is that the former can be nonzero even when there is no discrimination due to different distributions of risk scores in the two groups.
In addition, a variety of other fairness criteria have been used to evaluate algorithms. They
· Balance of positive or negative cases compares between groups the mean residuals among individuals whose true heath is above or below a certain threshold.
· Differential false positive compares the proportion of patients who do not have a condition (defined by being above a certain threshold of the qualification variable) given that they were identified to have a condition using the predictions (again, defined by being below a threshold of prediction scores).
· Differential false negative compares the proportion of patients who have a condition given that they were identified not to have a condition using the predictions.
· Differential positive predictive value (PPV) compares the proportion of patients who are predicted to have a condition among those who do have the condition.
· Differential negative predictive value (NPV) compares the proportion of patients who are predicted not to have a condition among those who do not have the condition.
It is important to realize that one algorithm can never be the best in all these metrics simultaneously. Classic examples of this issue are the balance of positive and negative classes. An algorithm that produces zero differential bias would almost, by definition, not generate zero balance in either positive or negative class. Similarly, an algorithm with a low differential PPV will likely not have a low differential NPV. Therefore, it is important to have a clear rationale for why any of these alternate metrics would be chosen to assess fairness in an algorithm. In addition to the primary discrimination metric, this study explores algorithm performance based on each alternative metric in the context of the simulations.
Even though at t = 0, the distribution of biological outcomes B10 and B20 was similar across race groups, these distributions begin to separate over time, with dynamically differential efforts (Fig. 2). Individuals belonging to race a (R = 1), facing a higher barrier to exerting efforts, continue to lag in biological outcomes production. Along with such dynamics, the association between biological outcomes changes too. Figure 3 illustrates the association between B21 and B11 (left) and B22 and B11 (right). A differential association between these variables by race emerges. But does this feature indicate whether the race variable should be used in developing the prediction models? This study explores the answers to this question through simulations. Simulation results from cases 1, 2, and 3 are provided in Tables 1, 2, and 3, respectively.
Fig. 2. The evolution of the distributions of biological outcomes over time in two race groups.Even though at t = 0, the distribution of biological outcomes B
10and B20were similar across race groups, these distributions begin to separate over time, with dynamically differential efforts. Individuals belonging to race a (R = 1), facing a higher barrier to exerting efforts, continue to lag in biological outcomes production.
Fig. 3. Race-specific associations of biological circumstance B
1at t = 1 as predictor with biological outcomes B2at t = 1 or t = 2 as outcomes.Along with such dynamics of differential outcome production in the two groups, the association between biological outcomes changes too. A differential association between these variables by race emerges.
With case 1 data (i.e., with no measurement errors), this study finds that when predicting B21 (corresponding to scenario #1 and scenario #2a decision-making), including race in the prediction algorithm generates substantially lower discrimination than the algorithm without race (Table 1, top). Including race in algorithms generates better calibration in both the race groups (race a: 0.00 and race b: 0.00 when interaction with race is included; race a: 0.09 and race b: 0.00 when only race main effect is included) compared to when race was not included (race a: 0.23 and race b: −0.20). These results carry over to the differential bias criterion, where algorithms, including race, perform far better. Algorithms that achieve overall lower bias for each race group and lower differential bias overall are not expected to maintain these results conditional on outcomes value below or above a certain threshold. The results in the +ve balance and −ve balance rows confirm these expectations. Therefore, the positive and negative balance measures, which are differences in such conditional biases between race groups, can go in arbitrary directions.
In terms of classification, algorithms with race perform much better in differential false positives or false negatives than algorithms without race. Differential PPV and NPV are higher with algorithms with race versus not. However, PPV was 89% for race a and 47% for race b with algorithms including a race interaction. The corresponding numbers for the algorithm without race were 81 and 68%, respectively, suggesting that the differential PPV alone does not tell the whole story. Differences can mask better performance in one group or the other.
When misclassified race variable is considered in predictions (corresponding case 4), all fairness metric deteriorates for algorithms with race, but they still perform substantially better than algorithm without race. Substantial misclassification of race variables would be required to erode the advantages of including race in these algorithms. It is important to note that under these data-generating processes, mismeasured race in algorithm development produces a positive bias in the calibration metric for race a (but does not affect calibration in race b).
The story changes in outcome B22 (corresponding to scenario #2b). Table 1 (bottom) shows that predictions from algorithms with race perform worse when compared with targets that are stripped of future efforts. Differential discrimination was higher for algorithms with race (−0.42 with race interaction and −0.50 with race main effects) compared to the algorithm without race (0.38). Differential bias tells the same story. As with above, the differential classification does not tell a coherent story about which algorithm is better as it illustrates fundamental trade-offs between prediction accuracy between the two groups.
When the algorithms were trained on a noisy measure of race (corresponding case 4), the bias generated appears to be protective in this case, improving the calibration within race a (but not affecting calibration for race b), generating lower differential discrimination. This is likely a manifestation of the positive bias that the mismeasured race variable generates in race a, which, when combined with the negative bias in calibration within race a when the proper race variable is used, produces a lower calibration bias. The main takeaway from these simulations is that mismeasured race will generate bias in calibration, especially in the mismeasured race group. There is no reason to believe that the direction of differential bias reported here will be generic with mismeasured race.
When developing algorithms with case 2 data (i.e., with measurement error in predictor B11), this study finds that when predicting B21, all algorithms generate differential discrimination, the lowest being with the algorithm with race interaction (Table 2, top). Differential bias shows a clear superiority for algorithms with race compared to those without race. However, as discussed earlier, this metric may not be ideal for detecting discrimination. Classification metrics show the same ambiguity as before in answering the question of whether to include race in algorithms. The misclassified race variable used in predictions shows the same additional positive bias in calibration with race a (as in Table 1), thereby reducing differential discrimination in these simulations by bringing the calibration bias in race a close to that in race b, where the calibration bias was positive.
When looking at outcome B22 (under scenario #2b), algorithms with race clearly perform much worse than those without race in terms of discrimination and differential bias (Table 2, bottom).
When developing algorithms with case 3 data (i.e., with measurement error in outcomes B2j
j = 1, 2), This study finds that when predicting B21, all algorithms generate high discrimination, with the algorithm without race producing the highest (Table 3, top). All other results are qualitatively similar, as in Table 2. When mismeasured race is used, the positive bias it imparts on calibration in race a increases discrimination since the calibration bias among race b was negative. Last, with predicting B22 for scenario #2b*,* all algorithms produce even higher discrimination, but the algorithms with race produce much higher discrimination than those without it (Table 3, bottom).
The clinical community has raised concerns about including race as a variable in clinical algorithms whose predictions are used for decision-making in clinical practices. Despite some valid concerns, there has not been enough specificity in the arguments as to how and when these algorithmic biases can creep in. This study approaches these issues using the lens of E.O. In this framework, the dynamic health production functions for outcomes, which are the prediction targets, include the legacy effects of biological and nonbiological outcomes and current individual efforts. Several social determinants and systemic racism shape these efforts. When invoking E.O., inequality of outcomes arising from (future) individual efforts should not be driving current prescriptive resource allocation at the individual level. This is because the opportunity to improve outcomes from baseline should be equally available to everyone, irrespective of whether anyone seizes those opportunities to produce better outcomes. This is a fundamental deviation from a utilitarian framework, which fails to incorporate any notion of fairness along the lines of opportunities.
On the basis of this rationale, prediction models can be broadly categorized into two types of prediction (i) diagnostic, which describes a patient’s clinical characteristics, and (ii) prognostic, which forecasts a clinical risk or treatment effect that a patient is likely to experience in the future. Using theoretical arguments and simulations, this study shows that in practical settings, failure to include race corrections will propagate systemic inequities and discrimination in any diagnostic model and specific prognostic models that inform decisions by invoking an ex ante compensation principle. In contrast, including race in prognostic models that inform resource allocations following an ex ante reward principle can compromise the equality of opportunities for patients from different races. In such settings, race is likely to proxy for differential efforts across these groups, which unobserved differential circumstances can shape. Even when race is not included, discrimination is likely to be part of any algorithms predicting future outcomes used to devise an ex-ante reward policy, although including race exacerbates this problem.
One concern about including race in algorithms is the risk of poor implementation of the algorithm in clinical practices due to the potential for mismeasurement of race. Such mismeasurement may come when the algorithms were developed using self-reported race while the observed race in the clinic is used for implementation. Looking at the literature that has directly studied agreement between self-reported and observed (proxy-reported) race (49, 50), this study finds that such misclassification is extremely low for non-Hispanic Black and whites. Mixed-race and Latinos have larger misclassification but are seldom identified as separate categories in algorithms. A non-Hispanic Black/non-Black classification, frequently used, lumps all other races with the white categories, increasing discrimination for non-Black minority groups that also face systemic racism and access issues. This raises interesting questions about how race should be coded as a variable. Research should attempt to identify sample sizes large enough to identify as many groups as possible to develop these algorithms.
Measurement errors pose problems with algorithmic discrimination. All algorithms continue to produce discrimination when features/predictors have measurement errors. However, classification errors increase for both groups and differential classification, especially for algorithms with race. Outcome measurement errors generate considerable discrimination and poor classification properties for all algorithms. This result aligns with the literature on label choice bias, which recommends against using proxies of desired outcomes for algorithm development.
The underlying dynamics of circumstances and efforts of individuals drive the associations observed in the empirical approach. While internalizing these contexts is critical to improving decision-making and reducing algorithmic discrimination contemporaneously, these underlying factors can change over time, making observed associations of today obsolete tomorrow. Therefore, it is essential to note that clinical algorithms that include race should be regularly updated. Hence, researchers should not wait for decades before updating these algorithms but rather do them with sufficient regularity to keep up with the changing socioeconomic dimensions of society and the health care system.
There are several limitations to this study. Algorithmic biases due to the inclusion of race were not studied, where one of the race groups is sparsely represented in the analytic data. The biases resulting in such cases would be driven by both generalizability and inferential issues. Last, human discrimination is not factored into this study.
In summary, for algorithms meant to predict any current outcome (already realized) for any decision or to predict a future outcome meant to inform compensatory decisions, the question of inclusion of race is empirical. In these cases, failure to consider race as a feature, even with race having no biological rationale, would propagate and increase discrimination in clinical practice. In contrast, algorithms predicting a future outcome that is meant to inform a reward-oriented decision should avoid using race as a feature. Otherwise, it will allocate treatments differentially based on all future efforts, including those driven by unobserved circumstances, that different groups exert to produce future outcomes, thereby denying the E.O. to avail of these treatments across the groups. Although such algorithms would generate discrimination even when race is not included, including race is likely to increase discrimination.
This study did not require an institutional ethics review board determination as no human subjects were involved in any parts of the analysis.
This study simulates a population of 1 million people with the specification for data generation for the population starting with t = {B10, B20, R} following Eq. 5. Then, {B1j, B2j, Ej} are dynamically generated for j = 1, 2 following Eqs. 6 and 7. Last, the proxy measures {B11,,B21,B22} are generated using Eqs. 9 and 10. Similarly, a mismeasured race variable, R, is generated. Next, an iterative process is followed.
Twenty thousand patients are sampled from the target population using a simple random sample.
Three different algorithms are estimated using these (i) linear regression with main effects and interactions between B11 (or B11) and R, (ii) linear regression with only the main effects of B11) and R, and (iii) linear regression with only the main effects of B11 (or B11 (or B~11).
Predictions are generated based on B11 and R (or R~) for the entire target population.
The discrimination metric and each of the alternative fairness criterion metrics are evaluated.
One to four are repeated 500 times and averages of all the fairness metrics are reported.
To evaluate the classification criteria, a threshold of c = 0 was used for either outcome, with outcome values less than 0 representing the presence of a health condition.
I thank C. Manski, A. Bansal, S. Khor, A. Jones, and A. Salas Ortiz for the very helpful comments. I also thank seminar participants at the Tulane University, the University of Washington, and the European Workshop in Health Economics and Econometrics.
Funding: I acknowledge support from a consortium of 10 biomedical companies to the University of Washington through an unrestricted gift.
**Author ** A.B. takes sole responsibility for all parts of this study.
**Competing ** The author declares that they have no competing interests.
**Data and materials ** All data needed to evaluate the conclusions in the paper are present in the paper and/or the Supplementary Materials.
View/request a protocol for this paper from Bio-protocol.