Authors: Daniel R. Kowal
Categories: Article, Discrete data, Interactions, Penalized Estimation, Regression analysis
Source: Journal of the American Statistical Association
Authors: Daniel R. Kowal
Categorical covariates such as race, sex, or group are ubiquitous in regression analysis. While main-only (or ANCOVA) linear models are predominant, linear models that include categorical-continuous or categorical-categorical interactions are increasingly important and allow heterogeneous, group-specific effects. However, with standard approaches, the addition of categorical interactions fundamentally alters the estimates and interpretations of the main effects, often inflates their standard errors, and introduces significant concerns about group (e.g., racial) biases. We advocate an alternative parameterization and estimation scheme using abundance-based constraints (ABCs). ABCs induce a model parameterization that is both interpretable and equitable. Crucially, we show that with ABCs, the addition of categorical interactions (a) leaves main effect estimates unchanged and (b) enhances their statistical power, under reasonable conditions. Thus, analysts can, and arguably should include categorical interactions in linear models to discover potential heterogeneous effects—without compromising estimation, inference, and interpretability for the main effects. Using simulated data, we verify these invariance properties for estimation and inference and showcase the capabilities of ABCs to increase statistical power. We apply these tools to study demographic heterogeneities among the effects of social and environmental factors on STEM educational outcomes for children in North Carolina. An R package lmabc is available. Supplementary materials for this article are available online, including a standardized description of the materials available for reproducing the work.
Interactions are remarkably valuable in linear regression analysis. In particular, interactions between a categorical (or nominal) variable and a continuous or categorical variable—which we refer to as cat-modifiers, short for “categorical effect modifiers”—are crucial for discovering and quantifying heterogeneous effects. A prominent example is due to structural racism and discrimination, the effects of many important variables on health and life outcomes vary by race (Williams, Lawrence, and Davis 2019), with race often interacting with sex or socioeconomic status (Schoendorf et al. 1992; Bauer 2014). Cat-modifiers are also relevant for studying gene-environment interactions (Miao, Wu, and Lu 2024) and appear broadly in the social and behavioral sciences (Krefeld-Schwalb, Sugerman, and Johnson 2024).
Yet there are significant obstacles to the inclusion of cat-modifiers in linear regression analysis. Broadly, cat-modifiers alter the interpretation of the main effects, introduce concerns about equity across categorical groups (e.g., for race, sex, etc.), change the main effect estimates, and typically inflate the main effect standard errors (SEs). Consequently, cat-modifiers are often omitted or misreported (Knol et al. 2009), which falsely suppresses heterogeneity.
We argue that, with the right parameterization, cat-modifiers can and arguably should be included in linear regression models with categorical covariates. To establish ideas, suppose we have p continuous covariates x=x1,…,xp⊤ and K categorical variables C=C1,…,CK⊤ with Lk levels for each categorical variable k=1,…,K. We consider regression models for data xi,ci,yii=1n parameterized by a linear regression function μ(x,c) that typically models the conditional expectation E(Y∣x,c) or a transformed version for generalized linear models. We distinguish between two classes of linear those that do not include cat-modifiers and those that do. First, the main-only model includes multiple continuous and categorical variables, but no
(1)μM(x,c)=α0M+x⊤αM+∑k=1Kβk,ckM
where ck indexes the Lk levels of the kth categorical variable. In Wilkinson notation, (1) is yx1+⋯+xp+c1+⋯+cK. Second, the cat-modified model expands (1) to allow categorical-continuous and categorical-categorical
(2)μ(x,c)=α0+x⊤α+∑k=1Kβk,ck+∑k=1Kx⊤γk,ck+∑k=1K-1∑k′=k+1Kγk,k′,ck,ck′
or equivalently, yx1+⋯+xpc1+⋯+cK+c1c2+⋯+cK-1*cK, using pairwise interactions for convenience. Our notation emphasizes that the parameters in (1) and (2) are fundamentally distinct, even though these models are nested.
The advantage of the cat-modified model is the ability to estimate heterogeneous, group-specific effects for each xj. While both models specify group-specific intercepts (consider (1) and (2) with x=0), only the cat-modified model features group-specific slopes: (3)μxj′(c)≔μxj+1,x-j,c-μxj,x-j,c=αj+∑k=1Kγj,k,ck. Heterogeneity arises because the effect of xj varies across different groups c. By comparison, this does not occur in the mainonly μxjM′≔μMxj+1,x-j,c-μMxj,x-j,c=αjM. The remaining difference between (1) and (2) is that the latter model also includes categorical-categorical interactions via γk,k′,ck,ck′, which impact the group-specific intercepts.
For concreteness, we consider two popular cases. Empirical examples are given in Tables 1 and E.1, respectively, and these cases are revisited subsequently.
Suppose we have p=1 continuous variable x∈R and K=1 categorical variable race with LR groups. The main-only model (1) is then
(4)μM(x,r)=α0M+xα1M+βrM
or equivalently, yx+race, with group-specific intercepts, μM(0,r)=α0M+βrM for each race group r, but a global (race-invariant) slope, μxM′=α1M. Thus, (4) produces parallel lines with race-specific vertical shifts. By comparison, the cat-modified model (2) is
(5)μ(x,r)=α0+xα1+βr+xγr
or equivalently, yx+race+x:race, with group-specific intercepts μ(0,r)=α0+βr
and group-specific slopes μx′(r)=α1+γr for each race group r.
Suppose we have K=2 categorical variables race and sex with LR and LS groups, respectively. The main-only model (1) is then
(6)μM(r,s)=α0M+β1,rM+β2,sM
or equivalently, yrace+sex, while the cat-modified model (2) is
(7)μ(r,s)=α0+β1,r+β2,s+γrs
or equivalently, yrace+sex+race:sex.
The central challenge is that expanding from the main-only model to the cat-modified model alters the interpretations, estimates, and inference for the main effects: the parameters α0M,αM,βk,ckM in (1) or α0,α,βk,ck in (2). If these impacts are detrimental, then a quantitative modeler may be reluctant to include cat-modifiers. The key determinant is the model parameterization or identification strategy used for the categorical variable coefficients. Specifically, both models (1) and (2) require additional constraints to interpret and estimate the model the main-only intercepts α0M,βk,ckM are overparameterized, while the cat-modified intercepts α0,βk,ck,γk,k′,ck,ck′ and slopes α,γk,ck are overparameterized. The identifications determine the interpretations of all main and interaction parameters and the statistical properties of their estimators.
The most popular identification strategies are problematic for cat-modified models. Reference group encoding (RGE) is the overwhelming default, including for all major software implementations of linear regression (R, SAS, Python, MATLAB, Stata, etc.). With RGE, a reference group is selected for each categorical variable Ck and βk,1M=0 for all k in (1) and βk,1=0,γk,1=0,γk,k′,1,ck′=γk,k′,ck,1=0 for all (ck,ck′) in (2) (using 1 for each reference group without loss of generality). This is equivalent to using Lk-1 “dummy variables” to encode each Ck. Despite the simplicity of RGE, the implied notion of “main effects” in (2) significantly impedes the use of cat-modifiers. For the main-only model, the jth main effect is a global slope, αjM=μxjM′, invariant of c; yet for the cat-modified model, RGE fixes γj,k,1=0 for all j so that αj=μxj′(1) is the group-specific xj-effect with all categorical variables set to their reference groups (c=1).
First, this main effect parameterization is statistically SEs for αˆj are typically larger than those for αˆjM—the intersection of all reference groups is a subset of the data with a much smaller effective sample size—while the group-specific xj-effect μxj′(c) may be smaller for the reference groups (c=1) than for other groups or globally (i.e., αjM). We illustrate this effect in Table 1: with RGE, the main effects for the cat-modified model are attenuated and sacrifice power compared to those for the main-only model (see also Sections 4 and 5). Similar effects occur for categorical-categorical interactions (see Table E.1). Of course, these parameters refer to different functionals of μ(x,c); yet crucially, they are presented identically as “main effects” in statistical software output and manuscript tables (Knol et al. 2009). In fact, fewer than half of recent social science publications even reported the reference category (Johfre and Freese 2021). This leads to misleading conclusions about effect magnitudes, directions, and heterogeneity (Kowal 2024).
Second, RGE is the main effect elevates a single group above the others. In Table 1, White is the reference the main effect α1=μx′(White) is the x-effect for the White group, while the interaction effect γr=μx′(r)-μx′(White) is the difference between the x-effect for race r and that for the White group. The choice of reference group also determines which estimates, SEs, and p-values are output by default (Table 1). RGE presents the reference groups (main effects) as “normal” and the other groups (interaction effects) as “deviations from normal”. This framing is known to bias the interpretations of results—often to the detriment of underrepresented groups (Chestnut and Markman 2018). The issues are compounded for regularized when coefficient estimates are shrunk toward zero, the group-specific slopes are statistically biased toward the reference group slope (γr→0impliesμx′(r)→μx′(White)). Beyond the obvious inequities—including (racial, gender, etc.) bias in the estimators—this shrinkage obscures potential differences between the x-effects for dominant (e.g., White, Male) and nondominant groups. These issues are explored for race in Kowal (2024), but also raise serious concerns for other protected groups (sex, national origin, religion, disability, etc.). Thus, RGE undermines progress toward statistical methods that promote equity (Chen et al. 2021).
Finally, RGE is difficult to each main effect and interaction in (2) must be traced back to all reference groups. Consider Example 2 (and Table E.1), using White and Male for the reference the main effects in the cat-modified model (7) are α0=μ(White,Male), β1,r=μ(r,Male)-μ(White,Male) for each race r, and β2,s=μ(White,s)-μ(White,Male) for each sex s. Each main effect is anchored at both reference groups, which then affects the interpretations of the interaction effects γrs. This may be reasonable in some settings with a natural reference group (e.g., a control group), but RGE becomes increasingly difficult to interpret with multiple categorical covariates and interactions as in (2).
An alternative identification strategy uses sum-to-zero constraints (STZ). STZ identifies the parameters by restricting the group-specific coefficients via ∑ℓ=1Lkβk,ℓM=0 for all k in the main-only model and ∑ℓ=1Lkβk,ℓ=0,∑ℓ=1Lkγj,k,ℓ=0 for j=1,…,p, and ∑ℓ=1Lkγk,k′,ℓ,ck′=0 and ∑ℓ=1Lkγk,k′,ck,ℓ=0 for all (ck,ck′) in the cat-modified model. STZ is common for ANOVA models (Scheffe 1999; Fujikoshi 1993) and has been incorporated into regularized regression (Lim and Hastie 2015). STZ eliminates the need for a reference group, and thus resolves the inequities of RGE. However, STZ does not offer any special statistical properties for estimation of (1) or (2), nor does it establish a clear connection between the “main effects” in (1) and (2). As a result, it is difficult to interpret the parameters under STZ, while the addition of cat-modifiers may have unpredictable or detrimental effects on the main effect estimates and inferences (see Section 4).
To address these limitations, we advocate and analyze abundance-based constraints (ABCs) for identification and estimation with cat-modified models. Broadly, ABCs identify parameters using group abundances (see Section 2). ABCs are sufficiently general and may be combined with ordinary least squares (OLS), maximum likelihood, and modern regularized estimation techniques. The benefits are summarized by “EEI”: E****fficiency: ABCs permit the inclusion of cat-modifiers (a) without altering the main effect OLS estimates and (b) either maintaining or increasing their statistical power, under reasonable conditions;E****quity: ABCs do not require a reference group and thus eliminate the potential for introducing alarming inequities under default approaches (RGE); andI****nterpretability: main effects are identified as group-averaged parameters, interaction effects are group-specific deviations from these group-averaged parameters, and both sets of parameters inherit meaningful notions of sparsity. ABCs effectively remove the impediments to cat-modifiers, thus, facilitating richer regression analyses of heterogeneous effects. Of course, ABCs cannot guarantee that cat-modifiers will be practically or statistically significant, especially when the effective sample sizes for interactions are small. Rather, with ABCs, there is virtually nothing to lose by expanding from the main-only model (1) to the cat-modified model (2); yet the potential gains include greater statistical power for the main effects and discovery of heterogeneous effects.
We emphasize that ABCs, in various forms and by other names, have deep historical roots, but have lacked sufficient motivation to encourage widespread adoption. Scheffe (1999) and Fujikoshi (1993) considered identification strategies for ANOVA models based on arbitrary group-specific weights. Ultimately, both adopted STZ. Sweeney and Ulveling (1972) suggested ABCs for the simple ANCOVA (4) so that the estimated intercept would equal the sample mean (see also Theorem 1). However, there was no consideration of cat-modifiers or multiple covariates and no case made for any of EEI. Among nonlinear models, Park et al. (2021) and Park et al. (2023) used an ABC-like approach to avoid estimating main effects, instead focusing exclusively on interactions to optimize individual treatment rules. For the two-way ANOVA (7), Wang and Lin (2024) briefly mentioned ABCs only to dismiss them, claiming they “complicate the interpretation of the model parameters and make it difficult to fit the model…especially when other covariates are present.” Here, we forcefully argue the opposite, embodied by EEI—each of which applies with multiple covariates present.
Contrasts provide an alternative perspective on identification of (1) and (2): dummy coding, effects coding, and weighted effects coding (WEC) are respectively linked to RGE, STZ, and ABCs. Although WEC has garnered recent support (Grotenhuis et al. 2017a, 2017b), this work did not consider general cat-modifiers or any EEI. Instead, WEC has been mainly limited to two-way ANOVAs (6) or (7) and only advocated in restrictive settings with “certain types of unbalanced data that are missing not at random” (Brehm and Alday 2022) or “categories of different sizes, and if these differences are considered relevant” (Grotenhuis et al. 2017b). Our case is much broader and more ABCs are ideal to identify coefficients on any categorical variables and cat-modifiers should be included in many, if not all linear models.
Lastly, we acknowledge additional perspectives on cat-modified models. A widely-used approach is subgroup analysis, which subsets the data into groups (for all combinations of c) and then fits separate regression models (e.g., Pocock et al. 2004). The appeal is that it estimates group-specific slopes without the complicated interpretations of the parameters in (2) under default approaches (RGE). However, subgroup analysis does not provide estimates or inference for the main effects, cannot incorporate regularization or borrow information across groups, and does not allow direct testing for interaction effects. ABCs offer the same (and more) benefits without any of these drawbacks. Related, Searle, Speed, and Milliken (1980) advocated for marginal means. These quantities, like group-specific slopes and fitted values, are identical for all (minimally sufficient) identifications under maximum likelihood estimation. Thus, it does not distinguish among identification strategies. However, the identification strategy remains key for (a) parameter interpretation, estimation, and inference and (b) regularized regression and variable selection.
The article is organized as follows. We introduce ABCs in Section 2 for parameter identification and statistical estimation. Our main results on the properties of estimation and inference with ABCs are in Section 3. Simulation studies are in Section 4 and a real data example is in Section 5. We conclude in Section 6. Supplementary material includes proofs of all results, details for generalized linear models, additional theoretical and simulation results, and supporting data analysis. An R package lmabc is available online with detailed documentation and examples at https://drkowal.github.io/lmabc/.
The goal of ABCs is to enforce model identifiability while maintaining EEI. We first describe the model parameterization and interpretation, and then show how to compute regularized regression estimators and inference using linearly-constrained optimization. The main properties for estimation and inference are in Section 3.
For motivation, consider (5) from Example identifiability is obtained by constraining ∑r=1LRπrβr=0 and ∑r=1LRπrγr=0 for some chosen nonnegative weights πrr=1LR. RGE sets π1=1 and πr=0 for r>0, while STZ sets all πr=1. Equivalently, we can write each constraint as an average, EπβR=0 and EπγR=0, where R is a categorical random variable with probabilities πrr=1LR. This perspective offers a path to interpret the model parameters under the identifying weights πrr=1LR; all parameters are still fixed and nonrandom. Now, we can identify the main x-effect as the average of the group-specific α1=Eπα1+γR=Eπμx′(R). Then ABCs adopt a natural choice for πrr=1LR: the (population or sample) group abundances.
For illustration, Figure 1 compares the model parameterizations under RGE, STZ, and ABCs using the same data as Table 1. In each case, the slope parameters α1 and γrr=1LR have a different link to the group-specific slopes μx′(r)=α1+γr, due to the differences in the implicit choices of weights πrr=1LR. Here, ABCs use the sample proportions, (πˆW,πˆB,πˆH=(0.587,0.351,0.062) to identify the main x-effect as the average of the group-specific x-effects: αˆ1=Eπˆμˆx′(R)=πˆWμˆx′(W)+πˆBμˆx′(B)+πˆHμˆx′(H)=(0.587)(-0.022)+(0.351)(-0.052)+(0.062)(0.016)=-0.030. Arguably, this a natural and interpretable parameterization for a global (main) x-effect.
More broadly, we define ABCs for the cat-modified model (2); special cases such as (1) simply omit the constraints for the omitted parameters. We express the ABCs in terms of πˆ, which is the joint proportions across all categorical variables C=C1,…,CK⊤ in the data cii=1n. ABCs may be defined using population or sample proportions; we prefer the latter because they are always available and estimation properties are tractable and favorable (Section 3). First, ABCs for categorical main effects and categorical-continuous interactions are (8a)Eπˆβk,Ck=0,k=1,…,K (8b)Eπˆγj,k,Ck=0,k=1,…,K,j=1,…,p. Equivalently, ABCs may be expressed marginally and with ∑ℓ=1Lkπˆk,ℓβk,ℓ=0, where πˆk,ℓℓ=1Lk are the sample proportions for each categorical variable Ck,k=1,…,K, and then similarly for each γj,k,ℓℓ=1Lk. The key implication is that, while the cat-modified model (2) incorporates heterogeneity via mutual, group-specific slopes (3), ABCs concisely identify each main xj-effect as the average of this group-specific (9)αj=Eπˆμxj′(C). ABCs parameterize each main xj-effect by aggregating the group-specific slopes (3), each weighted by its respective abundance in the data. Unlike RGE, ABCs do not elevate any single (reference) group, and thus avoid the accompanying inequities.
The identification in (9) also guides interpretation of the group-specific slope parameters γj,k,ℓℓ=1Lk. Consider (5) from Example γr=μx′(r)-Eπμx′(R) is the difference between the group-specific slope for group r and the group-averaged slope. For the general cat-modified model (2), isolating γj,k,ck requires averaging over the remaining categorical variables C-k with joint proportions πˆ-k with Ck=ck γj,k,ck=Eπˆ-kμxj′C-k,ck-Eπˆμxj′(C). Further simplifications are often available, since these averages only must include the categorical variables that act as cat-modifiers for xj. In contrast with RGE, these group-specific coefficients are parameterized relative to a global main effect term (9), rather than a single (reference) group (White, Male, etc.).
For categorical-categorical interactions, ABCs identify γk,k′,ck,ck′ by requiring (10)EπˆCk∣Ck′=ℓγk,k′,Ck,ℓ=0,ℓ=1,…,Lk′EπˆCk′∣Ck=ℓγk,k′,ℓ,Ck′=0,ℓ=1,…,Lk for all (Ck,Ck′) interactions based on the conditional proportions for categorical variable Ck given that the interacting variable Ck′ belongs to group ℓ (and vice versa). We illustrate (10) using model (7) from Example EπˆS∣R=rγrS=0 for r=1,…,LR and EπˆR∣S=sγRs=0 for s=1,…,LS, where πˆS∣R=r=πˆrs/πrs=1LS is the conditional probability for each sex given race = r (similarly for πˆR∣S=s). Equivalently, (10) may be expressed using the joint proportions πˆ: for Example 2, this is ∑s=1Lsπˆrsγrs=0 for all r=1,…,LR and ∑r=1LRπˆrsγrs=0 for all s=1,…,LS.
In conjunction, (8) and (10) constitute ABCs for the general cat-modified model (2). All ABCs can be written using expectations Eπˆ(⋅) under the joint probabilities πˆ for the categorical variables C. We use these expectations to concisely identify and interpret model parameters; they do not change the general modeling context or assumptions for (2).
There are several compelling reasons to identify the categorical-categorical interactions with (10). First, it guarantees a global, group-averaged identification for the
Lemma 1. Under ABCs, the intercept parameter in (2) satisfies Eπˆ{μ(0,C)}=α0.
ABCs produce clean expressions and simple interpretations for these main while cat-modifiers induce group-specific intercepts μ(0,c) and group-specific slopes μxj′(c), ABCs identify global intercepts and global slopes by averaging over the groups, α0=Eπˆ{μ(0,C)} and αj=Eπˆμxj′(C), respectively. This cannot occur for RGE and only occurs for STZ if the probabilities πˆ are exactly uniform. Second, (10) orthogonalizes the main and interaction categorical in fact, the OLS estimates of the main categorical effects βk,ck are identical between models that do (7) or do not (6) include cat-modifiers (see Theorem 2 and its proof). Finally, (10) offers the interesting result that, if we were to instead combine the interacted covariates (Ck,Ck′) into a single categorical variable (e.g., race-sex) with LkLk′ levels, the main effect ABCs (8) would be satisfied for this new categorical variable. Of course, doing so would sacrifice the ability to estimate the main effects βk,ck, but this internal consistency is reassuring.
For implementation, it is sufficient to enforce Lk+Lk′-1 of the Lk+Lk′ constraints in (10). The choice of omitted constraint is arbitrary, since all constraints (10) hold
Lemma 2. Suppose we apply (10) to all but one interaction EπˆCk∣Ck′=ℓγk,k′,Ck,ℓ=0 for ℓ=1,…,Lk′ and EπˆCk′∣Ck=ℓγk,k′,ℓ,Ck′=0 for ℓ=2,…,Lk. Then the same constraint holds for ℓ=1:EπˆCk′∣Ck=1γk,k′,1,Ck′=0.
Finally, we emphasize that ABCs (8) and (10) are designed for parameter identification in the general cat-modified model (2), which may be featured in generalized linear models (see the supplementary material, Section B), is readily generalizable for multivariate response regression, and includes many special cases, such as main-only models (1), ANCOVA models (Example 1), two-way ANOVA models (Example 2), etc.
ABCs are linear constraints and thus readily compatible with regularized regression. First, we consolidate the cat-modified model (2) into a traditional regression μx1,c1,…,μxn,cn⊤=Xθ, where X is the n×P matrix that includes an intercept, all (centered) continuous covariates, indicator variables for all levels of each categorical variable, and all specified interactions, and θ include all unknown regression coefficients. In the presence of at least one categorical covariate, X is rank deficient, say rank(X)=P-m. We represent all ABCs (8) and (10) generically as Aπˆθ=0, where is the m×P matrix of constraints with rankAπˆ=m. Then, for a loss function ℒ(y,Xθ) for data y=y1,…,yn⊤ and a coefficient penalty 𝒫(θ), we aim to solve (11)θˆ(λ)=argminθℒ(y,Xθ)+λ𝒫(θ)subject toAπˆθ=0 and λ≥0 is a tuning parameter. We primarily focus on squared error loss ℒ(y,Xθ)=‖y-Xθ‖2 and either unpenalized estimation (λ=0) or (group) lasso and ridge regression with λ selected by cross-validation. One way to solve (11) is to reparametrize to an unconstrained space with only P-m parameters. Let Aπˆ⊤=QR be the QR-decomposition with columnwise partitioning of the P×P orthogonal matrix Q=Q1:m:Qπˆ and similarly, R⊤=R1:m,1:m:0. By construction, AπˆQπˆ=0, so that for any (P-m)-dimensional vector θQ, the vector θ=QπˆθQ satisfies Aπˆθ=0. Then, letting XQ≔XQπˆ, (11) is equivalently (12)θˆ(λ)=QπˆθˆQλ,θˆQ(λ)=argminθℒy,XQθQ+λ𝒫QπˆθQ. Regularized regression with ABCs simply requires (a) computing the QR decomposition of Aπˆ⊤ and (b) solving an unconstrained regularized regression problem.
When ℒ is a negative log-likelihood, θˆQ≔θˆQ(0) is a maximum likelihood estimator (MLE) and so is θˆ. Hence, usual properties for MLEs apply to estimators with ABCs. Under standard regularity conditions, (12) satisfies n(θˆ-θ)→dNP0,QπℐθQ-1Qπ⊤ where ℐ is the Fisher information associated with θQ and π is the joint population probabilities for the categorical covariates C. Thus, it is straightforward to construct confidence intervals and conduct hypothesis tests for the coefficients θ. When the model errors yi-μxi,ci are Gaussian, uncorrelated, and homoscedastic, the OLS estimator under ABCs satisfies θˆ~NPθ,σ2QπˆXQ⊤XQ-1Qπˆ⊤, even in finite samples. Although this distribution does not account for the sampling variability in πˆ, this is typically quite small relative to the variability in θˆ. Our empirical analyses suggest that no further adjustments are needed (see Section 4).
Finally, we emphasize the unique challenges of regularization and selection for cat-modified models. Selection of interaction effects has primarily focused on high-dimensional, continuous-continuous interactions (Bien, Taylor, and Tibshirani 2013). For cat-modified models with RGE, coefficient shrinkage introduces (racial, gender, etc.) γj,k,ck→0impliesμxj′(c)→μxj′(1), so group-specific effects are pulled toward those for the reference (White, Male, etc.) groups. With ABCs, no such biases γj,k,ck→0impliesμxj′(c)→Eπˆμxj′(C) collapses to the group-averaged xj-effect, which produces a reasonable notion of parameter sparsity.
When λ>0, it is possible to omit constraints and still obtain unique estimators. However, these estimators do not target identifiable parameters and thus are difficult to interpret. For lasso estimation, Kowal (2024) observes that such “overparametrized” estimation tends to reproduce RGE by implicitly selecting a reference group, and thus inherits the same limitations as RGE.
A central nuisance with interactions is that they change the main effect estimates and SEs. Here, we show that ABCs circumvent these challenges for cat-modifiers. The main point is that, with ABCs, the addition of cat-modifiers is either (a) harmless, since it has little to no impact on main effects estimates and inference, or (b) beneficial, since it can reveal heterogeneity and improve statistical power for the main effects. Our results make no assumptions about the true data-generating process for Y and do not apply for other identifications (RGE, STZ, etc.).
We establish conditions under which main effect OLS estimates are invariant to the addition of cat-modifiers under ABCs. First, consider OLS estimation of the intercept. For an enormous class of linear models—with arbitrarily many continuous covariates, categorical covariates, and categorical-categorical interactions—ABCs ensure that the OLS-estimated intercept is always exactly equal to the sample mean, αˆ0=y‾≔n-1∑i=1nyi.
Theorem 1. For any linear model of the form (2) with (a) centered continuous covariates (x¯=0), (b) no categorical-continuous interactions (all γk,ck=0), and (c) ABCs (8a) and (10), the OLS estimate of the intercept is αˆ0=y‾.
Simple models such as yrace yield the same intercept estimate as more complex models like yx1+⋯+xp+race+sex+race:sex. This reaffirms the global interpretation of the under ABCs, αˆ0 targets the global intercept α0=Eπˆ{μ(0,C)} (Lemma 1). Of course, y‾ is a good estimator for the marginal expectation of Y, so αˆ0 is appropriately global—even in the presence of categorical variables and their interactions. For models with at least one categorical variable, this result cannot occur for any other identification (RGE, STZ, etc.). Theorem 1 extends Sweeney and Ulveling (1972) to allow for categorical-categorical interactions and arbitrarily many continuous covariates.
Next, consider the impact of adding categorical-categorical interactions on estimation of the main effects. For concreteness, we consider a two-way ANOVA (Example 2).
Theorem 2. Under ABCs (8a) and (10), the OLS estimates of all main effects are identical under the main-only model (6) and the cat-modified model (7): αˆ0M=αˆ0,βˆ1,rM=βˆ1,r for all r=1,…,LR, and βˆ2,sM=βˆ2,s for all s=1,…,LS.
This estimation invariance applies to all (1+LR+LS) main effects in (7). Implicitly, we assume that the OLS estimates exist and are unique (i.e., empty categories are not permitted), but otherwise there are no requirements on the data-generating process. In particular, there are no assumptions of independence or uncorrelateness between the categorical covariates and no assumptions about their relationship with Y. ABCs deliver a natural interpretation of the categorical variable the main effects are deviations from the global mean (Theorem 1), while the interaction effects are deviations from the main effects in the main-only model (Theorem 2). This result validates our choice of ABCs (8) and especially (10). Again, such invariance does not occur for other identifications. We suitably generalize Theorem 2 in the supplementary material (Section C) to include additional continuous and categorical main effects as in (1).
Finally, we consider the addition of categorical-continuous interactions to main-only models. For clarity, we focus on a single categorical variable (K=1) with levels r=1,…,LR, but showcase these principles empirically with multiple categorical variables (Section 5). Following Example 1, we begin with a single continuous covariate x(p=1). Let σˆx[r]2≔nr-1sx[r]2-x‾r2 be the (scaled) sample variance of xii=1n for each group r, where nr=nπˆr,sx[r]2=∑ri=rxi2 and x‾r=nr-1∑ri=rxi. If the continuous covariate has the same scale for each group, then the OLS estimate of the coefficient on x is the same whether or not the cat-modifier is included.
Theorem 3. Under ABCs (8) and the equal-variance condition (13)σˆx[r]2=σˆx[1]2for allr=1,…,LR, the OLS estimates for the main-only model (4) and the cat-modified model (5) satisfy estimation invariance, αˆ1M=αˆ1.
We apply Theorem 3 as an approximation, αˆ1≈αˆ1M whenever σˆx[r]2≈σˆx[1]2 for all r, which is reasonably robust to deviations from (13) (see Table 1 and the supplementary material, Section D). This condition makes no requirements on the true associations between Y and (x,r) and generally allows the distribution of x to vary by r—as long as the scale is approximately constant. In particular, strong dependencies between x and r are permissible. Note that this result concerns categorical-continuous interactions only, so ABCs for categorical-categorical interactions (10) are not needed.
It is clarifying to consider violations of the equal-variance condition (13), so that x varies substantially for some groups, but varies little for others. This scenario does not invalidate estimation with ABCs; rather, it decouples the coefficients on x (i.e., the main x-effects) between the models that do (5) or do not (4) include a cat-modifier. Arguably, the main-only model (4) is no longer appropriate in this setting. The x-effect in (4) is α1M=μM(x+1)-μM(x), which considers a one-unit change in x regardless of the group r. But when (13) is violated, the scale of x—and a “one-unit change in x”—is no longer comparable across groups. Thus, group-specific slopes μ(x+1,r)-μ(x,r)=α1+γr are appropriate, which mandates the cat-modified model. Instead of the global x-effect α1M from the main-only model, the cat-modified model with ABCs identifies a global x-effect via the group-averaged quantity α1=Eπˆ{μ(x+1,R)-μ(x,R)} as in (9). In this setting, distinctness between α1 and α1M is appropriate. Finally, we do not require Theorem 3 (or the subsequent results) to use ABCs for estimation or inference. We contest that ABCs still provide more equitable and interpretable model output than other approaches.
This result can be extended for p continuous covariates, each of which is cat-modified: y~x1+⋯+xp+c1+x1:c1+⋯+xp:c1. Here, the equal-variance condition (13) instead uses the (scaled) sample covariance between xj and xh in group r,cov^rxj,xh≔nr-1∑ri=rxij-x‾jxih-x‾h.
Theorem 4. Consider the main-only model (1) and the cat-modified model (2), each with K=1 categorical variable. Under ABCs (8) and the equal-covariance condition cov^rxj,xh=cov^1xj,xh for all r=1,…,LR and each j,h=1,…,p, the OLS estimates satisfy αˆ=αˆM.
Theorem 4 ensures estimation invariance for all p continuous main effects, each of which is cat-modified. This equalcovariance condition is stricter than (13). As with Theorem 3, we apply Theorem 4 as an approximation, so that αˆ≈αˆM when equal-covariance approximately holds.
Lastly, we establish a middle yx1+⋯+xp+c1+x1:c1, which is a cat-modified model with p continuous covariates and K=1 categorical variable, but now only x1 is cat-modified. Instead of covariances between all pairs of covariates, the equal-variance condition now involves only x1 and the residuals eˆ1 from regressing x1 on all other variables, x1x2+⋯+xp+c1.
Theorem 5. Consider the main-only model (1) with K=1 and the cat-modified model (2) with K=1 and interactions only with x1 (fix γj,r=0 for all j>1 and r=1,…,LR). Under ABCs (8) and the equal-variance condition cov^rx1,eˆ1=cov^1x1,eˆ1 for all r=1,…,LR, the OLS estimates satisfy αˆ1M=αˆ1.
To understand this modified equal-variance condition, we can equivalently express cov^reˆ1,x1=nr-1∑ri=rxi12-xi1xˆi1=σˆx1[r]2-cov^rxˆ1,x1, where xˆ1 are the fitted values from x1~x2+⋯+xp+c1. Theorem 5 requires that the variability in x1 explained by the remaining (continuous and categorical) covariates is the same within each group. When this condition is violated, a one-unit change in x1 holding all else equal among x2,…,xp is no longer comparable across groups r. As with Theorem 3, the main-only model x-effect α1M is no longer appropriate; group-specific x1-effects μx1′(r)=α1+γ1,r are preferred; and ABCs offer a substitute for the global slope parameter via the group-averaged x1-effect, α1=Eπˆμx1′(R).
For further context, we clarify the role of true (or population) parameters in main-only and cat-modified models. The parameters are identified by the chosen constraint (RGE, STZ, ABCs); estimates computed under the constraint target those identified parameters. A notable exception occurs for the continuous main effects αjM from the main-only model (1): these parameters are not anchored to any identification, αjM=μxjM′, and their OLS estimates and inferences are identical under RGE, STZ, and ABCs (e.g., Table 1). For cat-modified models (2), no such agreement occurs—unless all cat-modifiers are extraneous, in which case the estimates αˆj all target αjM. This underscores the appeal of ABCs afforded by Theorems 3–5: when cat-modifiers are extraneous, ABCs adopt the more efficient estimators αˆjM from the main-only model that omits the extra terms. Otherwise, even when cat-modifiers are necessary, ABCs still target the identification-invariant main effects αjM.
While Theorems 3–5 focus on estimation invariance for continuous main effects, similar properties emerge for categorical main effects. With continuous-categorical interactions, such estimation invariance is (approximately) enabled by the well-known approach of centering the continuous covariates (x¯=0). However, unlike for continuous main effects (Theorems 3–5) or categorical main effects with categorical-categorical interactions (Theorem 2 and the supplementary material, Section C), this property is not unique to ABCs.
A primary reason for the unpopularity of cat-modifiers is the loss of statistical power for the main effects. With RGE, cat-modifiers relegate the main effects to a single reference group, which shrinks the effective sample size and often attenuates global effects. Thus, quantitative modelers may be reluctant to include cat-modifiers for fear of larger p-values, wider confidence intervals, and less power to identify important effects. Consequently, potential race-, sex-, or other group-specific effects may remain hidden.
ABCs directly and uniquely address this challenge. With the addition of cat-modifier effects, we show that ABCs may actually reduce SEs for the main effects. The magnitude of this reduction increases with the effect size of the cat-modifier. Crucially, when the cat-modifier effect is unnecessary, the main effect SEs match, but do not inflate, those for a (correct) main-only model.
Consider two nested models, a main-only model and a cat-modified model. Our general result is that the cat-modified model with ABCs has smaller SEs for the main effects whenever the estimated residual variance is smaller for the cat-modified model, (14)Sˆ2≤SˆM2. For the MLEs Sˆ2=‖eˆ‖2/n and SˆM2=eˆM2/n, where eˆ and eˆM are the residuals from the cat-modified and main-only models, respectively, (14) is ‖eˆ‖2≤eˆM2, typically with strict inequality. More commonly, the unbiased estimators Sˆ2=‖eˆ‖2/n-dM-d and SˆM2=eˆM2/n-dM are used, where dM+d and dM are the number of identified parameters for the cat-modified and main-only models, respectively. In that case, (14) requires that the adjusted-R2 for the cat-modified model exceeds that for the main-only model, or equivalently, (eˆM2-‖eˆ‖2/eˆM2≥d/n-dM, so that the (guaranteed) reduction in sum-squared-residuals from main-only to cat-modified must be large enough to justify the addition of d parameters. This requirement is adjusted-R2 is well-known to prefer overparametrized models, and thus (14) is likely to hold even when the cat-modifiers are extraneous (see Section 4.2). When cat-modifiers are indeed necessary, the reduction from SˆM2 to Sˆ2 can be substantial.
We revisit each case from Section 3.1, beginning with a twoway ANOVA (Example 2).
Theorem 6. Under ABCs (8a) and (10) and (14), the OLS SEs of all main effects under the cat-modified model (7) are less than or equal to those under the main-only model (6): SEαˆ0≤SEαˆ0M,SEβˆ1,r≤SEβˆ1,rM for r=1,…,LR, and SEβˆ2,s≤SEβˆ2,sM for s=1,…,LS.
Remarkably, Theorems 2 and 6 confirm that ABCs deliver the best possible adding cat-modifiers to the main-only model (6) does not change the main effect estimates, but potentially decreases their SEs. Thus, analysts may include cat-modifiers “for free”—with no negative consequences for the main effects—and gain the ability to infer possibly heterogeneous, group-specific effects. This result occurs precisely because ABCs ensure orthogonality between main and interaction effects—which notably does not occur for other (RGE, STZ) identifications.
A similar result applies for categorical-continuous interactions, again with the equal-variance
Theorem 7. Under ABCs (8), equal-variance (13), and (14), the OLS SE for the main x-effect under the cat-modified model (5) is less than or equal to that under the main-only model (4): SEαˆ1≤SEαˆ1M. This result applies in the context of Theorem 3, but analogous extensions are available for Theorems 4 and 5; only the condition (14) must be added.
These results make minimal assumptions about the true datagenerating process and do not require independence or uncorrelatedness among the covariates. However, the OLS SEs are defined as usual, which implicitly refers to uncorrelated and homoskedastic error assumptions for both the main-only and cat-modified models. Thus, while Theorems 6 and 7 are direct statements about the SEs as statistics and do not require any assumptions on the error distributions, the utility of these results is clearly linked to these assumptions.
The first objective is to verify the theory of ABCs for estimation and inference invariance, focusing on the conditions in Section 3. We present results here for categorical-categorical interactions and consider categorical-continuous interactions in the supplementary material (Section D).
Given two categorical variables, say race and sex, what is the effect of including the race:sex interaction term on the estimates and SEs for the race and sex main effects? The theory of ABCs (Section 3) predicts that the estimates will be exactly the same, while the SEs may decrease if the interaction effect is sufficiently large. These results make no requirements on the data-generating process. Thus, we simulate data such that (a) race and sex are dependent, (b) the errors are non-Gaussian, and (c) ABCs are not satisfied.
Let race and sex be categorical variables with groups {A, B, C, D} and {uu, vv}, respectively; we use arbitrary labeling here to remain agnostic about particular race- or sex-specific effects in our synthetic data-generating process. For each of n=500 observations, we draw each race assignment with (πa,πb,πc,πd)=(0.4,0.3,0.2,0.1), and then draw the sex assignment conditional on race with πuu,πvv⋅∣r=A=(0.4,0.6),πuu,πvv⋅∣r=B=(0.6,0.4),πuu,πvv⋅∣r=C=(0.7,0.3), and πuu,πvv⋅∣r=D=(0.2,0.8). Thus, race and sex are dependent, and marginally πuu=πvv=0.5. The response variable y is simulated with expectation (7) with α0=1,βc=-1,γb,vv=γ, and all other coefficients zero, or equivalently, μ(r,s)=1-I{r=C}+γI{r=B,s=vv} plus t4(0,1)-distributed errors. Crucially, γ controls the magnitude of the sex we consider γ=0 (no interaction effect), γ=0.5 (moderate interaction effect; see the supplementary material), and γ=1.5 (large interaction effect). We repeat this process to create 500 synthetic datasets.
For each simulated dataset, we fit the main-only model (6) and the cat-modified model (7) and compare the estimates and SEs for each main effect between the two models. These models are fit using ABCs, RGE (references r=A,s=uu), and STZ. This setting is favorable for RGE: the data-generating process satisfies RGE (βa=0,βuu=0,γa,uu=0), but not ABCs, and the reference groups are the most abundant groups for both race and sex. To aid comparisons, we omit the main effects from the reference groups, resulting in four main effects (βb,βc,βd,βvv) to compare between the main-only and cat-modified models for ABCs, RGE, and STZ.
The estimates are in Figure 2. Under ABCs, all race and sex main effects are exactly identical between the models that do and do not include the race:sex interaction, confirming Theorem 2. This result persists regardless of the true data-generating process, including the magnitude of the interaction. No such invariance occurs for RGE or STZ: the inclusion of the interaction completely changes the estimates (and the interpretations) of the main effects.
The SEs are in Figure 3. Under ABCs, the addition of the race:sex interaction has virtually no impact on the main effect SEs. For larger γ, the main effect SEs are slightly smaller (about a 5% reduction) for cat-modified model, as expected. More substantial SE reductions occur for larger interactions (γ≥5), although such large interaction effects are not usually expected in practice. Again, no such results occur for RGE: the SEs are much larger for the model that includes the race:sex interaction, regardless of γ.
These results must be interpreted the “main effects” under ABCs, RGE, or STZ target different functionals of μ(r,s). In fact, the OLS fitted values for μˆ(r,s) are identical under each identification (this is not the case for regularized regression). However, each identification puts forth “main effects” in both the main-only and cat-modified model. We argue that the main effects under ABCs are the estimates are exactly invariant to the inclusion of (race:sex) interactions and the SEs may decrease slightly. Uniquely, ABCs circumvent the traditional roadblocks to including the interpretations remain simple (and equitable) and there is no loss of statistical power for the main effects.
In the supplementary material (Section D), we revise this analysis for categorical-continuous given categorical race and continuous x, what is the effect of including the x:race interaction on the main x-effect? The theory of ABCs (Section 3) predicts that invariance for estimation and inference is contingent on the equal-variance condition (13). We investigate the sensitivity to this condition as well as to the magnitude of the interaction effect. Broadly, our findings resemble those from the categorical-categorical interaction case above—even with mild violations of the equal-variance condition (13), strong dependencies between race and x, and some model misspecification.
We evaluate the practical impacts of the estimation invariance and enhanced power of ABCs. The goal is to quantify the extent to which ABCs (a) maintain accurate estimates and precise uncertainty quantification when extraneous cat-modifiers are included and (b) improve estimation and reduce uncertainty when necessary cat-modifiers are included. The simulation design has three main features, described below.
First, we generate multiple, dependent categorical and continuous covariates. Dependent categorical variables race and sex are generated as in Section 4.1. Next, p=10 dependent continuous variables are generated as follows. We incorporate dependencies between race and xj for j=1,3,5,7,9 by drawing from the conditional distributions xj∣race=r given by (15)xj∣race=r~5+N(0,1)r=A12Uniform(0,1)r=B-5+t8(0,1)r=CGamma(1,1)r=D For each race group, xj has a unique distribution, with expectations that also vary by race. The remaining xj for j=2,4,6,8,10 are drawn independently from N(0,1). Some xj are highly correlated with race, which induces correlations among those x-variables with each other and with sex, while others are uncorrelated.
Second, the regression coefficients are constructed to satisfy both RGE and ABCs. These include an intercept α0=1, active main x-effects αj=1 for j=1,…,5, race main effects βb=1 and βc=βd=-1, and cat-modifiers γb,j=γ and γc,j=γd,j=-γ for j=1,…,5; the remaining coefficients are all zero. RGE is enforced because all reference group coefficients are zero (βa=0,βuu=0,γa,uu=0, and γa,j=0 for all j=1,…,p), while ABCs are satisfied for the population proportions πa,πb,πc,πd=(0.4,0.3,0.2,0.1). Thus, it is meaningful to compare coefficient estimates and inference between RGE and ABCs because they target the same ground truth parameters (STZ is not satisfied and thus excluded). ABCs actually use the sample proportions and are at a slight disadvantage. All sex main and interaction effects are zero, since this is the only way to satisfy both RGE and ABCs for a variable with two groups.
Third, a parameter γ controls the magnitude of the cat-modifier (x:race) effect. We consider γ=0 for extraneous cat-modifiers and γ=1.5 for necessary cat-modifiers.
Using these covariates and coefficient values, the response variable y is simulated with expectation (2) plus Gaussian errors and a signal-to-noise ratio of one. We vary the sample size n∈{200,500,1000} and repeat this process to create 500 synthetic datasets.
For each synthetic dataset, we fit the main-only model yx1+⋯+xp+sex+race and the cat-modified model Yx1+⋯+xp+sex)*race that includes race interactions with all continuous covariates and sex, both under ABCs and RGE. The main-only model is favored when γ=0, while the opposite is true for γ>0. Either way, the true data-generating process is sparse, so both models include many extraneous parameters. The main-only model includes 15 identifiable parameters (9 true signals) while the cat-modified model includes 48 identifiable parameters. When γ=0, the cat-modified model estimates 33 extraneous (identified) parameters; even when γ>0, only 15 of those cat-modifier effects are nonzero.
Evaluation primarily focuses on the main x-effects (α1,…,α10), which isolates the impacts of including extraneous (γ=0) or necessary (γ>0) cat-modifiers on estimation and inference for the main effects. For benchmarking, we also include evaluations for all (main and cat-modifier) coefficients. Note that for the main-only model, RGE and ABCs produce main x-effect estimates and SEs that are identical and nearly identical, respectively.
Estimation accuracy is evaluated by root mean squared error (RMSE) for the regression coefficients (Figure 4). The cat-modified model with ABCs preserves estimation accuracy of the main x-effects, even when (all 33) cat-modifier effects are included unnecessarily (Figure 4, top left). When some cat-modifiers are necessary, the cat-modified model with ABCs delivers slightly more accurate main x-effect estimates than the main-only models (Figure 4, bottom left). Neither result holds for RGE. By comparison, estimation accuracy across all coefficients (Figure 4, right) overwhelmingly favors the correctly-specified model (main-only for γ=0, cat-modified for γ=1.5), regardless of RGE or ABCs. This result is not surprising, but rather serves as a contrast to emphasize the extraordinary robustness of main effect estimation accuracy for cat-modified models—but only under ABCs.
Inference is evaluated by mean interval widths and empirical coverage for 95% confidence intervals for the regression coefficients (Figure 5); narrow intervals are preferred, subject to nominal coverage. For ABCs, the cat-modified model offers nearly the same statistical power for the main x-effects as does the main-only model, even when (all 33) cat-modifier effects are included extraneously (Figure 5, top left). Compare that to inference for all coefficients (Figure 5, top right): here, the inclusion of extraneous cat-modifiers inflates interval widths by more than 300%. Clearly, this inferential robustness against extraneous cat-modifiers is a special property for (a) main effects and (b) ABCs. When some cat-modifiers are necessary, the cat-modified model with ABCs improves statistical power for the main x-effects compared to the main-only models. Again, no such results hold for RGE, for which the cat-modified model consistently sacrifices statistical power. Finally, as expected, main-only models fail to provide coverage for active cat-modified parameters (Figure 5, bottom right).
The supplementary material includes additional results for smaller (n=200) and larger (n=1000) sample sizes; predictive evaluations based on RMSEs for μ(x,c); and comparisons between ABCs and RGE for lasso and ridge regression, also including an “overparameterized” version that does not impose any constraints.
We apply cat-modified regression to assess heterogeneity among factors linked to STEM educational outcomes. Our dataset^1^ links three administrative datasets to provide individual-level data for n=27,638 children in North Carolina (NC): NC Detailed Birth Records, NC Blood Lead Surveillance, and NC Standardized Testing Data; details are provided elsewhere (Initiative 2020; Kowal et al. 2021; Bravo et al. 2022). The STEM educational outcome variable yi is the end-of-4th-grade standardized math score for student i, centered and scaled by year of test (2010, 2011, or 2012). These math scores are linked with a rich collection of demographic, social, and environmental exposure variables. The continuous covariates are racial (residential) isolation (RI), which is a measure of structural racism based on neighborhood information; blood lead level (BLL), which measures lead exposure; birthweight percentile (BWTpct); mother’s age at time of child’s birth (mAge); and exposure to the air pollutant PM2.5 during the year prior to the exam (PM2.5). The continuous covariates are centered and scaled. The categorical covariates are mother’s race (race), child’s sex (sex), mother’s education level (mEdu), and an indicator of economically disadvantaged (EconDisadv) determined by participation in the National Lunch Program; see Table 2 for categorical levels and proportions.
Our linear regression analysis spans from main-only models to a variety of cat-modified models, expanding significantly upon the earlier models from Tables 1 and E.1. First, the main-only model includes each of these covariates (RI, BLL, BWTpct, mAge, PM2.5, race, sex, mEdu, and EconDisadv) but no interactions. The main-only model features a variety of interesting demographic, socio-economic, maternal, and environmental exposure variables, with 16 regression parameters (12 identified). Next, the race-modified model adds an interaction between race and every other covariate. This expansion allows for heterogeneous effects of each variable by race, thus providing insights into the myriad impacts of race on each child’s life course and educational outcomes, with 52 regression parameters (30 identified). Finally, the cat-modified model adds all pairwise categorical-continuous and categorical-categorical interactions. This instance of (2) allows the fullest (pairwise) extent of heterogeneous effects across the rich collection of demographic and socio-economic variables (race, sex, mEdu, and EconDisadv), with 103 regression parameters (55 identified). We fit each of these models under ABCs and RGE (references White, Male, lowest mEdu (mEdu<HS), and not EconDisadv).
While each model offers potential for insight, a critical limitation of popular identification approaches, especially RGE, is that the estimates, inference, and interpretations of the main effects are highly sensitive to the choice of cat-modifiers. To see this, we present the main effect OLS estimates and 95% confidences intervals across these models in Figure 6. With RGE (right), the main effects shift and the intervals widen considerably—with increases from 160% to 230% in interval widths—upon adding race- (blue) and other (red) cat-modifiers. These main effects and accompanying interaction effects (not shown) are anchored at the reference groups and refer to different functionals of μ(x,c) under each model—even though the statistical output for the “main effects” is typically presented identically, regardless of any cat-modifiers. Thus, while cat-modified models are essential for heterogeneous effects, there is a cost incurred under RGE: each additional cat-modifier requires careful re-consideration of the main and interaction effects, which impedes statistical analysis and undermines interpretability.
The invariance of ABCs resolves these estimation and inference for the main effects (Figure 6, left) are nearly identical across these substantially different models. This occurs despite strong dependencies among the covariates (and interactions) with both continuous and categorical variables. ABCs effectively decouple the main effects from the cat-modifiers: even adding 87 parameters (43 identified) from the main-only model to obtain the cat-modified model does not lessen, and in some cases increases the statistical power for the main effects. With ABCs, the statistical analyst may consider these or other cat-modified models without compromising or complicating inferences for the main effects.
The full regression output from the cat-modified model with ABCs is in Tables 2 and E.3. Lower math scores are strongly (p<0.01) associated with racial (residential) isolation, lead exposure, lower (mother’s) education levels, and occur for non-Hispanic Black and economically disadvantaged students; higher math scores are strongly associated with birthweight percentile, mother’s age, and the opposing categories from above. ABCs provide output for all levels of all categorical variables, thus, eliminating the presentation bias of RGE that presents all output relative to the reference groups (White, Male, etc.). The regression output strongly supports heterogeneous effects, most notably via mother’s education level and with intersectionality of race and sex (e.g., Bauer 2014).
Finally, we simplify the heavily-parameterized cat-modified model by fitting a lasso regression under ABCs; λ is selected using 10-fold cross-validation and the one-standard-error rule (Hastie, Tibshirani, and Friedman 2009). The selected main effects (RI, BLL, BWTpct, mAge, race, mEdu, and EconDisadv) match the conclusions from Figure 6. Among interactions, coefficients from race:mEdu, race:mAge, race:EconDisadv, mEdu:mAge, and EconDisadv:mAge are selected. The accompanying coefficients of these cat-modifiers suggest that some positive effects are not as beneficial for minoritized the positive effect of mother’s education (mEdu>HS) are attenuated for Black and Hispanic students, while the benefits of mother’s age are less so for lower mother’s education, Black, or economically disadvantaged students.
To encourage and enable statistical analysis of heterogeneous effects, we analyzed and advocated ABCs—an alternative parameterization and estimation strategy for cat-modified models that include categorical-continuous or categorical-categorical interactions. Unlike default methods, ABCs allow the inclusion of cat-modifiers “for free”: there is virtually no impact on the main effect estimates, while main effect inference is stable or more powerful. We rigorously proved these estimation and inference invariance properties and validated them empirically with extensive simulation studies. We also provided strategies for estimation and inference, including both generalized and regularized regression. Finally, we applied these tools to analyze STEM educational outcomes and showed how ABCs facilitate identification and estimation of (demographic) heterogeneous effects without incurring any costs—in estimation, inference, or interpretation—for the main effects.
Despite these many advantages, we note several caveats. First, ABCs may increase susceptibility to p-hacking. Because ABCs facilitate the inclusion of interactions, and with a large enumeration of potential interactions, there is a heightened potential for both discovery and false discovery. Proper statistical analyses require careful consideration of hypothesis tests with multiple testing corrections as appropriate. Second, ABCs cannot guarantee that cat-modifiers will be (practically or statistically) significant. Detection of heterogeneous effects often requires well-designed studies or large sample sizes. Finally, many categorical variables, especially race, sex, and other protected groups, are susceptible to misinterpetation, inaccurate labelings, and exclusions of small groups.
Promising avenues for future work include generalizations for additive models, Bayesian regression, and high-dimensional regression, where the invariance properties of ABCs may prove especially useful for both variable selection and efficient computing strategies. It is also possible to use ABCs for higher-order interactions. With one categorical term (e.g., categorical-continuous-continuous interactions) or two categorical terms (e.g., categorical-categorical-continuous interactions), we may immediately apply (8b) or (10), respectively. For higher-order categorical interactions, suitable generalizations of the conditional expectations in (10) may be developed.
Supplementary material includes a PDF with proofs of all results, details for generalized linear models, additional theoretical and simulation results, and supporting data analysis; and R code files for reproducibility. An R package lmabc is available online with details and documentation at https://drkowal.github.io/lmabc/.
Supplementary materials for this article are available online. Please go to www.tandfonline.com/r/JASA.