Authors: Jennah Gosciak (https://ror.org/05bnh6r87Cornell University, USA), Daniel Molitor (https://ror.org/05bnh6r87Cornell University, USA), Ian Lundberg (https://ror.org/046rm7j60UCLA, USA)
Categories: Article, experimental design, multiple treatments, adaptive randomization, conjoint experiments, human preferences
Source: Political Analysis
Authors: Jennah Gosciak, Daniel Molitor, Ian Lundberg
Human choices are often both multi-dimensional and interactive. For example, a person deciding which of two immigrants is more worthy of admission to a country might weigh their education, and the weight placed on education may depend on other factors, such as their age, country of origin and employment history. We develop a response-adaptive experimental design that summarizes the range of effects of one attribute as a function of all other attributes. Our approach changes several aspects of the experimental design based on the ex ante choice to study the heterogeneous effects of one focal attribute (i.e., education). We update treatment assignment probabilities over the course of the experiment to search for the attribute vector at which the focal attribute has the most positive and most negative effects. By summarizing the full range of effects that exist, our approach complements existing approaches to conjoint experiments that typically aggregate over heterogeneity by marginalizing. We illustrate through two online experiments and provide customizable code infrastructure via a Docker container that other researchers can use to deploy adaptive randomization in online conjoint experiments.
Multi-dimensional choices abound in social life. When a voter chooses between two political candidates, those candidates may differ along several dimensions, such as age, gender and party affiliation. When a citizen considers which of two immigrants is worthy of admission to their country, they might weigh the immigrants’ education, country of origin and employment history. Conjoint experiments, which present respondents with a choice between two fictitious profiles comprised of several randomized attributes, have rapidly improved scientific understanding of multi-dimensional human choices (Bansak et al. 2021; Hainmueller, Hopkins, and Yamamoto 2014).
Often, these choices are interactive. For example, a long literature in political science claims that voters prefer co-ethnic candidates. In a study of Ugandan voters, Carlson (2015) shows that the preference for a co-ethnic candidate is stronger when that candidate has a good performance record. That may not be the only interaction; a reanalysis by Egami and Imai (2019) further demonstrates that the effect of co-ethnicity varies by the candidate’s platform (e.g., promoting jobs vs. improving education). What if co-ethnicity has effects that vary by an interactive function of even more attributes? When multi-dimensional choices involve interactions among several variables, these interactive preferences become difficult to summarize. A popular strategy is to summarize the interaction between two or three variables over the joint distribution of all other variables. We consider a different strategy to provide two summary the most positive and most negative effects of a focal attribute, conditional on values of the other attributes. We develop an experimental design that adaptively searches for these summary statistics. Our design summarizes the full range of effect sizes that an attribute may have, given particular values that the other attributes in the study can take.
To apply our design, a researcher first chooses a focal attribute of theoretical interest (e.g., candidate co-ethnicity). This choice differs from traditional conjoint experiments, which typically provide parallel evidence on the causal effects of many attributes. By prioritizing one attribute, our approach is more similar to a broad social science literature on human choices. Experiments often begin from a strong theoretical interest in one particular attribute. As one class of examples, audit studies have shown how race (a focal attribute) shapes human choices in hiring (Bertrand and Mullainathan 2004; Correll, Benard, and Paik 2007) or housing (Yinger 1986). For researchers most interested in one attribute, our approach enables a new form of how to search for strong causal interactions between the focal attribute and other attributes.
Our design is most helpful when the focal attribute interacts strongly with other attributes. These settings may be common in both industry and the social sciences. One historical example from industry is an experiment by Honda which revealed that consumers preferred sliding doors (as opposed to swing doors) in their vehicle only if that vehicle was a minivan. The attribute (door type) interacts with the attribute (vehicle type) (Sawtooth Software 2024). Interactions are also widespread in social science studies of human choices, as illustrated by an example study on racial inequality. Steffensmeier, Ulmer, and Kramer (1998) analyzed judicial sentencing decisions and showed that judges sentence young Black men more severely than would be predicted by an additive model for the effects of race, gender and age. By focusing on attribute interactions, researchers can discover new patterns in human judgments.Figure 1Standard and adaptive conjoint A comparison.
Our approach is complementary to existing strategies for interactive preferences that discover low-order interactions (Egami and Imai 2019) or that test for the presence of any effect of one attribute conditional on any vector of values for all other attributes (Ham, Imai, and Janson 2024). Once a study finds interesting interactions or rejects the null that an attribute has no effect, it is natural to ask questions about the total range of effects that attribute may have. Models involving two-way interactions may poorly approximate this total range if higher-order interactions exist (Appendix E). Our method searches efficiently for the total range of causal interaction using adaptive randomization.
The article proceeds in several sections. We first formalize our experimental design and causal estimands, allowing us to draw connections to past work on conjoint experiments (Figure 1). We then present our approach to adaptive randomization, drawing connections to past work on adaptive experiments. We illustrate our approach with two empirical examples before concluding with a discussion. ^1^
Our causal estimand and experimental design differ from those of a typical conjoint experiment in several ways, which we introduce in this section through a motivating example and mathematical notation.
We ground our method in a well-known example from the conjoint literature on attitudes among U.S. citizens toward immigrants. Hainmueller and Hopkins (2015) test the effects of nine different applicant attributes that each take on between 2 and 10 values. Hainmueller and Hopkins (2015) summarize each attribute’s average marginal effect, for example, estimating that a bachelor’s degree increases the likelihood of admission compared to no formal education by 20 percentage points.
We focus on what if the effect of having a bachelor’s degree differs by the values of the other attributes? Instead of averaging away the heterogeneity, one may want to summarize the range of heterogeneous effects. We answer this question with a three-part experimental design that we implement in a simplified replication of Hainmueller and Hopkins (2015), randomizing a binary indicator of education and four binary indicators of other attributes.Figure 2Elements of an adaptive conjoint design. Our design focuses on the causal effect of a randomized focal attribute, as it causally interacts with randomized context attributes. Each context attribute takes a value, and we refer to a vector of context attribute values as a context within which the focal attribute may have an effect.
We present each study participant i with a single pair of profiles for fictitious prospective immigrants and ask which is more worthy to be admitted to the country. Let indicate whether the participant selects the left or right profile. Profiles involve several elements, which we introduce mathematically here and visually in Figures 2 and 3. Let
indicate whether the left or right profile features a prospective immigrant with a college degree (vs. no formal education). Because we are interested in the causal effect of
on the choice
, we refer to
as the focal attribute. We refer to all other attributes as context attributes. Our context attributes include prior trips to the United States, profession, reason for application and country of origin. Let
denote the vector of context attributes assigned to respondent i, for example, an immigrant from Europe applying for employment reasons who has many prior trips to the United States and works in a skilled profession. Each element of
takes one of two possible values in our applications. It is straightforward to generalize to attributes with more than two values, but doing so produces more possible profiles. With two possible values for each of four attributes, there are
unique values of
. With three possible values, there are
unique values of
. The dimensionality of the task grows with both the number of attributes and the number of values per attribute.Figure 3Context attributes, values and signals. Every profile has a set of context all attributes other than the focal attribute. Our design adaptively randomizes a vector of attribute values across profile pairs; within a pair, the two profiles share identical attribute values. To ensure that the two profiles in the pair do not appear identical to each other, we randomly permute two signals of each value across the profiles. Our experimental design therefore requires the researcher to specify two signals for each attribute value.
In the experiment, respondents view two profiles that differ on the focal attribute
but share the same context attribute vector
. For example, both profiles may show immigrants from European countries of origin. To reduce suspicion, we create two signals of each context attribute value. One profile is from Poland and one is from Germany, for example (see Figure 3 for more examples). Each of the two profiles in a choice pair presents the respondent with a different vector of signals, even though the context attribute value is the same for both. Because signals are randomly assigned within contexts and our estimands (defined in the next section) marginalize over the random permutation of signals, these estimands are well-defined and identified even if signals have effects (e.g., respondents prefer immigrants from Poland over Germany). However, in practice, researchers should choose signals that indicate the same underlying context relevant to the choice. Optimally, the signals have minimal effects so that an estimand marginalized over random signals is marginalized over a nearly constant potential outcome.
When randomizing the signals within a context, there are at least two viable strategies. One strategy is to pre-specify two vectors of signals and make a single randomization for whether a particular vector appears on the left or the right,
. Advantages of this strategy include notational simplicity and potentially simpler deployment in software. We use this strategy for our main empirical illustration. A second strategy is to independently permute the signal possibilities for each attribute in the profile, so that the vector would be
for p attributes. An advantage of independent randomization is that causal effects of individual signal elements are identified. We use independent randomization in our second empirical illustration.
The respondent’s choice is a function of the focal attribute, context attributes and signals presented to that respondent. Let
denote the potential choice outcome for respondent i if exposed to profiles with context attribute vector
, signal permutation s and focal attribute permutation a. Each respondent has many potential outcomes, one for each combination of
,
and
. We assume consistency such that
. Let capital letters without subscripts refer to random variables that are random across respondents, so that
are the stochastic analogs of the particular values
taken for respondent i.
To preview our randomization (discussed in Section 3), a central phase of our experiment assigns new participants to a context
as a function of past data, and randomly permutes the focal attribute A and signal S across the two profiles in a pair. The purpose of the design is to study how the focal attribute A affects the respondent’s choice within a context
.
Define the choice probability within context
as marginalized over the random permutation of the focal attribute and signals, with what probability is the chosen profile Y the same as the profile A that signals a selected value of the focal attribute? (1)
In our motivating example,
corresponds to the probability that a respondent presented with a pair of profiles with context attribute values
would choose the applicant with a college degree over the applicant with no formal education. Because of the forced-choice design, a value of
would imply that the focal attribute has zero effect on the choice (analogous to a marginal mean, Leeper, Hobolt, and Tilley 2020). Appendix F formalizes connections to common conjoint estimands. A key difference is that our estimand involves the effect of A holding all other attributes
at a particular value
, rather than marginalized over a distribution of
. A benefit is that one need not specify a population distribution of
(as in De la Cuesta, Egami, and Imai 2022).
Our non-marginalized estimand creates a different many unique
values (many contexts) produce correspondingly many unknown parameters
. With four binary attributes, there are
parameters
. It may be costly to gather enough data to estimate them all precisely. An experiment with
participants would on average allocate only
participants to each context
. High-variance estimates
might leave the researcher with little confidence about which
is largest and which is smallest. Too few samples are spread across too many estimands. When the support of
is truly high-dimensional, for example, with hundreds of thousands of unique values, there is no hope of exploring the full space without assumptions that pool information across
-values. But when the support of
is only moderately large, such as 16 unique
-values, adaptive randomization can explore the space more efficiently than fixed randomization. We search for two summary the choice probabilities in contexts
and
that maximize and minimize
: (2) (3)
These estimands summarize the total range of effect sizes by which the focal attribute A may affect the choice Y across values of the context vector
.
Several key design decisions differ under our approach compared with a standard conjoint.
The researcher must choose the focal attribute. Theory may guide the choice, as when audit study experiments focus on the effect of race as an attribute of particular importance (e.g., Bertrand and Mullainathan 2004). Past research may also guide the choice, for example, if low-order interactions are present (Egami and Imai 2019) or if statistical tests point toward large, unexplored heterogenous effects (Ham et al. 2024). Ideally, the focal attribute will have large causal interactions with other attributes. But if it has minimal causal interactions or even zero effect, the result may still be interesting, as we show in our second empirical illustration. In this case,
. One can more confidently conclude that an attribute has nearly homogeneous effects if one searches explicitly for heterogeneity and finds nothing.
A second choice is how to select signals of context attributes. For each context attribute value (e.g., Eastern Europe), our design requires two signals (e.g., Poland and Germany) so that the two presented profiles are not identical. It is ideal to choose distinct signals that have minimal implications for the respondents’ choices. For example, it is simplest if respondents are thinking about Europe (the context) rather than about the particularities of Poland and Germany (the signals). Researchers should choose signals that avoid social desirability bias; Germany and Poland would be a bad signal pair if one was a more socially desirable preference than the other, alone or in combination with other attributes. It is of course not possible to know in advance when designing the experiment. Two considerations offer (1) the estimands
remain well-defined and identified even if signals have effects and (2) one can estimate signal effects directly (Appendix B and C). Researchers should choose signals that they believe will not affect the choice and should assess the effects empirically.
A third design choice further reduces we show each participant only one pair of profiles. A participant presented with a sequence of choices might come to recognize that they all differ on the focal attribute. Our choice comes at a cost, because traditional conjoints get much more data per respondent through repeated choices (Bansak et al. 2018). An advantage that partially compensates for this change is that the survey is shorter and thus costs less per respondent.
Recall that our goal is to discover the contexts
and
that maximize and minimize an unknown parameter
. The problem is analogous to the well-studied multi-armed bandit problem of choosing among many experimental arms to maximize an unknown payout. This section reviews existing work in adaptive experimentation and then formalizes the adaptive design we follow in our experiment.
Experimental research classically follows fixed (non-adaptive) designs. For example, a pre-specified number of units are randomized to treatment A or B by known, fixed probabilities. The fixed design brings many advantages. It is easy to understand. Researchers can pre-specify the design; nothing changes as a function of the data collected. Yet these advantages also correspond to limits of the fixed a researcher typically (though not always) selects a small number of treatment values (often two) and collects a pre-specified amount of data even if the answer gradually becomes clear before the end of the study.
In adaptive designs as we consider them in this article, the treatment assignment probabilities change over the course of data collection. In one of the earliest statements of adaptive designs, Thompson (1933) considered two interventions A and B. Over the course of an experiment that randomizes cases to A or B, imagine that it gradually becomes clear that A leads to better outcomes. Must we continue assigning new research participants to B? Thompson (1933) instead proposed a Bayesian updating procedure in which each participant’s probability of assignment to treatment A equaled the posterior probability that A was the more effective treatment, given evidence from all participants who had come before. Under the Thompson sampling design, the probability of assignment to treatment A rises smoothly to 1 as evidence builds that A is the more effective treatment.
Adaptive designs are especially helpful in randomized trials with high-stakes consequences, such as those in which assigning a person to an inferior treatment may result in death. The gains to social welfare that come from assigning units to effective (rather than ineffective) treatments in real time over the course of a study are one reason adaptive designs have become increasingly common in clinical trials (Chow and Chang 2008; Pallmann et al. 2018; Rosenberger and Lachin 1993; Villar, Bowden, and Wason 2015; Yao et al. 2021).
Even in settings with lower stakes, adaptive designs are desirable because they can yield substantial efficiency gains, reducing the cost of data collection and improving the selection of the best treatment arm. Thus, adaptive experiments have begun to appear in economics (Kasy and Sautmann 2021) and political science (Offer-Westort, Coppock, and Green 2021).
In our experimental design, the arms of adaptive randomization are the contexts
, and we seek to find the maximum and minimum of
, the choice probability which is defined for each arm. We carry out randomization in phases—warm-up, adaptive and validation—to avoid some difficulties that otherwise threaten inference in adaptive experiments. Appendix G formalizes the procedure in pseudocode. ^2^
At the beginning of the experiment, we have no knowledge of the unknown
probabilities. We begin by collecting a set of
observations equally distributed across the contexts. The warm-up phase provides initial evidence about all
parameters, enabling one to proceed to adaptive randomization with some evidence rather than solely a subjective prior.
After the warm-up phase, we begin assigning treatments with unequal probabilities adaptively as a function of past data. There are two adaptive one for
and one for
. One can carry out these searches in parallel, using only the warm-up data as the initial input to each adaptive randomization search. We follow this strategy in our second illustration with job candidate profiles. One can also carry out these searches in sequence, for example, searching first for
and then for
. We follow this strategy in our first illustration with immigrant profiles. We recommend the sequential (rather than parallel) approach because it makes the second search more the data from the first search support better treatment assignments in the second search. An open direction for future research would be to identify an optimal mixing of the two searches over time, using all cumulative data in each search.
The adaptive phase begins with a prior on the unknown choice probabilities
. We use a Beta distribution as the conjugate prior to a Bernoulli choice, but generalizations to categorical choices (Dirichlet prior, Categorical likelihood) or continuous ratings (Normal prior, Normal likelihood) are analogous. For an example of adaptive updating with a categorical response, see Deliu (2024). In our setting, we assume a uniform Beta(1,1) prior on each unknown choice (4)
Using conjugacy, we update the posterior probabilities after seeing a set of observations indexed by i, (5)where
and
correspond to the matrix of assigned contexts and to the vector of observed outcomes, respectively.
For each respondent, we simulate many values of each
from its posterior distribution given data from all previous respondents. For each simulation, we identify which context yields the highest and lowest values of the unknown probability. Averaging these results over all simulations provides our Monte Carlo estimates of the probabilities that each context maximizes or minimizes the value of
(6) (7)In our empirical illustration, we base these estimates on 100,000 simulated draws. ^3^ We use the estimated posterior probabilities
(or
when searching for the minimum context) to assign new survey participants to contexts.
The adaptive phase can be carried out in continuous fashion (updating posteriors after each respondent) or discrete fashion (updating posteriors after each batch of respondents). Our first empirical illustration and accompanying software support the continuous approach. Our second illustration uses discrete batches. We stop when a budgeted sample size has been reached. Future research could explore alternative stopping rules, such as halting the process when a target posterior probability threshold is reached.
At the end of the adaptive phase, we select two key contexts,
and
. These are the contexts with the highest probability of having the most positive and most negative focal attribute effects, respectively, across posterior draws given all data. Researchers should report these posterior probabilities as evidence for their degree of certainty. We also report posterior mean estimates
and
from all data collected to this point (warm-up + both adaptive phases). However, for reasons explained below, our preferred estimates of these parameters come from the validation phase, which we describe next.
The final phase of our experiment is a validation we collect new data in which respondents are assigned to the two selected contexts with fixed, equal probabilities. ^4^ The validation phase addresses a problem known as the winner’s curse.
After the adaptive phase, there remains statistical uncertainty about the true (and unknown) values of
. The variance
may be non-negligible. Suppose there were dozens of contexts and a comparably small sample size in the adaptive phase. The context with the highest estimated value is not necessarily the context with the highest true it is possible that
. In fact, which context is chosen depends partly on the true signals in each context and partly on the luck of which context happens to have a high estimated value in the particular adaptive sample analyzed. On average, over many trials, the winner has positive noise. The answer is to re-estimate the winner’s parameter
in validation data where the noise is independent of the choice. ^5^
The winner’s curse may seem unfamiliar, but it is the same as the well-known motivation for training and validation sets to evaluate the performance of predictive algorithms. In such settings, a researcher selects among many algorithms using a training sample. But predictive performance metrics in a training sample are biased estimators of out-of-sample performance due to the winner’s the fact that an algorithm wins in the training sample is correlated with the outcomes in the training sample. Just as an independent test sample solves the winner’s curse in evaluations of predictive algorithms, an independent validation sample solves the winner’s curse in our setting.
As an added benefit, the validation phase simplifies statistical inference. Inference on the adaptive phase is made challenging by the fact that treatment assignment probabilities are unequal across respondents as a function of past data. Unbiased estimators are possible through inverse probability of treatment weighting, yet challenges to inference can remain due to issues that arise with asymptotic normality (Hadad et al. 2021). Several approaches have been proposed to address this inferential challenge (Hadad et al. 2021; Zhang, Janson, and Murphy 2020, 2021). These complex solutions are not necessary in our the validation phase is a classic experiment with fixed assignment probabilities. Unweighted mean estimators are unbiased and asymptotically normal.
To summarize, the goal of our adaptive experimental design is to search for
and
. A central reason these estimands are of interest is that
for all
, so that these two estimands bound the total range. There are two important limitations that should inform design choices.
The first limitation arises because of statistical it is possible that the wrong contexts are selected, for example,
. If you pick the wrong context, then the validation estimates will be unbiased for the selected context while being downwardly biased for the parameter of the context that is the true
. By allocating more sample to contexts with high probabilities of being the true maximum, adaptive randomization reduces the risk of selecting the wrong context (simulation in Appendix D). Researchers can also be transparent about their uncertainty by reporting the posterior probabilities that the chosen contexts are the minimum and maximum.
A second limitation occurs even when the warm-up and adaptive phases lead to correct choices for the extreme contexts. After selection, the parameters of those contexts must be estimated in the validation phase. While validation estimates are unbiased, they will still be statistically uncertain. One can reduce this risk by allocating more samples to the validation phase and can make this risk transparent by reporting estimates of uncertainty (e.g., 95% credible intervals) for estimates from the validation phase.
These limitations can inform choices about how to allocate sample size across phases. On the one hand, allocating more samples to the warm-up and adaptive phases reduces the chance of selecting the wrong context. On the other hand, allocating more samples to the validation phase supports precise estimation for the selected contexts. The correct balance will depend on the application.
When allocating sample across the warm-up and adaptive phases, a few anchoring points are useful. To allocate all cases to warm-up is equivalent to running a fixed-randomization trial. To allocate all cases to the adaptive phase is a purely Bayesian approach, allowing the posterior to update treatment assignments from the start. Allocating some cases to warm-up and some to the adaptive phase is a middle ground. The warm-up phase then serves to provide initial observations so that the prior at the start of the adaptive phase is partially data-driven. We recommend this middle-ground approach.
A further design choice is when to stop the trial early. In our second empirical illustration (below), there was little evidence of variation in
across context values
at the end of the warm-up and adaptive phase. As discussed below, in that example, we decided not to carry out a validation phase because the range of point estimates from the first two phases suggested insufficient variation to be of scientific interest. In general, we recommend that researchers after the warm-up and adaptive phase consider whether the likely results of the validation phase are of sufficient interest to justify the cost of collecting data in that phase.
We demonstrate our approach in two applied one with respondents choosing which of two fictitious immigrants is more worthy for admission to the country, and one with respondents choosing which of two fictitious job applicants to hire for a hypothetical job. We discover substantial heterogeneity in the first illustration and almost no heterogeneity in the second.
We conducted an online experiment based on Hainmueller and Hopkins (2015). We recruited 10,000 participants using Prolific, a web-based survey platform with a non-probability sample. The Prolific data collection took place between June 24, 2024 and June 30, 2024. We administered the survey with a custom Python Shiny App deployed via AWS. ^6^ Our implementation is standalone software that researchers can use to administer an experiment. It replicates the survey functionality of tools like Qualtrics, but enables continuous adaptive updating, which Qualtrics does not easily support.
The average age in the sample was 35 years old, 68% were white and 56% were female. We allocated 2,000 respondents to the Warm-up Phase (approximately 125 per each of 16 arms). We then recruited 6,000 respondents for the adaptive 3,000 for finding
and 3,000 for finding
. We allocated 2,000 respondents to the validation phase (1,000 for each
and
).
Similar to Hainmueller and Hopkins (2015), we asked each “Please read the descriptions of the potential immigrants carefully. Then, please indicate which of the two immigrants you would personally prefer to see admitted to the United States.” Respondents then viewed a table similar to Figure 4. Figure 3 enumerates the full set of 16 contexts.Figure 4Example of two fictional immigrant profiles shown to respondents. The two values of education, the focal attribute, are college degree versus no formal education. The values of the other context attributes are identical for both immigrants 1 and 2 (e.g., both from Eastern Europe), though each immigrant has a different signal of each attribute value (e.g., Germany and Poland).
Figure 5 presents the results at the end of the warm-up and adaptive phases. For each context
, we report the posterior mean probability
of preferring the college-educated immigrant over the immigrant with no formal education, along with a 95% Bayesian credible interval. Recall that a value of
corresponds to no causal effect of the focal attribute, and that the signal of a college degree increases the probability that a profile is chosen to the degree that
exceeds 0.5. A preference for the college-educated profile is apparent in all contexts. There is also clear heterogeneity across contexts, with estimates ranging from below 65% to above 75%.Figure 5Results after the warm-up and adaptive Estimated preference for the college-educated immigrant profile. The x-axis depicts the estimated probability of choosing the college-educated immigrant within each of those contexts, along with 95% credible intervals. The y-axis shows the full set of context attributes for all 16 contexts. The contexts highlighted in red are the contexts discovered as having the highest and lowest posterior probabilities of a respondent choosing the more-educated profile.
Two estimates in Figure 5 are more important than the the contexts highlighted in red with the maximum and minimum estimates. The Bayesian credible intervals for these contexts are the most narrow; over the course of the experiment, the adaptive phase came to allocate sample to these contexts with high probabilities. The chosen maximum arm has the highest posterior draw of
in 79% of posterior draws. The chosen minimum arm has the lowest in 49% of posterior draws. These quantities summarize our degree of confidence that these are the arms of interest. With a larger budget, we could have continued the experiment to become more certain of our selections.
The validation phase collected new independent data on these selected contexts. Figure 6 compares estimates from the (warm-up + adaptive) and (validation) phases. Estimates are very similar, though validation estimates are slightly closer to 0.5 as one might expect given the winner’s curse from the adaptive phase. Overall, this illustration demonstrates that our design can successfully discover contexts across which the effect of the focal attribute is substantially heterogeneous.Figure 6Validation Preference for the college-educated immigrant. The y-axis shows the contexts
and
that were identified in the adaptive experimental phase as having the highest and lowest posterior probabilities that the respondent would choose the more-educated prospective immigrant within the pair. The x-axis depicts the estimated probability of choosing the more educated immigrant within each of those contexts. Warm-up and adaptive estimates are posterior mean estimates with 95% credible intervals, and validation estimates are frequentist mean estimates with 95% confidence intervals. Figure 7Example of two fictional resumes shown to respondents. Both applicants have similar backgrounds and amount of experience. The two resumes differ by the signal of motherhood. In the resume on the left, the applicant is a member of the parent–teacher association while the applicant on the right volunteers with a neighborhood association. Figure 8Context attributes, values and Motherhood illustration. Analogous to Figure 3, each profile has a set of context all attributes other than the focal attribute. Our design adaptively randomizes a vector of attribute values across profile pairs; within the two profiles share identical attribute values. To ensure that the two profiles in the pair do not appear identical to each other, we randomly permute two signals of each value across the profiles. Our experimental design therefore requires the researcher to specify two signals for each attribute value. Appendix H explains how we chose these signals.
Our second illustration shows a setting in which we estimate precise effect sizes near zero in all contexts.
An extensive literature has documented a negative association between motherhood and women’s labor market outcomes (Budig and England 2001; Kleven, Landais, and Søgaard 2019; Lundberg and Rose 2000). One source of this disparity may be employer discrimination against mothers compared with childless women. Some studies randomize signals of motherhood in lab studies and real-world audits to assess labor market effects (Correll et al. 2007; Ishizuka 2021). We designed an online experiment to investigate how the effect of motherhood differs across contexts defined by two background the applicant’s race and the ranking of the applicant’s educational institution. We collected data on Prolific from February 8, 2024 to March 13, 2024. An important caveat to our results is that our evidence among online survey respondents may not generalize to the population of employers making actual hiring decisions.
Each respondent is randomized to a pair of job applicant resumes who have degrees in marketing and are applying for a human resources position. One resume signals motherhood while the other does not. Figure 7 presents an example pair of resumes. There are two context attributes (race and educational institution rank) and one focal attribute (motherhood). Figure 8 shows all context attribute values and signals. Appendix H discusses how we chose these signals drawing on prior research. For this illustration, the adaptive phase proceeded in batches of 200 participants with posterior probabilities updated between batches.
Figure 9 presents results from the warm-up and adaptive phases. We report the estimated values of
, or the posterior mean probability of preferring the non-mother job applicant over the equally qualified mother. Estimates for all four contexts are close to
, indicating minimal effects of motherhood regardless of context. If we were to select
and
, we would be highly the context with white job candidates from mid-ranked educational institutions context has a 39% posterior probability of being
, and the context with the black job candidates from highly ranked educational institutions has a 47% posterior probability of being
. Because we detected effectively no evidence that the focal attribute had any effect in any context, we stopped the study after the adaptive phase and did not carry out a validation phase.Figure 9Estimated preference for the non-mother job candidate. The y-axis shows the labels for each context. The x-axis shows the posterior probability within each of those contexts, along with 95% credible intervals. The estimates in this figure are all from the Warm-up + Adaptive phases of our experiment.
The motherhood illustration shows both limitations and strengths of the adaptive randomization design. Adaptive randomization discovers heterogeneous effects. In scenarios where little or no heterogeneity exists, our method will not discover heterogeneity. But this is also a strength. In a single-context experiment, critics may argue that a null finding would have been significant in a different context. But when an adaptive design indicates homogeneity, there is compelling evidence that none of the contexts have especially large effects.
This article introduces an adaptive experimental design for conjoint experiments. By adaptively adjusting randomization probabilities via Thompson sampling, our method efficiently identifies the contexts where the focal attribute has the most positive and most negative effects.
The adaptive randomization design complements the standard conjoint design. The standard design produces an additive summary of marginal attribute effects, or perhaps marginal interactions. The adaptive design searches for the particular contexts (e.g., work experience attribute values) at which the effect of the focal attribute (e.g., education) is most positive and most negative. Because of their distinct goals, the standard and adaptive designs have very different structures. While the standard conjoint design considers all attributes in parallel, the adaptive conjoint design defines one focal attribute (e.g., education) and explores its heterogeneous effects across multiple contexts (e.g., different values of work experience and country of origin). The adaptive design randomizes the context values by Thompson sampling and then randomly permutes signals of the focal attribute and the context value signals across two profiles.
The adaptive conjoint solves the problem of a moderately high-dimensional estimand in a different way from a standard conjoint. A standard conjoint solves the curse of dimensionality by scatter sample uniformly over a large space of attribute values and then summarize by estimands that marginalize over most attributes. An adaptive conjoint solves the curse of dimensionality by changing the randomization begin scattering uniformly but gradually concentrate the sample on two chosen contexts, ultimately gathering enough information to produce two context-specific estimates. The adaptive design works best in settings, where the dimensionality is large but still small enough that all contexts can be explored with data. In these settings, the adaptive design searches for heterogeneity at a we do not obtain well-powered estimates for all attributes. Instead of the additive average effects of all attributes, we estimate the effect of one focal attribute in two specific contexts discovered data-adaptively.
To create the adaptive conjoint design, this article makes three main (1) We conceptualize conjoint experiments in a new way, with one focal attribute and other context attributes. (2) We define two signals (e.g., Germany and Poland) for each context attribute value (e.g., Europe), so that profiles differ within one context. (3) We provide a three-phase framework for randomization in which the warm-up reduces sensitivity to the prior, the adaptive phase discovers the estimands of interest, and the validation phase yields valid statistical inferences. We illustrate the utility of the framework through two illustrations and simulated evidence (Appendix D).
A noteworthy tradeoff involves participant suspicion. A traditional conjoint infers potentially sensitive preferences from observed choices. It avoids social desirability bias because preferences emerge only in the aggregate; any individual choice could be attributed to any of several differing attributes (Hainmueller, Hangartner, and Yamamoto 2015; Hainmueller et al. 2014). Our design risks greater suspicion; profiles within a pair are similar except the focal attribute. We reduce suspicion in two ways. First, we construct two signals of each context so the pair are not identical. Second, while a traditional conjoint participant may complete several choice tasks (Hainmueller et al. 2014), we present only one task to each respondent. Researchers must weight these conceptual and statistical difficulties in our design against the benefit of discovering causal interactions through adaptive randomization.
Our approach speaks indirectly to the broader literature on factorial and multi-arm experiments (Kasy and Sautmann 2021; Offer-Westort et al. 2021; Villar et al. 2015). This literature often emphasizes the search for an arm that maximizes an unknown parameter or payoff. Researchers might also search for the arm that minimizes the value of that parameter. By searching for both, researchers can summarize the total range of parameter values. Our approach also illustrates how to address inferential challenges of adaptive experiments through a simple yet powerful carry out an adaptive phase to discover estimands of interest, followed by a validation phase to estimate the values of those estimands.
The adaptive conjoint design points to many areas for future research within the domain of conjoint experiments. First, there are open questions about the optimal allocation of sample size to the warm-up, adaptive and validation phases. Second, adaptive randomization could apply in real-world audit studies to discover heterogeneous discrimination by applicant attributes. A difficulty is the time the researcher submits a fictitious resume to a job and then waits a month to see if there is a callback. Future methodological research that improves the speed of data collection in audit studies is needed to make these designs adaptive. Third, future research could explore model-based strategies to pool information across contexts, thereby extending our approach from moderately-large settings with
contexts to truly high-dimensional settings with hundreds or thousands of contexts.