Authors: Yanyao Yi, Ying Zhang, Yu Du, Ting Ye
Categories: Article, causal inference, data fusion, integrative data analysis, sensitivity analysis, 62-08, 62D20, 62P10
Source: Journal of causal inference
Authors: Yanyao Yi, Ying Zhang, Yu Du, Ting Ye
Leveraging external controls – relevant individual patient data under control from external trials or real-world data – has the potential to reduce the cost of randomized controlled trials (RCTs) while increasing the proportion of trial patients given access to novel treatments. However, due to lack of randomization, RCT patients and external controls may differ with respect to covariates that may or may not have been measured. Hence, after controlling for measured covariates, for instance by matching, testing for treatment effect using external controls may still be subject to unmeasured biases. In this article, we propose a sensitivity analysis approach to quantify the magnitude of unmeasured bias that would be needed to alter the study conclusion that presumed no unmeasured biases are introduced by employing external controls. Whether leveraging external controls increases power or not depends on the interplay between sample sizes and the magnitude of treatment effect and unmeasured biases, which may be difficult to anticipate. This motivates a combined testing procedure that performs two highly correlated analyses, one with and one without external controls, with a small correction for multiple testing using the joint distribution of the two test statistics. The combined test provides a new method of sensitivity analysis designed for data fusion problems, which anchors at the unbiased analysis based on RCT only and spends a small proportion of the type I error to also test using the external controls. In this way, if leveraging external controls increases power, the power gain compared to the analysis based on RCT only can be substantial; if not, the power loss is small. The proposed method is evaluated in theory and power calculations, and applied to a real trial.
RCTs are the gold standard for generating high-quality causal evidence of new treatments and have long been recognized as the standard method to support key decisions in the drug development process [1,2]. However, despite its clear advantages, the traditional paradigm of conducting RCTs has been increasingly criticized for failing to meet contemporary needs. In certain settings, for example, in HIV prevention [3,4], oncology [5], and neurology [6], randomizing patients to placebo may be difficult for ethical or feasibility reasons. Moreover, adequately powered RCTs are becoming more and more impractical as a growing number of new treatments are targeted toward rare diseases or biomarker-defined subgroups of patients in the era of precision medicine [7].
Meanwhile, a plethora of real-world data (RWD) have been curated for administrative or research purposes and are becoming accessible to researchers in the form of disease registries, administrative claims databases, and electronic health records. These rich data sources can produce valuable insights, i.e., real-world evidence (RWE), into the effect of treatments in routine, daily practice. However, researchers almost ubiquitously caution against possible bias from unmeasured confounding when using RWD.
Being well aware of the limitations of using either RCT or RWD alone, the idea of using RWD to supplement RCT has gained growing interest in recent years. As forcefully argued in the study of Eichler et al. [7], “the future is not about RCTs vs RWE but RCTs and RWE.” There are numerous opportunities in how the integration of RCTs and RWD can achieve fruitful results that using either RCT or RWD alone cannot [8–10]. Among those, an important theme is on augmenting the RCT with RWD to increase efficiency [11–16], and particularly, constructing an externally augmented control arm in the analysis of RCTs [17–20]. Leveraging external controls – relevant individual patient data under control from external trials or RWD – has the potential to reduce the cost of RCTs while increasing the proportion of trial patients given access to novel treatments.
Using external controls is not an entirely new idea. Criteria for evaluating what constitute an acceptable external control arm are proposed in the study of Pocock [21]. It was discussed 20 years ago by the [22, E10 Section 2.5], and also recognized by European Medicines Agency [23], US Food and Drug Administration [24], and National Cancer Institute [25] as one direction to modernize clinical trials. In fact, properly selected external controls (e.g., using propensity score matching) have shown early promise, and several drugs have already been approved based on external control groups [26–28].
Using external controls typically requires the exchangeability condition, i.e., all patient characteristics that affect the potential outcome under control and differ between the trial population and the external control population are measured [29]. While careful adjustment for observed covariates can probably render the exchangeability assumption to hold approximately, the analysis may still be biased due to unmeasured covariates related to “difficulties in reliably selecting a comparable population because of potential changes in medical practice, lack of standardized diagnostic criteria or equivalent outcome measures, and variability in follow-up procedures” [24]. To reduce the potential biases from using external controls, an intuitive frequentist approach is “test-then-pool” that first tests for the comparability of the external controls and internal controls before leveraging external controls [20]. Bayesian methods that rely on power priors have also been popular, which use the likelihood of the external data to a specified power as the prior distribution [30,31]. As such, one can use power priors to adjust the weight allocated to the external information according to the levels of comparability between the external control and the internal data. However, these methods lack formal statistical theory on how the unmeasured biases might affect the validity and efficiency of the proposed procedures.
In this article, we take a different perspective to this problem and propose a sensitivity analysis approach to quantify the magnitude of unmeasured bias that would be needed to alter the study conclusion that presumed no unmeasured biases are introduced by employing external controls [32]. With the unbiased RCT-only test as the benchmark, leveraging external controls increases power or not depends on the interplay between sample sizes and the magnitude of treatment effect and unmeasured biases, which may be difficult to anticipate. This motivates a combined testing procedure that performs both tests, one with and one without external controls, correcting for multiple testing using the joint distribution of the two test statistics. Because the two tests are highly correlated, this correction for multiple testing is small. Interestingly, the proposed combined testing procedure can be viewed as a new method of sensitivity analysis designed for data fusion problems that anchors at the unbiased analysis based on RCT only and “spends” a small proportion of the type I error (i.e., the cost of multiple testing) to also test using the pooled controls. In this way, if leveraging external controls increases power, the power gain compared to the RCT-only test can be substantial; if not, the power loss is small. Before introducing technical details, it is useful to consider a motivating example.
Consider a non-inferiority, phase 3 RCT (referred to as the internal trial, ClinicalTrials.gov number, NCT01894568) comparing a new basal insulin, insulin peglispro, to insulin glargine as the control in Asian insulin-naïve patients with type-2 diabetes using a noninferiority margin of 0.4% [33]. The primary endpoint is the change in hemoglobin A1c (HbA1c) from baseline to 26 weeks of treatment. HbA1c is a continuous-valued measure of average blood glucose in the past 3 months. Before this trial, a phase 3 RCT of similar design (referred to as the external trial, ClinicalTrials.gov number, NCT01435616) has been conducted in the North America and Europe [34], whose control arm will be used as the source of external controls.
We focus on the overweight and obese population, which are, respectively, defined as 23≤ body mass index (BMI) <25 and BMI ≥ 25 for the internal trial according to the Asia-Pacific guidelines, and 25 ≤ BMI < 30 and BMI ≥ 30 for the external trial according to the World Health Organization classifications [35]. There are in total 159 patients under treatment and 150 patients under control in the internal RCT, and 486 patients under control in the external trial. We match 159 similar external controls to the 159 treated patients in the internal RCT using optimal matching based on a robust Mahalanobis distance and a caliper on the propensity score. See Rosenbaum [36, Part II] for discussion of these matching techniques. Table 1 describes covariate balance in the 159 matched pairs. All variables have standardized differences less than 0.13 and are considered sufficiently balanced [37].
Using only the internal RCT, 159 patients under treatment and 150 under control, we conduct a Z-test with the noninferiority margin of 0.4% and obtain a one-sided p-value 7.92 × 10^−7^. In this analysis, the evidence that the new insulin treatment is noninferior to insulin glargine is strong enough when only using the internal controls. On the other hand, under the exchangeability assumption, which implies that the 159 matched external controls are comparable to patients in the internal RCT, we construct an augmented control arm of 309 patients in total and obtain a one-sided p-value 1.88 × 10^−7^. Again, we find strong evidence of noninferiority; however, an investigator may be in doubt about the exchangeability assumption due to the influence of regions on the outcome. Then a natural question is could the one-sided p-value of 1.88 × 10^−7^ be due to regions rather than the effect of treatment? If the study conclusion from using external controls can be altered by a plausible effect of regions and because the RCT-only test is already powerful enough, the RCT-only test would be a better choice. However, it would be difficult to know this before examining the data. Motivated by the advice of performing multiple analyses with an appropriate correction for multiple testing given by Rosenbaum [38], we propose a combined testing procedure that performs both analyses, controlling for multiple testing using the joint distribution of the two test statistics. In this article, we will demonstrate that the combined test avoids making an inapt choice about whether to use external controls or not, and only has a small loss of power compared to knowing a priori which is the better choice.
Section 2 presents a test that uses only the internal controls and another test that also leverages the external controls, and discusses controlling type I error and comparing power without the exchangeability assumption. Section 3 proposes a combined test that performs both tests and studies in detail its statistical properties. Section 4 presents power calculations. Section 5 returns to the real data applications. Section 6 concludes with a discussion.
There is an RCT denoted as D=1. Let A be a binary treatment, where A=1 denotes treatment and A=0 denotes control, X a vector of observed baseline covariates, Y(a) the potential outcome under A=a, for a=0,1. Throughout the article, we assume consistency and Stable Unit Treatment Value Assumption (SUTVA) so that the observed outcome satisfies Y=AY(1)+(1-A)Y(0) [39]. Our estimand of interest is the average treatment effect in the RCT population θ⋆=EY(1)∣D=1-EY(0)∣D=1. In particular, we consider testing a one-sided H0:θ⋆=θ0versusHA:θ⋆>θ0. The other direction can be considered in the same way. Combining both one-sided tests and applying Bonferroni correction give a two-sided test [40, Section 4.2], and by inversion, a confidence interval.
Write the RCT sample as Yi,Xi,Ai,Di=1,i=1,…,nr, which is assumed to be independent and identically distributed according to the joint law of Y(1),Y(0),X,A∣D=1. Randomization in the RCT guarantees that A⊥Y(1),Y(0),X∣D=1 and P(A=a∣D=1)=πa>0 for a=0,1, with πa known and π0+π1=1. Let Y¯a and Sa2, respectively, be the sample mean and sample variance of the responses Yi’s from RCT subjects under treatment a, for a=0,1. Hence, the null hypothesis H0 can be tested using a simple Z-statistic: T1=Y¯1−Y¯0−θ0n1−1S12+n0−1S02, where n1 and n0 are, respectively, the number of RCT patients under treatment and control. Based on T1, we reject H0 when T1≥z1-α, where z1-α is the (1-α)th quantile of the standard normal distribution.
To supplement the RCT using external controls, one approach is to first extract external data for patients under control based on the inclusion/exclusion criteria of the RCT and then proceed by matching these external patients to the RCT patients based on their similarity in observed baseline information X [27]. Let D=0 denote the matched external controls, and thus D=0 implies A=0. Write the matched external controls as Yi,Xi,Ai=0,Di=0,i=1,…,ne, which is assumed to be independent and identically distributed according to the joint law of Y(0),X∣A=0,D=0. Suppose that matching has rendered the baseline observed covariates comparable between the RCT and external controls, i.e., D⊥X, and that these baseline covariates X explain all differences between the RCT and external controls, i.e., the exchangeability assumption D⊥Y(0)∣X holds. This implies D⊥Y(0),X and thus EY(0)∣D=1=EY(0)∣D=0. Let Y¯e be the sample mean of the responses Yi’s from the external controls, and wY¯0+(1-w)Y¯e be a weighted average of mean responses for the two control groups, where w∈[0,1] is a pre-specified weight, which could reflect the proportion of the internal control in the two control groups combined. Therefore, the null hypothesis H0 can also be tested borrowing information from the external controls using T2w=Y¯1-wY¯0+1-wY¯e-θ0n1-1S12+w2n0-1S02+(1-w)2ne-1Se2, where Se2 is the sample variance of the responses Yi’s from external controls. We make two remarks about T2(w). First, T2(w) is constructed assuming independence between the RCT and external controls, which means that T2(w) may be conservative due to correlation induced by matching [41] but usually to a small extent as the correlation is typically small [42]. Second, T2(w),w∈[0,1] defines a family of statistics that includes T2(1)=T1 as a special case. Among those, the exchangeability assumption implies the optimal w that maximizes the efficiency of T2(w) is proportional to the sample size, i.e., the optimal w equals nrπ0/nrπ0+ne. One can also choose different values of w to reflect the weights allocated to the two control groups.
The aforementioned approach of leveraging external controls relies on the exchangeability assumption, which may not hold because the RCT patients and external controls may differ with respect to covariates that may not have been measured. Without exchangeability, Y¯1-wY¯0+(1-w)Y¯e is not necessarily centered at θ0 under H0 and rejecting the null hypothesis when T2(w)≥z1-α may inflate type I error.
Define Δ⋆=EY(0)∣D=1-EY(0)∣D=0, which may be nonzero when exchangeability does not hold. This could occur, for example, if an important prognostic variable is unobserved and left uncontrolled, or if a variable that differs in distribution between D=0 and D=1 (such as region) cannot be matched. The correct rejection region for a size-α test based on T2(w) is T2w-1-wΔ⋆n1-1S12+w2n0-1S02+(1-w)2ne-1Se2>z1-α, which is infeasible because Δ⋆ is unknown. To deal with this issue, a tempting choice is to estimate Δ⋆ and “debiase” the numerator of T2(w) to make it mean zero. Nonetheless, estimating Δ⋆ introduces additional variation, which negatively affects the efficiency of the test. In particular, if one estimates Δ⋆ by Y¯0-Y¯e, the debiased numerator of T2(w) becomes Y¯1-Y¯0-θ0, and the resulting test statistic (after appropriately adjusting for its denominator to reflect the variability of the numerator) becomes equivalent to T1, the test statistic without using any external controls.
In order to borrow information from external controls while still controlling type I error, we consider departures from the exchangeability through the lens of a sensitivity analysis [36]. Specifically, we consider a sensitivity parameter Δ0 such that it bounds the magnitude of bias Δ*, i.e., Δ0≥Δ*. Define T2,Δ0w=Y¯1−wY¯0+1−wY¯e−θ0−1−wΔ0nrπ1−1S12+w2nrπ0−1S02+(1−w)2ne−1Se2. Because Δ⋆≤Δ0, the reject region T2,Δ0(w)≥z1-α controls type I error at level α. As a special case when Δ⋆≤Δ0 holds with Δ0=0 (e.g., under exchangeability), T2,Δ0(w)≥z1-α becomes T2(w)≥z1-α, the reject region under exchangeability. As Δ0 increases, there is greater uncertainty about how the exchangeability might be violated, leading to more stringent rejection criterion to control type I error. The reject region T2,Δ0(w)≥z1-α is sharp under Δ⋆≤Δ0 in the sense that they are of size-α when Δ⋆=Δ0, so it cannot be improved unless further information is provided.
Write σa2=VarY(a)∣D=1), for a=0,1, and σe2=VarY(0)∣D=0. Under the alternative hypothesis HA:θ⋆>θ0, the power of T1 is the probability of event T1≥z1-α, which is asymptotically equal to (1)1-Φz1-α+nrθ0-θ⋆π1-1σ12+π0-1σ02, where Φ(⋅) is the standard normal cumulative distribution. In parallel, the power of T2,Δ0(w) is the probability of event T2,Δ0(w)≥Z1-α, which is asymptotically equal to (2)1-Φz1-α+nrθ0-θ*+nr1-wΔ0-Δ*π1-1σ12+w2π0-1σ02+(1-w)2nrne-1σe2.
Several remarks are in order based on the above power formulas. First, the power of T2,Δ0(w) is larger than that of T1 if and only if θ0-θ⋆+(1-w)Δ0-Δ⋆π1-1σ12+w2π0-1σ02+(1-w)2nrne-1σe2≤θ0-θ⋆π1-1σ12+π0-1σ02. For instance, when Δ0=Δ⋆, i.e., the specified upper bound for Δ⋆ is tight, and σ02=σe2, i.e., the variance of Y for the two control groups are equal, simple algebra reveals that the power of T2,Δ0(w) is always larger than that of T1 for any w satisfying max0,nrπ0-ne/nrπ0+ne≤w<1.
Second, we can derive the oracle w that maximizes the power of T2,Δ0(w). Let κ=π0-1σ02/π1-1σ12+π0-1σ02, the optimal w takes the following (3)wopt=1,whenΔ0-Δ⋆≥κθ⋆-θ0>0,1-Δ0-Δ⋆π1-1σ12+π0-1σ02+θ0-θ⋆π0-1σ02θ0-θ⋆nrne-1σe2+π0-1σ02+Δ0-Δ⋆π0-1σ02,whenκθ⋆-θ0>Δ0-Δ⋆≥0, where the first case is when Δ0 is specified too large, the power of T2,Δ0(w) is maximized at w=1, which means that using the external controls does not lead to efficiency gain. As an illustration, under the special case that π1-1σ12=π0-1σ02=nrne-1σe2, when Δ0-Δ⋆>θ⋆-θ0/2, the optimal w is 1, whereas when θ⋆-θ0/2>Δ0-Δ⋆>0, the optimal w is 1-θ0-θ⋆+2Δ0-Δ⋆/2θ0-θ⋆+Δ0-Δ⋆. Under another special case when Δ*=Δ0 and σ1=σ0=σe, wopt becomes nrπ0/nrπ0+ne, which agrees with the optimal w under exchangeability discussed in Section 2.1. The proof of (3) is given in the supplementary material.
Finally, we compare the two tests T1 and T2,Δ0(w) in terms of their limiting power as the sample sizes grow to infinity. When θ⋆>θ0 and limnr→+∞nrθ⋆-θ0=+∞ (e.g., when θ⋆,θ0 are two constants), then the power of T1 goes to 1 as nr→∞. In contrast, the limiting power of T2,Δ0(w) depends on specifications of w and Δ0. Specifically, as can be seen from the power formula in (2), when minnr,ne→+∞, the power of T2,Δ0(w) tends to 1 if θ0-θ*+(1-w)Δ0-Δ*<0, i.e., when Δ0<θ*-θ0/(1-w)+Δ*, and to 0 if θ0-θ*+(1-w)Δ0-Δ*>0, i.e., when Δ0>θ⋆-θ0/(1-w)+Δ⋆. In other words, there exists a w-dependent number Δ~(w)=θ⋆-θ0/(1-w)+Δ⋆ that characterizes the limiting behavior of T2,Δ0(w) under the the power of T2,Δ0(w) tends to 1 if Δ0<Δ~(w) and to 0 if Δ0>Δ~(w) as minnr,ne→+∞. In plain language, since the maximum bias Δ0 counteracts with θ⋆-θ0, the difference we would like to detect, when Δ0 exceeds a certain threshold Δ~(w), the maximum bias starts to dominate the true difference θ⋆-θ0, resulting in no power to detect the difference. This number Δ~(w) is analogous to the design sensitivity in the literature of sensitivity analysis [36,43].
Should we leverage external controls? In other words, is it better to use the test statistic T1 constructed solely based on the RCT or the test statistic T2,Δ0(w) that additionally leverages the external controls? We know from the above theory and analysis that the answer to this question depends upon the context, specifically upon the nature and size of the treatment effect, and the specification of w and Δ0, that might be difficult to anticipate prior to examining the data. As Motivated in Section 1, we propose a combined testing procedure that performs both T1 and T2,Δ0(w), correcting for multiple testing using the joint distribution of the two test statistics.
Under H0, the joint distribution of (T1,T2,Δ*(w)) is asymptotically bivariate normal, satisfying T1T2,Δ*(w)→dN00,1ρρ1, where ρ=π1-1σ12+wπ0-1σ02π1-1σ12+π0-1σ02π1-1σ12+w2π0-1σ02+(1-w)2nrne-1σe2. Again, for illustration, consider the special case that π1-1σ12=π0-1σ02=nrne-1σe2, then ρ increases as w increases from 0 to 1, and thus ρ ranges between 0.5 and 1.
Consider the testing procedure that, for any specified Δ0 and w, rejects H0 if (4)maxT1,T2,Δ0(w)≥c1-α;ρ, where c1-a;ρ satisfies Φ2,ρc1-a;ρ=1-a,Φ2,ρ(x,y) is the probability of the two-dimensional lower orthant (-∞,x]×(-∞,y] for a bivariate normal distribution with expectation (0,0)T, unit variances, and correlation coefficient ρ, and write Φ2,ρ(x)=Φ2,ρ(x,x). This combined testing procedure is able to control the type I error for any Δ*∈-∞,Δ0 because PH0maxT1,T2,Δ0(w)≥c1-a;ρ≤PH0,Δ*=Δ0maxT1,T2,Δ*w≥c1-a;ρ=α.
In what follows, we establish several attractive features of the combined test. Note that under the alternative hypothesis, the power of the combined test – the probability of event (4) – is
(5)PmaxT1,T2,Δ0(w)≥c1-α;ρ≈1-Φ2,ρc1-α;ρ+nrθ0-θπ1-1σ12+π0-1σ02⏟B1,c1-α;ρ+nrθ0-θ+nr1-wΔ0-Δ*π1-1σ12+w2π0-1σ02+(1-w)2nrne-1σe2⏟B2,
where ≈ means asymptotic approximation. This leads to the first observation that the power of the combined test is generally larger than the worst of the two component tests, i.e., Powerc≥min(Power1,Power2), where Power1, Power2, Powerc are, respectively, the asymptotic power of T1,T2,Δ0(w), and the combined test. This can be seen from noting that
1-Powerc=Φ2,ρc1-α;ρ+B1,c1-α;ρ+B2≤Φ2,ρc1-α;ρ+minB1,B2,+∞=Φc1-α;ρ+minB1,B2=Φz1-α+maxB1,B2-B1-B2-c1-α;ρ-z1-α≤Φz1-α+maxB1,B2=maxΦz1-α+B1,Φz1-α+B2=1-minPower1,Power2,
where the second inequality holds when B1-B2≥c1-α;ρ-z1-α, i.e., when the power of the two component tests are not too similar.
Moreover, not only is the power of the combined test better than the worst of the two component tests in finite sample, it is also close to the better of the two component tests in finite sample, and equal to the better of the two component tests in the limit. To see this, we bound the difference in power as maxPower1,Power2-Powerc=Φ2,ρc1-α;ρ+B1,c1-α;ρ+B2-Φz1-α+minB1,B2≤Φc1-α;ρ+minB1,B2-Φz1-α+minB1,B2≤1-2Φz1-α-c1-α;ρ/2. It is helpful to anchor several values of c1-α;ρ and the upper bound 1-2Φz1-α-c1-α;ρ/2 in terms of different α and ρ. When α=0.025 and for ρ=0.5,0.7,1, the critical values are c1-a;0.5=2.21,c1-a;0.7= 2.18, c1-α;1=1.96, and respectively, the upper bounds are 0.100, 0.088, 0. This means that because of the high correlation between T1 and T2,Δ*(w), the price paid for multiple testing is generally small. With regard to the limiting power, it is also easy to see that for fixed θ0 and θ⋆>θ0,B1→-∞ as the sample size nr increases. Hence, the combined test always has its power approaching 1 as nr→∞, just like the test T1 that only uses RCT data, which is not the case for T2,Δ*(w) as discussed in Section 2.3. This further shows the advantage of the combined test.
For implementation of the sensitivity analysis (either T2,Δ0 or the combined test), practitioners are not required to specify the value of the sensitivity parameter Δ0. Following the pioneering work by Cornfield et al. [44] and the sensitivity analysis literature [36], results from the combined test can be summarized by the “tipping point” – the magnitude of Δ0 that would be needed such that the null hypothesis can no longer be rejected. If such a value of Δ0 is deemed implausible, then we still have evidence to reject the null hypothesis based on the combined test. In Section 5, we illustrate the method using a real example.
We investigate three factors when conducting power calculations. The first factor concerns the true treatment effect θ⋆=0.2,0.3, and 0.4. The second factor is the specified value of the maximum bias Δ0=0.2,0.3,0.4,0.6. The third factor is the sample size n1=50,100,150,200, with n1:n0:ne=2:1:3. Additional parameters are θ0=0,Δ⋆=0.2,σ1=σ0=σe=1, and α=0.025.
Table 2 summarizes the power of T1,T2,Δ0(w) and the combined test Tc,Δ0(w)=maxT1,T2,Δ0(w), calculated, respectively, using (1), (2), and (5). For T2,Δ0(w) and Tc,Δ0(w), we consider two choices of w: the oracle w in (3) that maximizes the power (denoted as wopt), and its value under exchangeability n0/n0+ne=1/4. In the supplementary material, we check powers by simulation, finding good agreement. In the supplementary material, we also include a check of the type I error, which are all close to or below the nominal level, indicating validity of all the tests. In contrast, a naive combined test without correcting for multiple testing cannot control the type I error.
The following is a summary of results in Table 2.
We revisit the example introduced in Section 1.2 and illustrate how the proposed methods can be applied. Formally, we test the hypothesis that H0:θ⋆=θ0 versus HA:θ⋆<θ0, with θ0=0.4, which can be equivalently implemented using the tests described in Sections 2 and 3 with Yi’s replaced by -Yi’s and θ0 replaced by -θ0. We set the significance level α=0.025.
Using only the internal RCT, T1=4.80 with p-value 7.92 × 10^−7^, based on which we reject the null hypothesis H0. This result is solely based on internal controls and thus is invariant to the value of Δ0.
Leveraging external controls and let w=n0/n0+ne=0.485,T2(w)=5.08 with p-value 1.88 × 10^−7^ when Δ0=0. Therefore, under the exchangeability assumption, we can also reject the null hypothesis H0. To gauge the robustness of this conclusion to violation of the exchangeability, we apply the proposed sensitivity analysis. As discussed at the end of Section 3, results of our sensitivity analysis can be summarized by the “tipping point” – the magnitude of Δ0 that would be needed such that the null hypothesis can no longer be rejected. In this example, as Δ0 increases, the adjusted p-value associated with T2,Δ0(w) increases but remains below α=0.025 for any Δ0≤0.62. Namely, two patients with the same observed characteristics (as listed in Table 1), one in the internal RCT and the other in the external trial, may differ in their expected potential outcome under control by up to 0.62, under which the adjusted p-value is still below the significance level α. This means that the significant effect we observe cannot be explained away by unmeasured biases of magnitude up to Δ0=0.62. If such a large unmeasured bias is deemed implausible, then there is no real doubt that the rejection based on T2,Δ0 provides evidence of noninferiority.
Finally, using the combined test, maxT1,T2,Δ0(w)=5.08 with adjusted p-value 3.41 × 10^−7^ when Δ0=0. As Δ0 increases, the adjusted p-value for the combined test increases but plateaus at 1.41 × 10^−6^ when T1≥T2,Δ0(w). This means that rejection based on the combined test is insensitive to any value of Δ0, i.e., similar to T1 that only uses the internal RCT, rejection based on the combined test is insensitive to any violation of the exchangeability assumption.
It is also interesting to see the relative performance of T1,T2,Δ0(w),Tc,Δ0(w) when the internal RCT is underpowered, and thus the combined test may be more useful. For this purpose, we randomly sample with replacement 100 patients from the internal RCT, with a target ratio of 4/5 from the treated arm and 1/5 from the control arm. Then T1 is computed using this subsample from the RCT, while T2,Δ0(w) and Tc,Δ0(w) additionally use the external controls that were matched to the sampled treated patients with w=n0/n0+ne calculated using the subsample. This procedure is repeated 1,000 times. Among these repetitions, T1 rejects the null hypothesis 71.5% of the time, i.e., the power of T1 is 71.5%, while the combined test Tc,Δ0(w) has power 82.4%, 74.2%, 71.1% when Δ0=0.1,0.2,0.25, respectively. Hence, the sensitivity parameter Δ0 can be as large as 0.25 before the combined test starts to lose power compared to T1. In comparison, T2,Δ0(w) has worse performance, with power equal to 80.4%, 64.8%, 45.7% when Δ0=0.1,0.2,0.25, respectively. Taking a closer look at the results, we note that if T1 is larger than c1-α;ρ defined in (4), then both T1 and the combined test can reject H0 regardless of the value of Δ0. If T1<z1-α, then T1 cannot reject H0 while the combined test can still reject 27.7% of these cases at Δ0=0.2. The potential loss of using the combined test is when T1 is between z1-α and c1-α;ρ, in which cases using T1 alone can reject H0 but the combined test is sensitive to a certain value of Δ0. However, this scenario is relatively rare and occurs in 8.4% of the repetitions; furthermore, even in this scenario, the combined test can still reject H0 at Δ0=0.2 around half the time.
The last step of a sensitivity analysis is to reason about whether a value of Δ0=0.2 is plausible given that we have already controlled for baseline covariates listed in Table 1. For this task, an intuitive strategy is to judge the plausibility of Δ0 in reference to some observed covariates [45]. Specifically, we can omit observed covariates one at a time during matching and calculate Y¯0-Y¯e using the resulting matched external controls. Using this procedure, we estimate the amount of bias from not matching on one of the observed covariates and to benchmark the plausibility of Δ0, the amount of bias from not being able to match on the region variable. The results show that omitting the baseline HbA1c leads to the largest Y¯0-Y¯e that is equal to 0.14, while omitting any other observed variables in Table 1 leads to Y¯0-Y¯e ranging from −0.05 to 0.04. Based on the prior knowledge in the study of Home et al. [46] that the baseline HbA1c explains most of the variability in the change in HbA1c, particularly in comparison to the geographical region, we view that Δ=0.2 is implausible.
In summary, before looking at the data, the choice between T1 and T2,Δ0(w), would be difficult to make or justify on the basis of a priori considerations. In some cases, T1 may not be powerful enough due to the small sample size of the internal RCT, while leveraging external controls leads to a more powerful test. In some other cases, T2,Δ0(w) may be sensitive to unmeasured biases while T1 is already powerful enough. Under these circumstances, the combined test Tc,Δ0(w) is often preferable as it performs both tests with a small correction for multiple testing by taking into account the high correlation of the two test statistics.
We propose a sensitivity analysis approach for using external controls in clinical trials to examine the robustness of study conclusion to remaining unmeasured bias after controlling for measured covariates. Results from the sensitivity analysis can be summarized by the “tipping point” – the magnitude of Δ0 that would be needed such that the null hypothesis can no longer be rejected. If Δ0 is deemed plausible (or implausible), the conclusion based on using external controls is sensitive (or robust) to unmeasured bias.
When in doubt about whether the use of external controls increases power, we propose a combined testing procedure that performs both tests, one only using the internal controls and one additionally using the external controls, correcting for multiple testing using the joint distribution of the two test statistics. Because the two test statistics are highly correlated, this correction for multiple testing is small, and thus the combined test only has a small loss of power compared to knowing a priori which test is best. Moreover, the combined test provides a new method of sensitivity analysis designed for data fusion problems, which anchors at the unbiased RCT-only analysis and spends a small proportion of the type I error to also test using the external controls. In this way, if leveraging external controls increases power, the power gain compared to the RCT-only analysis can be substantial; if not, the power loss is small.
Our work is motivated by the literature of sensitivity analysis, in which testing a hypothesis multiple times has been shown to be useful in enhancing the robustness to unmeasured bias [38,47–49]. Nonetheless, we focus on a distinct context and have shown that testing multiple times using both a known unbiased test and potentially biased tests can be particularly attractive for data fusion problems. We also have developed various properties of the combined procedure that has not appeared in the existing literature.
Finally, a remaining question is how to choose w for the combined test. The power of the combined test depends on w in a complicated way as w not only affects the definition of T2,Δ0(w) but also the correlation ρ, which makes finding the optimal w a cumbersome task. In practice, a reasonable choice is w=π0nr/ne+π0nr, which minimizes the variance of wY¯0+(1-w)Y¯e when VarY(0)∣D=1=VarY(0)∣D=0. Another way is to pre-specify several values of w, calculate the corresponding test statistics, and combine all the test statistics using their joint null distribution. Because of the high correlation between these test statistics, the price paid for multiple testing will generally be small.