Authors: Adina Harri, Laurence S. Magder, Patricia D. Franklin, Jay S. Magaziner, Laura M. Kernan, Carol A. Lambourne, Vincent D. Pellegrini, Jr
Categories: Methods, Pilots, and Protocols
Source: JBJS Open Access
Authors: Adina Harri, Laurence S. Magder, Patricia D. Franklin, Jay S. Magaziner, Laura M. Kernan, Carol A. Lambourne, Vincent D. Pellegrini
Pragmatic trials have gained popularity in recent years because of their applicability to practical clinical situations. In contrast to traditional explanatory trials, which are tightly controlled and designed to define the effects of treatments under ideal circumstances, pragmatic trials are intended to reveal differences between established treatments in real-world situations. Although the intention-to-treat principle remains the cornerstone of explanatory trials, it may not reliably identify the treatment effect of greatest relevance in pragmatic trials. The estimand approach to trial design and analysis provides for specification and handling of various important intercurrent events that characterize pragmatic trials and arguably allows clearer definition of the treatment effects of interest for assorted real-world populations. Since pragmatic trials have considerable relevance to orthopaedics, we share the rationale for design and upcoming analysis of the Pulmonary Embolism Prevention after Hip and Knee Replacement Trial using an estimand framework.
Comparative effectiveness research (CER) trials, also known as “pragmatic trials,” are typically designed to compare treatment effects of existing guideline-approved interventions within a broad population in a real-world setting. Since its creation in 2010 and Congressional reauthorization in 2019, the Patient-Centered Outcomes Research Institute (PCORI) has been an advocate and leading funder of patient-centered comparative clinical effectiveness trials in the United States^1^. Such comparative effectiveness trials comprise the vast majority of clinical trials in orthopaedics, where the objective is to ascertain relative benefits and harms of existing treatment alternatives that involve medications, procedures, and/or implantable devices. By contrast, traditional “explanatory” trials are designed to demonstrate whether a new drug or intervention exhibits a treatment effect under ideal conditions^2^. To isolate and analyze this causal relationship, the efficacy trial tightly controls study population inclusion and exclusion criteria, intervention dose, and protocol adherence. Thus, generalizability to clinical practice is compromised because actual clinical settings, real patients, and intervention adherence are rarely so controlled. Indeed, specific criteria have been proposed to facilitate the distinction between efficacy (explanatory) and effectiveness (pragmatic) trials^3^. Hence, the primary strength of CER trials—their broad generalizability to practical clinical circumstances—is also their major weakness because they are, by definition, conducted in loosely controlled real-life environments.
In 2021, in an effort to clarify specific treatment effects addressed by clinical trials, concurrent with the increasing popularity of comparative effectiveness trials, the International Committee for Harmonization of Technical Requirements for Pharmaceuticals for Human Use (ICH) underscored the importance of estimands. They questioned “whether estimating an effect in accordance with the Intention to Treat principle always represents the treatment effect of greatest relevance to regulatory and clinical decision-making”^4^. This issue is of particular importance when considering comparative effectiveness trials. The distinction between CER and explanatory trials largely boils down to differences in treatment effects (i.e. “estimands”) that are targets of estimation and, importantly, how intercurrent events are reflected in the clinical question of interest^3^. The ICH suggests the first step in planning a clinical trial is to precisely define the treatment effects, or causal “estimands,” to be estimated.^4–6^ The abundance of real-world factors resulting in intercurrent events in CER trials presents a plethora of estimands; as such, the estimand strategy for targeting treatment effects in varied populations is particularly suited to CER trials and the ICH outlines a specific framework for defining estimands and their analytic implications. To demonstrate the utility of the estimand framework and its applicability to the universe of CER clinical trials, we share our rationale for the upcoming analysis of the Pulmonary Embolism Prevention after Hip and Knee Replacement (PEPPER) CER Trial: Balancing Safety and Effectiveness (NCT02810704, 11/05/16).
The PEPPER trial is a multicenter comparative effectiveness trial funded by PCORI and designed to study the effectiveness and safety of 3 guideline-approved agents in patients undergoing primary or revision total hip or knee replacement as prophylaxis for venous thromboembolism (VTE)^7^. Nearly 19,000 patients scheduled for hip or knee replacement at 32 centers across the United States and Canada were randomized between December 2016 and November 2024 to receive aspirin, warfarin, or rivaroxaban for 4 weeks and followed for 6 months. Primary effectiveness outcomes are all-cause mortality plus clinical VTE (deep vein thrombosis [DVT] and pulmonary embolism [PE]), as well as patient-reported functional outcomes; primary safety outcome is bleeding, inclusive of major, clinically important nonmajor, and wound-related bleeding. With both effectiveness and safety event rates of approximately 2%, the trial was powered to study both effectiveness and safety outcomes. All patients at least 21 years of age undergoing either primary or revision total hip or knee arthroplasty were eligible for the trial and were allowed enrollment for only 1 eligible procedure. Patients with a medical contraindication to one of the 3 study regimens were only eligible for randomization to the other 2 regimens; those with a personal history of prior PE were ineligible to receive aspirin and were randomized to only warfarin or rivaroxaban, and those on chronic preoperative anticoagulation were excluded from the trial. Randomization to warfarin ended in May 2023, with 5,495 patients enrolled, due to challenges of INR monitoring post-COVID. Randomization to both aspirin and rivaroxaban arms attained enrollment targets on November 30, 2024, with 6,603 and 6,785 patients, respectively. Final primary and secondary outcomes data collection and analyses are pending. Study drug costs were not funded by the trial, and patient out-of-pocket drug costs, most often for rivaroxaban, were variably subsidized at some US sites by availability of a 340B program. Similarly, INR monitoring of warfarin was often unpredictably not available at some enrolling sites. These 2 issues were the most common factors unrelated to VTE risk that affected both medication adherence and cross-over between study drugs.
Estimands represent true causal effects of treatments studied in a population of interest^8^. The causal estimand of one exposure vs. another (treatment “A” vs. treatment “B”) is a summary function of what would happen if everyone in the population was exposed to treatment A vs treatment B. Because CER trials cannot study the entire population with both treatments, statistical inference is used to estimate effects on the entire population derived from a representative study sample. By assuming random assignment to the 2 treatments, we derive unbiased estimates of the causal estimand of an intervention in a specific population. To specify an estimand of interest, one must define the a) population, b) treatments to be compared, c) clinical outcome (endpoint of interest), and d) summary measure to be used^9^ (Fig. 1).

The choice of estimand will influence the study design and analytic strategy. Traditional efficacy trial estimands are straight-forward and require few assumptions because randomization in a predefined population minimizes confounding. However, intercurrent events (e.g., postrandomization events such as a change in treatment due to clinical condition or inability to pay for one treatment) compromise randomization in any type of trial. In CER trials where patients with intercurrent events are included, more complex compensatory methods and/or assumptions are needed to obtain valid estimates of the specified estimands. The ICH guidelines identify strategies that can be used to approximate valid estimates of causal estimands in the presence of such confounding intercurrent events.
By way of example, we define several causal estimands that will be targeted in the PEPPER trial, a large pragmatic trial that includes diverse patients, clinical settings, and postrandomization intercurrent events encountered in clinical practice. We discuss how we plan to estimate them, including the handling of intercurrent (postrandomization) issues and the assumptions required for unbiased estimates. We will focus on the comparison of aspirin with rivaroxaban (Table I), but analogous estimands in our analysis will compare aspirin with warfarin, rivaroxaban with warfarin, and warfarin or rivaroxaban with aspirin. For simplicity in this illustration, we focus on one primary the composite of death and clinical VTE.
The effect of a clinical guideline or policy recommending aspirin or rivaroxaban among those undergoing hip or knee replacement.
The reference population consists of Americans and Canadians who satisfy inclusion criteria; i.e., individuals over 21 years of age undergoing total hip or knee replacement and clinically eligible to receive either aspirin or rivaroxaban.
Treatments to be Being recommended to receive aspirin vs rivaroxaban. Note the causal effect of a treatment guideline is not the same as the effect of a treatment because not everyone will be prescribed the recommended treatment and those prescribed might not take it. In PEPPER, some participants randomized to receive rivaroxaban were prescribed or independently elected to take aspirin due to the high copay for rivaroxaban required by some insurance companies. Thus, despite recommendations, drug costs can influence protocol adherence independent of any risk of experiencing the specified outcome.
All-cause mortality and/or clinical PE or DVT, confirmed by imaging and/or resulting in readmission within 182 days of discharge.
Six-month risk of the clinical outcome between treatment groups.
Analysis plan and handling of intercurrent Only those medically eligible to receive either aspirin or rivaroxaban will be included in the analysis. We will compare patients randomized to aspirin with those randomized to rivaroxaban using a Kaplan-Meier approach to estimate 6-month risk. The analysis will include all patients in each randomization group except a small number who ultimately did not undergo surgery since they no longer meet eligibility criteria of having experienced the event associated with increased VTE risk. Patients lost to follow-up will be censored at the time of their last follow-up. This “modified intention to treat” analysis uses a “treatment policy” (intended behavior) approach to the intercurrent event of cross-over and a “composite outcome” approach to the intercurrent event of death^4–6^.
Assumptions required for unbiased estimation of the Being assigned to receive rivaroxaban (or aspirin) will result in similar uptake as if there was a recommendation to treat with rivaroxaban (or aspirin). Participants at clinical sites who agree to participate are an unbiased representation of the population, there is no effect of treatment assignment on the decision to cancel surgery, those lost to follow-up have the same risk as those completing follow-up, and there is an unbiased ascertainment of patient clinical outcomes. Unlike subsequent estimands, due to randomization, this estimand does not require adjustment for potential group differences.
The effect of being treated in hospital and sent home with a prescription for aspirin vs. rivaroxaban, among those undergoing hip or knee replacement.
For this estimand, aspirin treatment corresponds to receiving aspirin in the hospital (day of surgery and thereafter) and rivaroxaban treatment corresponds to receiving rivaroxaban in the hospital (starting approximately 24 hours after surgery and daily thereafter) and each group is discharged with a 30-day prescription for the same drug. The effect of being sent home with a prescription is not the same as being treated because, for various reasons pertinent to a pragmatic trial such as medication cost, side effects, or noncompliance, some prescribed a treatment will not take it after discharge.
Analysis plan and handling of intercurrent The Kaplan-Meier approach will be used to estimate 6-month risk in each group. Possible confounding events must be addressed in refining the 2 comparison groups. Patients who never received the randomized treatment in hospital will be excluded; the minimum requirement is to have been treated in hospital with the randomly assigned drug. Those prescribed a different treatment on discharge will be included but censored at the time of discharge. Because some participants will be excluded or censored because of failure to receive the randomized treatment in the hospital, treatment groups may be imbalanced about prognostic variables, introducing possible confounding. These imbalances will be adjusted using weights in the Kaplan-Meier analysis^10^. Candidate variables for weighting (i.e. adjustment) will include age, sex, clinical site, date of surgery, education, and socioeconomic status. This estimand represents a real-world patient cohort after discharge and most closely affords generalizability of study results to clinical practice. As such, it is the primary analysis for the PEPPER Trial and represents a “modified per-protocol” analysis.
Assumptions required for unbiased estimation of the In addition to the same assumptions as for Estimand 1, this analysis requires the assumption that participants discharged with a different treatment than that randomly assigned did not have a different risk for the primary outcome of death or clinical VTE than those discharged with the randomly assigned treatment, conditional on the variables adjusted with the weights.
The effect of actually receiving treatment with aspirin vs. rivaroxaban as defined by the study protocol among those undergoing hip or knee replacement.
Treatments to be For this estimand, being treated with aspirin (or rivaroxaban) corresponds to receiving aspirin (or rivaroxaban) per protocol in the hospital, being discharged with aspirin (or rivaroxaban), and taking that medication at home.
Analysis plan and handling of intercurrent The Kaplan-Meier approach is used to estimate 6-month risk in each group. All participants who receive one of the study treatments in-hospital and at discharge will be included, regardless of whether they were randomly assigned or crossed-over to the treatment. Those who switch treatments will be censored from analysis when they change treatments such that all included in the analysis will have taken only a single treatment as defined by the protocol for the analysis period. Because some participants will be excluded or censored because of cross-over, groups may be imbalanced about prognostic variables, resulting in confounding. These group-specific imbalances will be adjusted using the same weights as for Estimand 2. This most closely resembles a traditional “as-treated” analysis as seen with an observational study.
Assumptions required for unbiased estimation of the These analyses require the same assumptions as Estimand 1. Additional assumptions include that participants who switched from randomly assigned treatments did so for reasons unrelated to their primary outcome risk, did not have a different primary outcome prognosis than those who did not switch treatments, and the same Estimand 2 assumptions (noninformative censoring) and conditional on all variables being adjusted with weights.
Beyond providing a framework for comparisons between different treatments, the estimand approach can be used for within-group analyses based on duration of treatment on the protocol from 1 to 4 weeks. This would be of interest to groups formulating clinical practice guidelines specifying durations of prophylaxis based on related differences in outcome risk.
Defining the population of interest, treatments to be compared, and important clinical outcomes are integral to defining estimands, which should be identified in advance of conducting a pragmatic clinical trial. Once specific estimands are identified, the analytic plan for comparative outcomes can be designed. Furthermore, critical assumptions required for unbiased estimation of treatment effects, which may be compromised by intercurrent events common in pragmatic trials, must be carefully assessed. Our estimands illustrate an important tension between explanatory and pragmatic (CER) estimating the “pure” physiologic effect of a treatment intervention in a tightly controlled explanatory trial vs. the real-world effect of being recommended, prescribed, or receiving that intervention in a pragmatic trial such as PEPPER.
Conventional discourse around analytic strategies for pragmatic CER trials focuses on whether the methodology should follow an “intention to treat” versus a “per protocol” framework^11-17^. Such a binary approach affords only a limited number of interpretations relevant to a limited number of clinical conditions. By contrast, an estimand framework provides greater precision and flexibility in specifying clinical trial goals when uncontrolled confounding “real-world” events are more likely to be encountered. Different clinical questions beget different causal estimands, and each question is of interest to different stakeholders, such as patients, clinicians, hospitals, payers, professional advisory groups, and industry.^9,18^ Considering these stakeholders facilitates choosing appropriate causal estimands and congruent analytical methods.
Potential limitations of pragmatic (CER) trials derive from the critical assumptions that must be relied on to develop valid analyses for interpretation of results. Accordingly, we list the assumptions and adjustments required for unbiased estimation of each estimand. Although such adjustments constitute a standard approach to outcome analysis, one must acknowledge that unmeasured patient factors associated with the outcome may bias interpretation of results. If crossover from rivaroxaban to aspirin is associated with unidentified factors that were neither measured nor controlled, the analyses may inadvertently be biased. Such concerns are minimized by the randomization and tight controls inherent in an explanatory clinical trial.
The estimand framework is only recently endorsed by the ICH and is making its way into publications of clinical trials in other fields. Recent reports of clinical orthopaedic trial results typically did not specify the causal estimands being targeted. Especially in comparative effectiveness trials, there are multiple causal effects of interest (such as the effect of being prescribed vs. actually taking a treatment), and it may not be obvious which is the target of a study. It is anticipated that with increasing investigator awareness of the ICH guidelines and growing popularity of CER, the estimand framework for clinical trial design and analysis will become more commonplace in orthopaedics. Explicit specification of target causal estimand(s) helps investigators design a study and identify assumptions required for valid analysis while simultaneously guiding readers to proper interpretation of results.
In conclusion, the estimand approach to formulating an analytic plan provides researchers conducting comparative effectiveness trials a framework within which to address various stakeholder questions and develop a trial design and analysis plan that can consider different populations. This approach avoids potential pitfalls present in choosing a single analytic plan before defining clinical trial goals. Most importantly, it ensures that analyses are guided by the central questions that various stakeholder groups hope to answer in their quest for results that are generalizable and can be applied to real-world clinical medicine.