Authors: Aidan M. Campbell, A. James O’Malley, Inas S. Khayal
Categories: Article, Mathematics and computing, Health services
Source: NPJ Digital Medicine
Authors: Aidan M. Campbell, A. James O’Malley, Inas S. Khayal
Measures of healthcare quality and equity often overlook when care is delivered, potentially masking important disparities. We present a novel, time-aware approach to detect racial inequities in hospice use among 100,480 Medicare beneficiaries with advanced cancer. Tracking hospice initiation over the final 200 days of life, we introduce a daily “difference signal” that shows how utilization patterns change over time. Using a Bayesian framework to quantify uncertainty and applying clinically meaningful thresholds, we pinpoint when disparities arise and how large they are. Compared to a conventional, time-agnostic benchmark, our approach reveals substantial day-level disparities that aggregate measures can miss. Notably, while overall measures often show greater hospice access for white patients, our temporal analysis frequently found earlier hospice use by patients of color, demonstrating how timing can reverse apparent patterns. By identifying when disparities emerge, our method offers more actionable targets for interventions and future digital health equity efforts.
A patient experiences healthcare delivery as a sequence of events that occur over time. Notably, the relative timing and duration of these events have an important effect on the patient’s experience of their care. Consider two patients with ultimately terminal cancer who both receive aggressive chemotherapy followed by hospice care and then death. Neglecting the temporal component, these two patients appear to have utilized chemotherapy and hospice identically. However, suppose that one ceased chemotherapy 100 days before death and was in hospice for the final 50 days of life, while the other ceased chemotherapy only 5 days before death and was only in hospice for the final 4 days of life. We have reason to believe these kinds of differences in timing matter. For example, Choi et al. found that early referral to hospice and palliative care has been shown to improve the quality of dying and death, with patients receiving hospice and palliative care for at least 22 days experiencing better outcomes^1^. Hence, the temporal dimension is essential in sufficiently understanding the patient experience. Yet, most methods in health services research for understanding hospice and palliative care utilize a population quality measurement approach whereby only the population-level utilization of the service is measured, while the timing of care is neglected^2^. Current approaches answer the question ‘What fraction of patients ever received hospice?’, but don’t answer the question, ‘When did patients receive hospice, if ever?’
Healthcare providers and administrators recognize the need to address limitations in the current quality measurement approach as it applies to end-of-life care for patients with cancer^3^. The deeper insight provided by measures that elucidate the temporal aspect of care delivery can provide targets for more actionable quality improvement interventions that are obscured by currently used quality measures^4^.
Quality measures are often used to compare care delivery across patient populations and can help identify racial inequalities in care delivery^5^. Racial and ethnic disparities in healthcare delivery represent a significant threat to quality care, as evidenced by the Institute of Medicine’s recognition of equity as one of the six core domains of healthcare quality^6^. Evidence demonstrates that timing-related inequities–a critical element of care quality^6^-disproportionately affect marginalized populations across diverse cancer types, including breast^7–9^, lung^10^, prostate^11^, colon^12^, melanoma^13^, endometrium^14^, bladder^15^, and brain cancers^16,17^, as well as in end-of-life services such as hospice care^18–21^.
However, as discussed, current approaches are insufficiently descriptive of the temporal aspect of care delivery^22^. As a result, inequalities in the temporal dimension of care may be systematically undetected and, therefore, left unaddressed. Previous work has demonstrated a novel approach to quality measurement that incorporates the temporal aspect of care^2^. In this paper, we add to this methodology by quantifying uncertainty to identify specific days with clinically meaningful racial disparities in hospice care utilization in cancer patients at the end of life. Specifically, the novel contribution of this paper is the methodology used for uncertainty quantification and ensuing determination of statistical inferences.
At the core of our analysis is the difference signal (diffSignal), which gives, for each day before death, the fraction of white patients who received hospice by that day minus the fraction of patients of color who received hospice by that day. We present a method for quantifying the uncertainty of the difference signal, specifically by computing the posterior probability distribution of the value of the difference signal for each day, along with related measures derived from the difference signal. We also apply four conventional tests of statistical significance to test for disparities in longitudinal care delivery that serve as a comparison to our developed procedure.
We apply this battery of measures and significance tests to the analysis of racial inequalities in the receipt of hospice care for white patients and patients of color with cancer at the end-of-life on a per-hospital basis. In this study, patients are categorized as patients of color if they are not classified as white patients. We then compare the findings of the various methods presented and include detailed results from two hospitals to exemplify our approach.
Our analysis included a total of 100,480 Medicare beneficiaries (patients) with advanced cancer across 773 U.S. hospitals. Per CMS data suppression rules, we are only allowed to report on hospitals that have 11 or more patients. Since we are comparing two groups, the 773 hospitals analyzed include only those with at least 11 people of color and 11 white people who meet the inclusion criteria. Of the patients, 22,355 (22.25%) were people of color (POC), while 78,125 (77.75%) were white people (WP). The median number of white patients per hospital was 81, with a range of 11–659 and an interquartile range (IQR) of 44–135. The median number of patients of color per hospital was 21, with a range of 11–219 and an interquartile range (IQR) of 14–34.
Of the 773 hospitals, 90 were Academic Medical Centers (AMC), 30 were National Comprehensive Cancer Network Centers (NCCN), 22 were National Cancer Institute–Designated Cancer Centers (NCI) but not NCCN, and 631 were community hospitals. Hospitals were located in metropolitan (739/773), micropolitan (33/773), and small town (1/773) urban areas. Table 1 shows a summary of demographic and clinical characteristics of the patients and hospitals included in our analysis.Table 1Demographic and clinical characteristics of patients and hospitalsCharacteristicWhite patientsPatients of colorp-value(N = 78,125)(N = 22,355) Patient characteristics Age at death, years, median [IQR]77 [71–83]75 [70–82]0.00 Sex, n (%)0.01 Female37,000 (47.4%)10,818 (48.4%) Male41,125 (52.6%)11,537 (51.6%) Prevalence of selected cancer diagnoses, n (%), Cancer of the… Lung29,834 (38.19%)7562 (33.83%)<0.001 Pancreas8433 (10.79%)2623 (11.73%)<0.001 Brain and nervous system7410 (9.48%)1672 (7.48%)<0.001 Liver9068 (11.61%)3400 (15.21%)<0.001 Esophagus3344 (4.28%)738 (3.30%)<0.001 Hospital characteristics (overall setting, N = 773) ** Hospital type** n (%) Academic medical center90 (11.6%) NCCN center30 (3.9%) NCI-designated (non-NCCN)22 (2.8%) Community hospital631 (81.6%) ** Hospital location (RUCA)** Metropolitan739 (95.6%) Micropolitan33 (4.3%) Small town/rural1 (0.1%)Data are presented as Median [Interquartile Range] or n (%). Percentages may not sum to 100 due to rounding. P-values compare the White Patient and Patient of Color cohorts and are derived from Mann–Whitney U tests for continuous variables and Chi-squared tests for categorical variables. P-values are not applicable to Hospital Characteristics, which are descriptive of the entire set of included facilities.IQR interquartile range, NCCN National Comprehensive Cancer Network, NCI National Cancer Institute, RUCA rural-urban commuting area.
We applied our methodology, which included six disparity tests, to analyze racial differences in hospice utilization at the end of life. The descriptions of each test, along with the results from two hospitals, as illustrative examples, are shown in Fig. 1.Fig. 1Statistical framework for analyzing racial disparities in hospice care timing, illustrated with two hospital examples.Six approaches are a A Fisher Exact (FE) test is used to compare overall hospice utilization rates between white patients and patients of color, providing a time-agnostic measure of access disparity. b First Day of Hospice (FDH) test uses a Mann–Whitney U to compare timing of hospice initiation among patients who received hospice care. c Number of Hospice Days (NHD) test uses Mann–Whitney U to compare total hospice duration across all patients, with non-hospice patients contributing zero days. d The diffSignal shows, for each day, the fraction of white patients who had at least one day of hospice by that day, minus the fraction of POC patients who had had at least one day of hospice by that day, with shaded 95% credible intervals. e Probable Clinically Significant Difference (PCD) analysis identifies days when the diffSignal exceeds 5 percentage points in magnitude with >95% WMPCD(+) favors white patients, PMPCD(+) favors patients of color. f Day-by-Day Fisher Exact test (DBDFE) applies a Fisher Exact test to compare the fraction of white and POC patients who have had at least one day of hospice by each day, with DBDFE(+) indicating statistical significance (p < 0.05).
In Fig. 1a, we assessed disparities using a standard quality measures approach with a Fisher Exact test (FE) to compare the proportion of patients of color and white patients who received hospice care at any time during the observation period. A hospital is FE(+) if p < 0.05 and FE(−) if p ≥ 0.05. Next, in Fig. 1b, we compare the time relative to death of the first day of hospice between white patients and patients of color who received hospice care using a Mann–Whitney U Test (FDH). A hospital is FDH(+) if p < 0.05 and FDH(−) if p ≥ 0.05. In Fig. 1c, we compare the number of days spent in hospice by patients of color and white patients using a Mann–Whitney U Test (NHD). Patients who did not receive hospice spent 0 days in hospice. A hospital is NHD(+) if p < 0.05 and NHD(−) if p ≥ 0.05.
In Fig. 1d, we visualize the diffSignal (dark blue), which gives, for each day before death, the fraction of white patients who received hospice by that day minus the fraction of patients of color who received hospice by that day. We show a 95% credible interval (tan) around the difference signal for each day—computed using a Bayesian approach to approximate the posterior probability distribution of the difference signal value for each day. A critical point is that these computations are performed independently for each day, which, for pointwise inference, requires no assumptions about the statistical independence of the values for each day.
In Fig. 1e, we visualize two (1) the conditional probability given the data of a clinically significant difference with POC receiving more hospice (PMPCD, pink) and (2) the conditional probability given the data of clinically significant difference with white people receiving more hospice (WMPCD, green). We shaded days that showed a high (>0.95) PMPCD or WMPCD value; we refer to such days as PMPCD(+) and WMPCD(+), respectively. In the remainder of the paper, we generically refer to days that are either PMPCD(+) or WMPCD(+) as PCD(+).
Finally, in Fig. 1f, we show the p-values of a Fisher exact test applied on a day-by-day basis comparing the proportion of patients of color and white patients who received hospice by each day before death (DBDFE, light blue). Days with a statistically significant (<0.05) DBDFE value are shaded; we refer to such days as DBDFE(+).
Across the 773 hospitals, 44 (5.69%) hospitals were FE(+), while 46 (5.95%) were FDH(+), though only 4 hospitals were both FE(+) and FDH(+). Of the FE(+) hospitals, 32 (4.14%) had WP receiving more hospice, while 12 (1.55%) had POC receiving more hospice. Of the 154, 600 days (200 days contributed for each of the 773 hospitals) included in our analysis, 1485 (0.96%) were PCD(+), of which 1214 (0.79%) were PMPCD(+) and only 271 (0.18%) were WMPCD(+). In addition, 3204 days (2.07%) were DBDFE(+).
No hospitals had both PMPCD(+) and WMPCD(+) days; in other words, if a hospital had one or more PMPCD(+) days across all of its days, then it had no WMPCD(+) days across any of its days, and vice versa. Additionally, if a hospital had at least one PCD(+) day and was FE(+), then if that day was WMPCD(+), the disparity detected by the FE test had WP receiving more hospice, while if that day was PMPCD(+), the disparity detected by the FE test had POC receiving more hospice. In other words, for the hospitals we analyzed, if both the FE test and the PMPCD series agree that the hospital has a disparity, then they agree on the direction of the disparity with respect to WP or POC receiving more hospice. Of the 148 (19.15%) hospitals that had at least one PCD(+) day, 78 (10.09%) had at least one PMPCD(+) day, while 70 (9.06%) had at least one WMPCD(+) day. In addition, 205 (26.52%) hospitals had at least one DBDFE(+) day. Table 2 shows how many hospitals had at least 1, 5, 10, 15, 20, 25, 30, or 35 positive days for each day-level test.Table 2The number of hospitals with at least n PMPCD(+), WMPCD(+), PCD(+), and DBDFE(+) days, for n ∈ {1, 5, 10, 15, 20, 25, 30, 35}daysPMPCD(+)WMPCD(+)PCD(+)DBDFE(+)≥178 (10.09%)70 (9.06%)148 (19.15%)205 (26.52%)≥551 (6.60%)19 (2.46%)70 (9.06%)125 (16.17%)≥1035 (4.53%)6 (0.78%)41 (5.30%)86 (11.13%)≥1524 (3.10%)2 (0.26%)26 (3.36%)63 (8.15%)≥2021 (2.72%)0 (0%)21 (2.72%)52 (6.73%)≥2519 (2.46%)0 (0%)19 (2.46%)43 (5.56%)≥3014 (1.81%)0 (0%)14 (1.81%)35 (4.53%)≥3510 (1.29%)0 (0%)10 (1.29%)28 (3.62%)The percent of all hospitals that each count represents is shown in parentheses. A similar number of hospitals had at least one PMPCD(+) day as had at least one WMPCD(+) day, but many more hospitals had a large number of PMPCD(+) days than had a large number of WMPCD(+) days.
For each day before death, Fig. 2 shows the number of hospitals that had a PMPCD(+), WMPCD(+), and DBDFE(+) result for that day. Greater disparities were detected closer in time to death according to the PMPCD and WMPCD series, and WMPCD(+) days were concentrated nearer to death than PMPCD(+) days.Fig. 2Number of hospitals with a positive result for WMPCD (green), PMPCD (pink), and DBDFE (light blue) for each day before death.Day 0 represents the day of death. Although there were about the same number of hospitals with at least one PMPCD(+) day as with at least one WMPCD(+) day, there were many more days detected as PMPCD(+) than as WMPCD(+).
Figure 3 illustrates the relationship between four the number of PCD(+) days, the timing of those PCD(+) days relative to death, whether they were classified as PMPCD(+) or WMPCD(+), and whether or not the hospital was FE(+) or FE(−). This figure illustrates several interesting findings. In Fig. 3a, we see that the majority of hospitals that have at least one WMPCD(+) day and are additionally FE(+) (meaning that there is a statistically significant difference in the overall proportion of WP and POC patients who received any hospice care, with WP patients being more likely to receive hospice care) have fewer than five PCD(+) days with the average timing of these days being fewer than four days before death. Meanwhile, we observe in Fig. 3b a notable cluster mostly composed of hospitals with at least one PMPCD(+) day that are FE(−). This is the result of the fact that the set of hospitals with a large number of PCD(+) days is dominated by those with PMPCD(+) days (Fig. 3c), and these PMPCD(+) days are, on average, further from death (Fig. 3d).Fig. 3Hospitals with at least one PCD(+) day annotated to emphasize key observations.a Among FE(+) hospitals with at least one WMPCD(+) day (where white patients show significantly higher hospice utilization than patients of color), most have fewer than five PCD(+) days, occurring on average within four days of death. b Distinct cluster of mostly FE(−) hospitals, each with many PMPCD(+) days. c Hospitals with a very large number of PCD(+) days have almost only PMPCD(+) days. d PMPCD(+) days occur further from death on average. The average days before death of PCD(+) days is shown incremented by one, in order to display data where the average day before death was zero. A small amount of noise has been added to the scatter plot to prevent overlapping of points; this noise is only present in the scatter plot portion of the figure and is not reflected in the histograms.
Table 3 compares the FE test to the FDH test, showing the number of hospitals positive and negative for each. We found that there was no strong relationship between the results of the two tests, with a Fisher Exact test applied to Table 3 yielding p = 0.32.Table 3Number of hospitals with each combination of FDH and FE resultsFE(+)FE(−) FDH(+)442 \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ P({\bf{FE}}(+),| {\bf{FDH}}(+))=4/(4+42)=0.087
**FDH**(−)40687\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ P({\bf{FE}}(+)\,| {\bf{FDH}}(-))=40/(40+687)=0.055 $$\end{document}P(FE(+)|FDH(−))=40/(40+687)=0.055There is no strong relationship between a hospital’s **FE** result and a hospital’s **FDH** result. Odds of 0.087 and 0.055 are comparable. The contingency table Fisher Exact test *p* = 0.32. Table 4 compares the **FE** and **PCD** tests by showing the number of **FE**(+) and **FE**(−) hospitals as well as whether those hospitals had at least one or no **PCD**(+) days. Of the hospitals with at least one **PCD**(+) day, the number for which those days were **PMPCD**(+) and **WMPCD**(+) is also shown. We found that there was a statistically significant (*p* = 1.97 × 10^−28^) relationship between a hospital being **FE**(+) and having at least one **PCD**(+) day. However, there were 107 hospitals with at least one **PCD**(+) day but that were **FE**(−), illustrating how approaches that neglect the timing of care delivery can overlook disparities.Table 4The number of hospitals with each combination of FE and PCD results0 days PCD(+)≥1 day PCD(+)≥1 day PMPCD(+)≥1 day WMPCD(+) **FE**(+)3411229 **FE**(−)6221076641The first two columns show complementary results for hospitals with 0 days versus ≥1 day of **PCD**(+). The third and fourth columns provide a drill-down of the second column, showing the breakdown of hospitals with ≥1 day **PCD**(+) by the direction of the detected disparity (**PMPCD**(+) vs. **WMPCD**(+)). 107 of the 773 hospitals with at least 1 **PCD**(+) day had no significant difference in the overall hospice utilization. This illustrates how approaches that neglect the timing of care delivery fail to identify existing disparities. Table 5 compares the Bayesian **PCD** test with the **NHD** test, which considers the number of days spent in hospice by patients of each group. A highly significant association was found between the two tests (*p* = 6.09 × 10^−41^). Of the 54 hospitals flagged as having a disparity by the **NHD** test, 53 also had at least one **PCD**(+) day. However, though there was a significant association between the two tests, the **PCD** analysis identified disparities missed by the **NHD** test. Specifically, 95 hospitals that showed no significant difference in the overall number of hospice days (**NHD**(−)) were found to have at least one day with a probable clinically significant difference (**PCD**(+)). This finding illustrates how even analyses that investigate the temporal dimension of care delivery may miss disparities revealed by the more detailed methods we present.Table 5The number of hospitals with each combination of NHD and PCD results0 days PCD(+)≥1 day PCD(+)≥1 day PMPCD(+)≥1 day WMPCD(+) **NHD**(+)1532330 **NHD**(−)624955540The first two columns show complementary results for hospitals with 0 days versus ≥1 day of **PCD**(+). The third and fourth columns provide a drill-down of the second column, showing the breakdown of hospitals with ≥1 day **PCD**(+) by the direction of the detected disparity (**PMPCD**(+) vs. **WMPCD**(+)). 95 of the 773 hospitals with at least 1 **PCD**(+) day had no significant difference in the number of hospice days. Like Table 4, *these results also illustrate how approaches that neglect the timing of care delivery can fail to identify disparities*. The contingency table Fisher Exact test comparing **NHD** and **PCD** results yields *p* = 6.09 × 10^−41^. Table 6 compares the results of the **FE** test, which assesses differences in the proportion of patients ever receiving hospice, with the **NHD** test. The results show a strong positive relationship between the two measures (Fisher Exact test *p* = 2.21 × 10^−32^). Among the 44 hospitals with a significant disparity in hospice utilization rates (**FE**(+)), a majority (32) also showed a significant disparity in the number of hospice days (**NHD**(+)). Similarly, of the 729 hospitals with no disparity in utilization rates (**FE**(−)), the vast majority (707) also had a negative **NHD** result. The strong correlation between the **FE** and **NHD** results should be contrasted with the failure to find a correlation between the **FE** and **FDH** results. This suggests that the correlation may be driven by the number of patients in each group receiving no hospice, who are omitted from the **FDH** analysis, but are included as having 0 days of hospice in the **NHD** analysis.Table 6Number of hospitals with each combination of NHD and FE resultsNHD(+)NHD(−) **FE**(+)3212 **FE**(−)22707There is a strong positive relationship between a hospital’s **NHD** result and a hospital’s **FE** result. The contingency table Fisher Exact test *p* = 2.21 × 10^−32^. Figure 4 compares the number of hospitals with **PMPCD**(+) and **WMPCD**(+) days. While a comparable number of hospitals had at least one **PCD**(+) day of either type, those with **WMPCD**(+) days were more than twice as likely to be **FE**(+). This disparity reflects the different temporal patterns of **PMPCD**(+) and **WMPCD**(+) days shown in Figs. 2 and 3.Fig. 4Hospitals with at least one PCD(+) day categorized by whether the PCD(+) days were PMPCD(+) or WMPCD(+) and then further subdivided by whether the hospital was FE(+) or FE(−).Note that no hospitals had both **WMPCD**(+) and **PMPCD**(+) days. Additionally, note that only two hospitals were **FE**(+) but had no **PCD**(+) days and that if a hospital was **PCD**(+) and **FE**(+), then the disparity detected by the one test favored the same group as that detected by the other test. Approximately the same number of hospitals had at least one **PMPCD**(+) day as had at least one **WMPCD**(+) day, but hospitals with at least one **WMPCD**(+) day were more than twice as likely to be **FE**(+) than hospitals with at least one **PMPCD**(+) day. Table 7 compares the **DBDFE** test to the **PCD** test, showing the number of days categorized as positive and negative for each. We found that there was a very strong positive relationship between the results of the two tests, with a Fisher Exact test applied to Table 7 yielding *p* ≈ 0.Table 7The number of days with each combination of **PCD** and **DBDFE** resultsPCD(+)PCD(−) **DBDFE**(+)13841820 **DBDFE**(−)101151,343There is a strong positive relationship between the two tests, with the contingency table Fisher Exact test yielding *p* ≈ 0. To assess the robustness of our findings, we conducted a sensitivity analysis by varying the threshold for a clinically significant difference in the **PCD** approach to 3% and 7%. Lowering the threshold to 3% increased the number of hospitals with at least one **PCD**(+) day to 234 (30.27% of hospitals), while raising it to 7% decreased the number to 96 (12.42% of hospitals), compared to 148 (19.15% of hospitals) at the original 5% threshold. Despite these expected shifts in volume, our core conclusions remained unchanged. The vast majority of disparity days consistently favored earlier hospice initiation for patients of color (**PMPCD**(+)), accounting for 81.3% of **PCD**(+) days at the 3% threshold, 81.8% at 5%, and 81.1% at 7%. Moreover, the key finding that our temporal method identifies inequities missed by traditional, frequency-based tests (**FE**) held firm, with 190 hospitals having at least one **PCD**(+) day being classified as **FE**(−) at the 3% threshold and 63 at the 7% threshold. This sensitivity analysis confirms that, while the absolute number of identified disparities depends on the selected threshold, the fundamental patterns and conclusions of our study remain robust. As an additional robustness check within the existing cohort, we repeated the day-level analyses after truncating the observation horizon to the final 100 days of life. Truncating the observation horizon to 100 days slightly but negligibly increased the number of hospitals with at least one **PCD**(+) day to 153 (19.79% of hospitals), in contrast to 148 hospitals (19.15% of hospitals) using the full 200 days. This difference was due to 4 additional hospitals having at least one **PMPCD**(+) day (82 hospitals compared to 78 hospitals with the full 200-day look-back) and one additional hospital having at least one **FDH**(+) day (71 with a 100-day window versus 70 with the 200-day window). There was also a minimal increase in the number of hospitals found to be **FDH**(+): 51 for 100 days versus 46 for 200. Despite these small changes, our main observations at 200 days hold for the 100-day look-back data. The vast majority of days with detected disparities still favored earlier hospice initiation among people of color, with 79.7% of **PCD**(+) days being **PMPCD**(+) days with the 100-day look-back, similar to the 81.8% for the 200-day look-back. The number of **PCD**(+) days didn’t substantially decrease with the truncated 100-day window, 1435 days for 100 days versus 1485 days for 200 days. This is because, as shown in Fig. 2, most of the **PCD**(+) days are concentrated close to death. The finding that our temporal method identifies inequities missed by traditional, frequency-based tests could be appreciated under either look-back 112 hospitals had at least one **PCD**(+) day but where **FE**(−) for the 100-day look-back compared to 107 hospitals for the 200-day look-back. This sensitivity analysis shows that although we observe small changes in the results with a truncated dataset comprising only the final 100 days of life, these fluctuations do not affect the fundamental patterns observed or the main conclusions drawn from our primary analysis. ## Discussion In this paper, we demonstrate how the difference signal and Bayesian measures of probable clinically significant differences (**PMPCD,** **WMPCD**, and **PCD**) allowed us to identify when, in the final days of life, clinically meaningful disparities emerged. The frequentist analog of the **PCD** approach, the **DBDFE** test, also allowed us to identify days with a likely disparity. There was a strong positive correlation between the results of both approaches. Both are suitable for gaining a better understanding of the timing of care delivery, with each having distinct advantages and disadvantages. The **DBDFE** approach is simpler to implement, while the **PCD** is more complex, but allows for a more nuanced and straightforward interpretation of the data. Our findings suggest that (1) incorporating temporal and probabilistic measures, such as the difference signal and the **PCD** approach, can reveal racial disparities that static, aggregate measures may overlook; (2) disparities in hospice access (**FE**) and timing (**FDH,** **PCD**) often operate independently but can interact in complex ways; (3) understanding when in the care trajectory disparities emerge provides richer insight into plausible mechanisms, aligning with and extending prior work; and (4) combining time-aware, multi-measure approaches yields more actionable targets for quality improvement, methodological innovation, and policy design. Our analysis compared results from the **FE** and **FDH** tests to investigate whether a systemic relationship exists between disparities in hospice access and disparities in timeliness. The test for association, based on the data in Table 3, found no statistically significant relationship. This finding of independence suggests that the failures leading to inequitable access are distinct from those leading to inequitable timeliness. A second, crucial observation from our results helps to contextualize this independence. We found that, overall, **FE**(+) hospitals were more likely to favor white patients, while our time-aware analyses found that timing disparities often involved patients of color initiating hospice earlier. The existence of these differing, often opposing, disparity patterns is consistent with the statistical independence we observed. The non-correlation of disparities in timeliness and access underscores the need for a multi-faceted analytical approach, as different dimensions of healthcare equity can operate independently and even in opposing ways. The broader literature on end-of-life care provides a compelling context for our finding that disparities in whether patients received hospice at all tended to find that white patients received hospice at a higher rate, while the time-aware methods we present reveal that, among patients who received hospice, there was more likely to be a disparity where patients of color initiated hospice earlier. Mirroring the direction of disparity often seen in our **FE**(+) hospitals, the most common line of research focuses on barriers to access, frequently concluding that minority patients are less likely than white patients to utilize hospice at all. For example, Fairfield et al. found that Black women with ovarian cancer were more likely to never receive hospice care^23^, and Kwak et al. reported that Black nursing home residents were significantly less likely to use hospice services^24^. In contrast, a separate body of work examining the timing of care for enrolled patients aligns with our temporal analyses, often finding that minority patients experience longer, not shorter, hospice stays. An early, large-scale study by Christakis and Iwashyna found that non-white Medicare beneficiaries enrolled in hospice a median of four days earlier than white beneficiaries^25^. This pattern has been consistently observed since Park et al. reported longer hospice stays for Hispanic and African-American patients^19^, and Yu and Brown found a similar pattern for African Americans in a national sample^26^. Several factors may explain why enrolled patients of color sometimes experience longer hospice stays. One possible cause relates to the intensity of treatment preceding hospice enrollment. White patients may have greater access to, or be more likely to pursue, aggressive, life-extending treatments such as novel chemotherapies or clinical trials. This pursuit of curative-intent care until very late in the disease trajectory would naturally result in a delayed, and therefore shorter, hospice admission. Conversely, if patients of color have less access to these late-stage options or transition away from them earlier, they would enter hospice sooner relative to death, leading to a longer length of stay. The finding by Park et al. that referral source moderates these timing differences further suggests the influence of specific healthcare system structures rather than just patient-level factors^19^. Temporal analysis of care delivery can improve our understanding of the care trajectory and refine our understanding of care quality^27^. While traditional comparisons, such as applying a Fisher Exact test to compare the overall rate of hospice utilization of two populations, do detect disparities in aggregate utilization, these methods fail to detect prevalent and meaningful disparities in the timing of care delivery. In contrast, the time-aware approaches we present can not only identify the presence of inequality in the timing of care delivery but can also pinpoint when disparities begin to arise in the care trajectory. Our finding that disparities where white patients received more hospice tended to be closer to the end of life, while disparities where patients of color received more hospice tended to arise earlier relative to death underscores the temporal complexity of care delivery and the corollary that to understand care delivery well; it must be investigated longitudinally. The combination of the traditional and time-aware measures revealed a more nuanced picture than could have been gained by applying either alone. Some hospitals showed no difference in overall utilization but exhibited statistically and clinically significant differences in the timing of care delivery, while other hospitals had no meaningful difference in the timing of care delivery and yet had significant differences in aggregate hospice utilization (though only two hospitals were **FE**(+) but had no **PCD**(+) days). These results bolster the conclusion that healthcare stakeholders who rely solely on non-temporal measures may overlook substantial inequalities in care. Our findings suggest that even hospitals with relatively balanced aggregate metrics could be failing certain patient subgroups at specific times in their care trajectories. The analysis comparing our **PCD** approach with the **NHD** test results further demonstrates that approaches neglecting details in the timing of care delivery can overlook disparities. While the **NHD** test showed a significant association with our **PCD** results, 95 hospitals that showed no significant difference in total hospice days (**NHD**(−)) were found to have at least one day with a probable clinically significant difference (**PCD**(+)). This finding confirms that the timing of hospice care delivery contains informational value that cannot be captured by examining total duration alone, reinforcing the utility of our temporal approach. From a practical perspective, the methods we present can guide more targeted interventions. Understanding when inequalities in hospice utilization emerge-and-widen enables healthcare providers and administrators to optimize the timing of hospice-related services and conversations. For example, disparities that manifest primarily in the final month of life may indicate a need for earlier, culturally sensitive discussion of hospice care options. Conversely, disparities that emerge further from death might suggest systematic barriers to timely hospice referral, such as inadequate advanced care planning processes or structural impediments, including language accessibility, established therapeutic relationships, and logistical constraints. A granular temporal understanding of where disparities exist can help healthcare systems more effectively allocate resources to address specific mechanisms driving inequitable hospice utilization. Our Bayesian approach used distinct probability parameters and non-informative \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\mathcal{B}}(1,1) $$\end{document}B(1,1) priors for each day to preserve simplicity and avoid inadvertent bias in the detection of disparities. A more sophisticated survival-time modeling approach could pool information across time to form a hospice admission day variable that is subject to censoring (some patients are not admitted to hospice by their date of death). By accounting for the fact that each patient can only enter hospice once, this approach would likely improve statistical precision and better capture the temporal structure of end-of-life care. However, it would introduce the more onerous task of specifying a survival-time model and the need to account for the clustering of observations by hospital (e.g., a hierarchical survival-time model would be needed); the latter is avoided by our current approach because the dependent variable is a hospital-level binomial count. In addition, our current specification yields parameter estimates that are more amenable to interpretability for clinical audiences. Future methodological work should systematically compare these approaches to determine whether the increased complexity of hierarchical survival-time modeling yields meaningfully different conclusions about racial disparities while preserving the actionability that is crucial for clinical implementation. Several limitations warrant consideration in interpreting our findings. While administrative claims data provide longitudinal information about service utilization, they lack the clinical detail necessary to fully elucidate the mechanisms driving temporal differences in hospice enrollment. We can identify that disparities exist, but not why. This approach provides insights that quality improvement personnel can use to guide a detailed search for the why, typically in electronic medical records. Our restriction to Medicare fee-for-service beneficiaries introduces potential selection bias, as this population excludes Medicare Advantage enrollees who may have different hospice utilization patterns. Additionally, our findings may not generalize to patients younger than those in our cohort, including Medicaid beneficiaries and those with private insurance, who may exhibit different end-of-life care patterns. Future research should reproduce these results in these populations. Additionally, while our claims data include diagnostic and procedure codes, we did not incorporate detailed clinical factors such as specific cancer staging, disease progression markers, or cause-specific mortality patterns in our analysis of disparities. These might interact with both racial identity and hospice initiation behavior, and future research incorporating these clinical factors could provide deeper insight into the mechanisms underlying the disparities we identify. Medicare fee-for-services claims data are also subject to well-documented data quality limitations. We did include an additional RTI race algorithm^28^ that begins to address this limitation. Future research will include any new algorithms to alleviate this limitation. Furthermore, although our selection of temporal thresholds and observation windows was informed by scientific norms and clinical relevance, these parameters remain inherently arbitrary. Our use of a 5% threshold should be understood as a pragmatic interpretive choice rather than a model assumption. While the Bayesian inference is independent of this threshold, summarizing results relative to such a standard helps distinguish changes that are likely to be regarded as clinically or policy relevant. Readers applying this framework are free to adopt alternative thresholds according to their own judgments of meaningful change. The independent daily modeling also represents a statistical simplification. By not formally modeling the autocorrelation between days, our analysis is limited to identifying a collection of individual days where a disparity is likely present. A more complex time-series model at the hospice-day level or hierarchical survival-time model at the individual patient level could make formal probabilistic statements about the disparity as a continuous process, such as estimating the probability that a disparity persists for a certain duration (e.g., “for at least 10 consecutive days") or identifying the single time point at which a sustained period of disparity is most likely to have begun. While our approach does not offer this type of inference on temporal periods, it does not alter our main conclusions, which are focused on identifying the specific days and temporal windows where these disparities are most evident. While the cohort was defined using a 200-day look-back window, exploratory truncation to 100 days (not a full cohort-level sensitivity analysis) produced results that were substantively unchanged. This supports the robustness of our conclusions to the choice of observation horizon. Additional research incorporating alternative temporal parameters and qualitative methodologies could provide deeper insight into the decision-making processes of both providers and patients. Our focus on Medicare fee-for-service beneficiaries may also constrain the extent to which our results can be generalized, as these may not reflect patterns observed in other insurance populations or geographic regions. Having acknowledged these limitations, this study demonstrates the critical importance of incorporating temporal information in quality measurement approaches. The methodology presented here provides a statistical framework for detecting racial inequalities in longitudinal care delivery that conventional population-level quality measures may fail to identify. By characterizing when racial disparities in hospice utilization emerge and evolve, this approach enables more precise identification of temporal windows where targeted interventions may be most effective in promoting equitable access to end-of-life care across racial and ethnic groups. In the future, we hope that this enhanced temporal understanding of care delivery patterns can be integrated into digital health analytics systems, enabling day-by-day monitoring of patient trajectories. By pinpointing exactly when disparities arise and intensify, the methodology we present can inform the development of evidence-based policies and quality improvement initiatives aimed at reducing disparities in hospice utilization in specific healthcare systems. ## Methods The “Methods” section first details the ***Patients and Hospitals*** and continues with a description of the analytical approach with sections detailing the ***Measurement Window and Data Construction***, the calculation of the ***Difference Signal***, the ***Bayesian Inference for Uncertainty Quantification***, ***Assessing Clinical and Statistical Significance***, ***Conventional Significance Tests***, and ***Implementation of Analysis***. Figure 5 visualizes a summary of the key variables and statistical tests involved in the methodology applied to each hospital.Fig. 5Summary of key variables and statistical tests.involved in the methodology applied to each hospital. ### Patients and hospitals We conducted a retrospective analysis of Medicare fee-for-service beneficiaries with advanced (poor-prognosis) cancers, drawn from Medicare fee-for-service claims data from a retrospective study of decedents completed by Wasp et al.^29^ Specifically, we used a 100% sample of Medicare fee-for-service beneficiaries drawn from 2016 to 2017 Centers for Medicare and Medicaid Services (CMS) files, (1) the Master Beneficiary Summary file, and (2) the Hospice file. Advanced, poor-prognosis cancers were defined following the methods of Iezzoni et al.^30^, as adapted to ICD-10-CM codes by Wasp et al.^29^. Under these criteria, patients were included if they were diagnosed with primary or metastatic cancers known to have high near-term mortality risk. From the identified population, we selected patients who (1) died between January 1st, 2017 and December 31st, 2017, (2) were aged 66–99 years at the time of death, (3) had at least one admission for cancer in the final six months of life or at least two outpatient oncologist visits, and (4) had complete claims data for a six-month look-back period (June 1st, 2016–June 30th, 2017) to ascertain hospice utilization. Each patient was attributed to the hospital where they received the preponderance of their cancer-related care in the final six months of life^29^. Per CMS data suppression rules, we restricted our analysis to hospitals with at least 11 white decedents and 11 decedent people of color (POC). People of color include Black or African-American, Asian/Pacific Islander, Hispanic, American Indian/Alaska Native, and Other decedents. While the POC group conflates heterogeneous population subgroups with unique healthcare experiences, they share a history of marginalization in the U.S. Patients were categorized as white or POC based on enhanced CMS race and ethnicity codes, which refine Social Security Administration data with algorithms (i.e., RTI race algorithm) to more accurately identify Hispanic and Asian beneficiaries^28,31^. The prevalence of selected cancer diagnoses shown in Table 1 was computed by aggregating all ICD-10-CM diagnostic codes associated with each patient across all Medicare fee-for-service files used in our analysis, matching these diagnostic codes with the corresponding CCS codes, then counting the number of unique patients with each CCS code. Hospital characteristics included hospital type (National Cancer Institute–Designated Cancer Centers (NCI), National Comprehensive Cancer Network Centers (NCCN), Academic Medical Centers (AMC), and Community Hospitals), rurality, and the number of white patients and patients of color with advanced cancer treated who died between January 1st, 2017 and December 31st, 2017, and met all inclusion criteria. We identified hospital rurality from publicly available 2010 Rural-Urban Commuting Area (RUCA) codes from the U.S. Department of Agriculture website^32^. The Dartmouth Committee for the Protection of Human Subjects (CPHS) approved this study (CPHS# STUDY00033050). ### Measurement window and data construction We recorded hospice utilization during each patient’s final 200 days of life, indexing time relative to death^33^. Specifically, we defined day 0 as the day of death, day 1 as one day before death, and so on until day 199. This indexing means that a lower-numbered day is chronologically closer to death, while a higher-numbered day is earlier in the disease trajectory. For each hospital, we recorded the total number of white patients (*n*~*w*~) and POC patients (*n*~*p*~). We also noted how many patients from each group utilized hospice at least once in the final 200 days of life (*u*~*w*~ and *u*~*p*~). For patients who utilized hospice care, we identified the first day they entered hospice and used this to create two vectors *X*~*w*~ and *X*~*p*~:*X*~*w*~ contains the day index of the first hospice day for each white patient who received hospice during the 200-day period. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \dim ({X}_{w})={u}_{w} $$\end{document}dim(Xw)=uw)*X*~*p*~ similarly contains the day index of the first hospice day for each POC patient who received hospice. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \dim ({X}_{p})={u}_{p} $$\end{document}dim(Xp)=up) For example, if a hospital had four POC patients who began hospice 20, 15, 15, and 10 days before death, respectively, *X*~*p*~ for that hospital would be [20, 15, 15, 10]. We also recorded the number of days spent in hospice by each patient, regardless of whether they received hospice or not. For each hospital, we used this data to create the vectors *D*~*w*~ and *D*~*p*~:*D*~*w*~ contains the number of days spent in hospice care by each white patient during the 200-day period. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \dim ({D}_{w})={n}_{w} $$\end{document}dim(Dw)=nw)*D*~*p*~ similarly contains the number of days spent in hospice care by each POC patient during the 200-day period. (\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \dim ({D}_{p})={n}_{p} $$\end{document}dim(Dp)=np) To represent hospice utilization, we constructed two binary matrices **W** and **P**. For each **W** is an *n*~*w*~ × 200 matrix, where **W**~*i*,*j*~ = 1 if the *i*-th white patient has initiated hospice by day 200 − *j* and is 0 otherwise.**P** is an *n*~*p*~ × 200 matrix, where **P**~*i*,*j*~ = 1 if the *i*-th POC patient has initiated hospice by day 200 − *j* and is 0 otherwise. Patients who never utilized hospice during the observation window correspond to rows of all zeros in the matrices just defined. ### The difference signal Traditional quality measures of hospice frequently focus on overall utilization rates, ignoring the timing of care delivery. These are computed by summing across the rows of **W** and **P** and forming a binary indicator variable of whether the row-sum exceeds 0. To capture how disparities evolve over time, we defined a daily difference signal in the style of Khayal et al.^2^:1\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\bf{diffSignal}}}_{j}=\left(\frac{1}{{n}_{w}}\mathop{\sum }\limits_{i=1}^{{n}_{w}}{{\bf{W}}}_{i,j}\right)-\left(\frac{1}{{n}_{p}}\mathop{\sum }\limits_{i=1}^{{n}_{p}}{{\bf{P}}}_{i,j}\right),\quad j\in [1,200]. $$\end{document}diffSignalj=1nw∑i=1nwWi,j−1np∑i=1npPi,j,j∈[1,200]. This difference signal describes, for day 200 − *j*, the difference between the fraction of white patients and the fraction of POC patients who had received hospice by that day. A positive value means a higher fraction of white patients had initiated hospice by that time, while a negative value indicates the opposite. By examining this signal across the 200-day window, we can see when disparities in hospice utilization emerge or intensify. Within the fixed 200-day cohort, we repeated the entire analysis after restricting the observation horizon to the final 100 days of life. This probes dependence on horizon length but is not a cohort-level sensitivity analysis to alternative look-back requirements. ### Bayesian inference for uncertainty quantification While the difference signal provides a point estimate of disparities by day, it does not directly quantify the uncertainty of these estimates. To address this, we applied a Bayesian approach to estimate the probability distribution of the value of the difference signal for each day. For each *j* ∈ [1, 200], we modeled the probability that a given group (white or POC) entered hospice by day 200 − *j* using a beta distribution, the conjugate prior for binomial data^34^. For day 200 − *j*, let \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{w} $$\end{document}pjw be the underlying probability that a randomly selected white patient at the hospital in question had received hospice by day 200 − *j*. Similarly, let \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{p} $$\end{document}pjp be the analogous probability for a randomly selected POC patient. Using a uniform \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {\mathcal{B}}(1,1) $$\end{document}B(1,1) prior, which is a non-informative prior that imparts no initial bias toward any particular value of the probability underlying the difference signal^34^, the posterior distributions 2\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{w} \sim {\mathcal{B}}\left(1+\mathop{\sum }\limits_{i=1}^{{n}_{w}}{{\bf{W}}}_{i,j},\,1+{n}_{w}-\mathop{\sum }\limits_{i=1}^{{n}_{w}}{{\bf{W}}}_{i,j}\right) $$\end{document}pjw~B1+∑i=1nwWi,j,1+nw−∑i=1nwWi,j3\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{p} \sim {\mathcal{B}}\left(1+\mathop{\sum }\limits_{i=1}^{{n}_{p}}{{\bf{P}}}_{i,j},\,1+{n}_{p}-\mathop{\sum }\limits_{i=1}^{{n}_{p}}{{\bf{P}}}_{i,j}\right) $$\end{document}pjp~B1+∑i=1npPi,j,1+np−∑i=1npPi,j We emphasize that our model constructs one such pair of distributions for each hospital, for each day. Critically, we are modeling the probability that a patient has already entered hospice by day 200 − *j*, not the probability of admission occurring on that day, which distinguishes our approach from time-to-event modeling. This beta-binomial framework assumes each patient’s hospice initiation (by a given day) is a statistically independent Bernoulli trial both cross-sectionally and longitudinally, with a probability that may change from day to day but that is invariant within patient groups within the same day, and the posterior updates our belief about the underlying probabilities \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{w} $$\end{document}pjw and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{p} $$\end{document}pjp given the observed data. The assumption that statistical independence holds between the status of the same subject across days is clearly unreasonable and so serves as a working model designed to yield pointwise credible intervals, not credible interval bands. This simplification is important because the cumulative nature of hospice initiation creates strong positive autocorrelation between adjacent days. Our pointwise approach allows us to sidestep the complexities of formally modeling this temporal dependence, ensuring the analysis remains focused on identifying when disparities emerge. While this means we cannot make joint probability statements across multiple days, it does not impact the validity of the inference for any single day, which is the foundation of our conclusions. Therefore, while the credible interval on a given day is valid in isolation of the other days, the probability that the difference signal is fully contained within the credible-interval band is not calibrated to a nominal probability level. Thus, our methodology provides a posterior probability distribution around the value of the difference signal for each day, rather than a posterior probability distribution for the difference signal as a whole. To approximate the sampling distribution of the difference signal for day 200 − *j*, we drew 20, 000 samples from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{w} $$\end{document}pjw to form the vector \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathfrak{w}}}_{k} $$\end{document}wk and 20, 000 samples from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\rm{p}}}_{j}^{p} $$\end{document}pjp to form the vector \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathfrak{p}}}_{k} $$\end{document}pk, pairing them to form the 4\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathcal{D}}}_{j}=\{{{\mathfrak{w}}}_{k}-{{\mathfrak{p}}}_{k}| {{\mathfrak{w}}}_{k} \sim {{\rm{p}}}_{j}^{w},{{\mathfrak{p}}}_{k} \sim {{\rm{p}}}_{j}^{p},k\in \{1,\ldots ,20000\}\}. $$\end{document}Dj={wk−pk∣wk~pjw,pk~pjp,k∈{1,…,20000}}. From \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathcal{D}}}_{j} $$\end{document}Dj, we derived a 95% credible interval (CI) for the difference signal at day 200 − *j* by taking the 2.5-th and 97.5-th percentiles of the sampled differences. These CIs provide a range of plausible values for the difference signal for each day, reflecting uncertainty due to sample size and variability in patient-level hospice utilization with respect to that day ignoring information known about any other day. ### Assessing clinical and statistical significance Statistical significance alone does not imply clinical relevance, and the distinction between the two is crucial for interpreting the applicability of results to clinical decision-making^35,36^. In order to impose a threshold representing a potentially impactful disparity in patient care, we defined a difference of at least 5 percentage points in hospice utilization as clinically significant. We believed this to be an appropriate threshold, though we recognize that, like the choice of 0.05 as the threshold for a statistically significant *p*-value, it is to some extent arbitrary. To determine the sensitivity of the results to this choice, we also ran the analysis with the threshold set at a difference in utilization of 3 and 7%. We computed two time series, **WMPCD** and **PMPCD**, with values defined for each day 200 − *j*, *j* ∈ [200], as **WMPCD**~*j*~, *The Probability of a Clinically Significant Difference with White Patients Receiving More Hospice*, defined as the probability that **diffSignal**~*j*~ > 0.05.**PMPCD**~*j*~, *The Probability of a Clinically Significant Difference with Patients of Color Receiving More Hospice*, defined as the probability that **diffSignal**~*j*~ < −0.05. These probabilities were derived directly from \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathcal{D}}}_{j} $$\end{document}Dj; **WMPCD**~*j*~ is the fraction of samples in \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathcal{D}}}_{j} $$\end{document}Dj exceeding the positive clinical significance threshold (0.05) and **PMPCD**~*j*~ is the fraction of samples in \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {{\mathcal{D}}}_{j} $$\end{document}Dj less than the negative clinical significance threshold (−0.05). We mark days where **WMPCD**~*j*~ > 0.95 as **WMPCD**(+) and days where **PMPCD**~*j*~ > 0.95 as **PMPCD**(+). If a day is either **WMPCD**(+) or **PMPCD**(+), we conclude that there is a high probability of a clinically significant difference on that day. Such days are marked as **PCD**(+)—having a *probable clinically significant difference*. ### Conventional significance tests To benchmark and contextualize our Bayesian, time-aware approach, we also applied four conventional statistical tests at the hospital level and one conventional statistical test adapted to application at the day **Fisher Exact Test** (**FE**) We compared the overall hospice utilization fraction by applying the Fisher Exact Test^37^ to the 2 × 2 contingency table comparing hospice utilization (received/not received) across racial groups (white/POC) over the entire 200-day period. The null hypothesis assumes equal 5\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \frac{{u}_{w}}{{n}_{w}}=\frac{{u}_{p}}{{n}_{p}} $$\end{document}uwnw=upnpA significant result (*p* < 0.05) is denoted as **FE**(+).**First Day of Hospice Test** (**FDH**) We applied a two-sided Mann–Whitney U test^38^ to the distributions of the first hospice day (*X*~*w*~ and *X*~*p*~) for those patients who received hospice. The null hypothesis is that these two distributions have the same central tendency. If *p* < 0.05, we label the result **FDH**(+). Unlike the Fisher Exact test, the First Day of Hospice test incorporates timing by comparing when (on average) patients entered hospice. Because this test only examines the timing of hospice for patients who received hospice, it would fail to detect a disparity where the rate of hospice utilization was different but the timing of care for those who did receive hospice was the same. This issue can be addressed by using both the Fisher Exact test and the Mann–Whitney test together.**Number of Hospice Days Test** (**NHD**) We applied a two-sided Mann–Whitney U test^38^ to the distributions of the number of days spent in hospice by white patients and patients of color (*D*~*w*~ and *D*~*p*~). The null hypothesis is that these two distributions have the same central tendency. If *p* < 0.05, we label the result **NHD**(+).**Day-by-Day Fisher Exact Test** (**DBDFE**) We performed a Fisher exact test for day 200 − *j* before death to compare6\documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \frac{\mathop{\sum }\nolimits_{i = 1}^{{n}_{w}}{{\bf{W}}}_{i,j}}{{n}_{w}}\,\,{\text{and}}\,\,\frac{\mathop{\sum }\nolimits_{i = 1}^{{n}_{p}}{{\bf{P}}}_{i,j}}{{n}_{p}}. $$\end{document}∑i=1nwWi,jnwand∑i=1npPi,jnp.A day with *p* < 0.05 is marked as **DBDFE**(+). This test serves as a frequentist analog to the Bayesian day-by-day assessment and can be compared and contrasted with the **PMPCD,** **WMPCD**, and **PCD** measures. Statistical adjustments were not made to reduce the type 1 error inflation created when using the **DBDFE** test to identify hospitals with a disparity rather than to identify patterns in the temporal pattern of the disparity. We chose not to make adjustments for multiple comparisons to the **DBDFE** for two (1) patient utilization is correlated across adjacent days so standard correction methods that assume independence may be overly conservative (2) the **DBDFE** approach is presented alongside our Bayesian, day-level analysis to illustrate how conventional frequentist tests compare rather than to provide a definitive, fully corrected measure of if a statistically significant disparity exists at the hospital level (the combined use of the **FDH** and **FE** tests are well suited to that task). While it may initially appear that we are juxtaposing fundamentally different approaches by comparing time-aggregated measures (**FE,** **FDH**, and **NHD**) with our day-level Bayesian and **DBDFE** analyses, we do so purposefully to emphasize the critical role of temporal information, as well as to present the reader with more than one option for implementing an analysis that takes into account the temporal dimension of care delivery. We wish, especially, to illustrate how time-agnostic analyses can inadvertently mask meaningful temporal patterns in hospice utilization. By contrasting the aggregate measures with those that provide day-by-day detail, it is possible to show that disparities in end-of-life care exist not only in *whether* hospice is utilized, but also in *when* it is initiated. ### Implementation of analysis All analyses were implemented in Python (v3.10) using standard scientific computing libraries. For each hospital, we The difference signal and its 95% credible intervals for each day.The **WMPCD,** **PMPCD**, and **PCD** series.The **FE,** **FDH,** **NHD**, and **DBDFE** *p*-values. In creating the visualization for this paper, we used a color palette based on those proposed by Okabe and Ito in their ‘Color Universal Design’ guidelines to ensure accessibility for individuals with color vision deficiencies.