Authors: Aline Scohy, Tina Lesnik, Brecht Devleesschauwer, Romana Haneef
Categories: Comment, Causes of death, Mortality, Ill-defined deaths, IDDs, Garbage codes, Redistribution, GBD
Source: Archives of Public Health
Authors: Aline Scohy, Tina Lesnik, Brecht Devleesschauwer, Romana Haneef
Causes of death (CoDs) statistics are an important source of information for epidemiological research. Some causes of death (CoDs) are not regarded as sufficiently specific causes of death for public health purposes, they are named “ill-defined deaths” (IDDs) or “garbage codes”. Redistribution is the process of reallocating IDDs to plausible underlying causes (i.e., target cause) as defined in the Global Burden of Disease (GBD) study. The main objectives of this commentary were to describe the various steps taken for the redistribution of IDDs by applying the Belgium redistribution approach to national mortality data sets and to underline key challenges with potential solutions.
A four-step probabilistic redistribution method was developed in the context of the Belgian national Burden of Disease (BeBOD) study including (1) predefined ICD codes, (2) package redistribution using multiple causes of death data (MCOD), (3) internal redistribution, and (4) redistribution to all causes. The Belgian, French and Slovenian mortality databases included a share of 34%, 36% and 20% of IDDs in 2017, respectively. The majority of the IDDs were redistributed using predefined ICD-10 codes (14%), followed by package redistribution using MCOD data (11%) in Belgium and French database, whereas this redistribution was 7% and 10% for Slovenia database, respectively. The main challenges encountered were lack of computational capacity to run 100 simulations and lack of MCOD.
This commentary highlighted the importance of the harmonized approach of redistribution that allowed us to get an in-depth knowledge of various steps applied for the redistribution process transparently and allowed comparing these results with other European countries. Further research is recommended considering creating a common/standardized consensus list of IDDs based on a comprehensive definition and aligned with the context of national mortality databases in European countries.
The online version contains supplementary material available at 10.1186/s13690-025-01652-x.
Text box 1. Contributions to the literature • There is limited evidence on the application of a harmonized approach to the redistribution of IDDs, for BoD analysis.• The method of redistributing IDDs, their definitions and variations in ICD-coding practices/systems, contribute to the ranking of leading causes of death.• A standardized consensus list of IDDs based on a comprehensive definition and tailored to the context of mortality databases of European countries, is needed.
Causes of death (CoDs) statistics are an important source of information for epidemiological research and can be used in public health policy decisions to develop interventions and prevention policies [1]. Some causes of death (CoDs) are not regarded as sufficiently specific causes of death for public health purposes, they are named “ill-defined deaths” (IDDs) or “garbage codes” [1]. The presence of IDDs limits the utility of death statistics, undermining their importance as a primary source of information for planning and assessing health policies and interventions [2]. Depending on the amount of IDDs, the CoD statistics may not accurately reflect a country’s mortality, hampering comparisons or leading to biased priorities [3, 4]. The official mortality statistics are based on the underlying cause of death, the disease or injury that started the chain of events leading to death [5]. To illustrate an accurate assessment of national disease burdens, it is essential to deal with IDDs.
Redistribution is the process of reallocating IDDs to plausible underlying causes (i.e., target cause) as defined in the Global Burden of Disease (GBD) study [6]. There are a number of methods to redistribute IDDs. One of the most well-known methods within the framework of GBD, used to redistribute IDDs, such as regression models, redistribution based on fixed proportions, proportional reassignment, and fractional assignment of death due to multiple causes of death [6]. In Australian Burden of Disease study, three methods were developed (1) direct evidence on more plausible causes of death from data linkage studies or other sources; (2) redistribution algorithms based on the distribution of underlying causes of death where the ill-defined cause was recorded as an associated cause of death and (3) reassignment of deaths across a specified range of target diseases according to patterns of causes of death observed in the mortality data for the ABDS disease list [7]. In Germany Burden of Disease study, the proportional redistribution method was applied [3].
General redistribution methods may generate biased estimates due to differential patterns in IDDs observed between and within countries and may lead to risk of misclassification, potentially distorting cause-of-death statistics [4, 8]. The GBD methodology used to redistribute the IDDs is complex, builds on several assumptions, which may not necessarily reflect the national context and is difficult to replicate [1]. The most recent redistribution method is developed in the context of the Belgian national Burden of Disease (BeBOD) study by the researchers from Sciensano [9]. The Belgian method of redistribution is more transparent to understand and can be adopted and applied to different European countries contexts. This Belgian approach is a four-step probabilistic redistribution method using (1) predefined ICD codes, (2) package redistribution using multiple causes of death data (MCOD), (3) internal redistribution, and (4) redistribution to all causes [9] (Fig. 1).
Fig. 1Flowchart to visually demonstrate the four-step IDD redistribution process
In the first step, the selected underlying CoDs of IDDs were proportionally redistributed to predefined ICD-10 target codes. For example, malignant neoplasm of the uterus, part unspecified (ICD-10 C55) is redistributed to two groups of target codes – malignant neoplasm of cervix uteri (ICD-10 C53) and of corpus uteri (ICD-10 C54) – pro rata to the occurrence of both diseases as underlying causes of death (Devleesschauwer B et al.). In the second step, packages of selected IDDs were created and MCOD were used to define targets and redistribute them. For example, the assigned underlying cause of death is “Unspecified kidney failure” (ICD-10 N19). This code cannot be assigned to a specific GBD cause (it is an IDD) and there are no predefined target codes available (step 1). For these cases, the defined packages, i.e., sets of IDDs that are similar and considered to have a similar redistribution target, were applied. In this example, the package created is called “Acute kidney failure” and is made of the following ill-defined ICD-10 N19, N17.0, and N17.9 [9]. In the third step, the internal redistribution of IDDs, when one or more specific codes were present as MCOD on the death certificate, one of these specific codes was selected as the target code. For example, « Sequelae of inflammatory diseases of central nervous system (ICD-10 G09) » is an IDD and the corresponding MCODs are CoD1, CoD2, CoD3 (specific CoD: Encephalitis, myelitis and ICD-10 G04), was used as a target code. In the final step, all remaining IDDs were proportionally redistributed to all specific CoDs. For example, respiratory arrest (ICD-10 R092) was redistributed to all specific CoDs. All these redistribution steps respected the proportions observed in the specific CoDs stratified by six age groups (0–4, 5–14, 15–44, 45–64, 65–84, 85+) and sex.
We selected the Belgian approach due to two main first, under the burden-eu (European Burden of Disease [BoD]) network, technical support was provided for carrying out BoD studies by the BoD experts. As Sciensano researchers have developed this approach, we carried out this study under STSM (short term scientific mission), to apply a method that has already been applied and tested by a country. As redistribution is complex process, this scientific collaboration allowed us in-depth understanding of this process and to learn from each other. Second, similarities in mortality databases and ICD-10 coding practices among these countries facilitated the adoption of this approach to other European countries.
The main objectives of this comment were to describe the various steps taken for the redistribution of IDDs by applying the Belgium redistribution approach to national mortality data sets and to underline key challenges with potential solutions.
Before the redistribution process of IDDs, we undertook several steps of data preparation and adjustments of the algorithm based on the available mortality database.
We obtained the GBD mapping file and the list of IDD from the GBD 2019 study (2019 was the most recent available data file at the time of the study), to map these data files on the available pre-pandemic mortality database for each country. We followed several steps to prepare these data First, we created a data file containing individual-level mortality data including age, sex, year, place of residence, place of death and ICD-10 codes. Second, we mapped every ICD-10 code of the underlying CoD of each individual to the GBD 2019 cause list and identified the IDDs in each country-specific mortality data. Third, we created a separate list of all ICD-10 codes of the IDDs and defined for each code the corresponding target codes with the redistribution steps. Fourth, we developed the master cause list for the respective country, which classified all specific causes of death into the GBD cause levels (i.e., 1, 2, 3, 4). We made sure that each cause of death in the national mortality database could be matched with the three reference data files (i.e., GBD cause list, master cause list and IDDs list if relevant) before the redistribution process.
After preparing the required data files, we replicated the Belgium redistribution approach to the country-specific mortality data sets of France and Slovenia. We made adjustments in the algorithm to redistribute the IDDs according to the respective mortality databases. For instance, in Slovenia’s mortality database, there are no MCODs. Therefore, in the package redistribution, the target proportions were defined based on Belgian MCODs and the algorithm was modified accordingly. Whereas for France, the algorithm was updated to account for a maximum of 35 MCODs per person. The internal redistribution step was not applied to the Slovenian mortality database as during the process of coding the underlying cause of death, ambiguous cases were investigated and if needed linked to the hospitalization database to determine the true cause of death.
The commentary was based on anonymous and aggregated data and does not require the ethics approval and consent to participate.
included a share of 34%, 36% and 20% of IDDs in 2017, respectively (Fig. 2). The majority of the IDDs were redistributed using predefined ICD-10 codes (14%), followed by package redistribution using MCOD data (11%) in Belgium and French database, whereas this redistribution was 7% and 10% for Slovenia database, respectively. In Belgium, France and Slovenia mortality databases, the most common IDDs as underlying causes were I50.9 (Heart failure, unspecified), followed by J18.9 (Pneumonia, unspecified), I64 (stroke, not specified as haemorrhage or infarction), and R99 (Ill-defined and unknown cause of mortality). The most frequent IDD (i.e., 20) are enlisted in the Supplementary material 1: Additional file 1, corresponding to their share to total number of deaths and to the redistribution steps applied.
Fig. 2Comparing the proportion of ill-defined deaths (IDDs) across Belgium, France and Slovenia in 2017
After the redistribution process, we compared the age-standardized mortality rates (ASMRs) by ranking approximately 30 of the most important causes of death across the three countries (Supplementary material 1: Additional file 2). In general, the ASMRs between Belgium and France were more similar to each other than to Slovenia. However, notable differences were observed. For example, hypertensive heart disease was ranked 5th in Slovenia, 21st in France, and 32nd in Belgium. Chronic obstructive pulmonary disease ranked 6th in Belgium, 9th in Slovenia, and 12th in France. Atrial fibrillation and flutter ranked 11th in Belgium, 13th in France, and 27th in Slovenia. Alcohol use disorders were ranked 13th in Slovenia, 33rd in Belgium, and 37th in France. Cardiomyopathy and myocarditis ranked 19th in Belgium, 36th in Slovenia, and above the 50th position in France (Supplementary material 1: Additional file 2).
There were two main challenges applying this approach to the French and Slovenia mortality data sets. First, the probabilistic nature of the redistribution method requires increased computational capacity, which was limited by the portal of the French national health database (i.e., SNDS [Système National des Données de Santé]) to run 100 simulations. Therefore, we ran the highest number of iterations (i.e., 30) that passed successfully through the SNDS system and took into account only one full year. Other alternative computational strategies could be to adopt the algorithm that converge faster or are more lightweight to be more efficient or to break down the dataset into small and equal parts, using machine learning approach. Second was the lack of MCOD in Slovenia’s mortality database. It would be difficult to distinguish the effect of not having MCOD vs. the effect of having fewer IDDs in Slovenia, for example. One way to assess this is by examining the percentage of IDDs redistributed at the last stage (ALL). One would expect this percentage to be higher if the MCOD step was skipped. However, due to lower number of IDDs, this is not the case. It is expected to introduce the MCOD in Slovenia mortality in the near future.
This exercise highlights the critical role of the redistribution method, emphasizing its interdependence with ICD coding practices and the definitions of IDD codes.
This redistribution method is transparent and can be replicated and adopted to apply to other national contexts, depending on the available mortality data. This approach allows using all available data on multiple causes of death, to redistribute IDDs by 6 age groups and sex and validated by burden of disease experts [9]. By applying a probabilistic redistribution approach, the resulting uncertainty intervals were relatively narrow, indicating higher precision in the estimates. After redistribution of IDDs, we observed a few major shifts in leading causes of death among three Belgium: Lower respiratory infection ranked 28th (before redistribution) to 5th position (after redistribution), and stroke ranked 6th (before redistribution) to 3rd position (after redistribution). France: Lower respiratory infection ranked 22nd (before redistribution) to 6th position (after redistribution), and other cardiovascular and circulatory diseases ranked 26nd (before redistribution) to 8th position (after redistribution). Slovenia: Lower respiratory infections ranked 31st (before redistribution) to 4th position (after redistribution), atrial fibrillation and flutter ranked lower than 50th (before redistribution) to 27th position (after redistribution) and cardiomyopathy and myocarditis ranked 43rd (before redistribution) to 36th position (after redistribution). These shifts could be due to narrow definitions of IDDs.
This approach allows us to understand the in-depth knowledge of various steps applied to the redistribution process in a clear and transparent way. These steps ensured that the IDDs were accurately redistributed to specific causes of death and that the targets were updated at the end of each step, which decreases the chance of over- or under-represented causes of death in the final ranking [9].
The ICD (International Classification of Diseases) coding practices can significantly vary from country to country. For example, variations in interpretation of causes of death, updates/revisions in ICD versions, which may add or reclassified into more specific categories, some coding systems may choose more specific codes while others might use a broader category, some coding systems may emphasize on contributing cause while others may focus on the underlying cause of death, presence of a centralized coding system for underlying causes of death (i.e., Slovenia, with a low proportion of IDD). Overall, all these factors should be considered in the context of the identification of IDDs, which may result in differences in the ranking of leading causes of death between regions, countries, and over time, even when mortality rates are similar.
The criteria used to define what constitutes an ill-defined death is important and it can vary between countries or regions, affecting the number of deaths categorized under IDD. At present, only GBD updates the list of IDDs at the cycle of each GBD study. The definition of IDD could have a narrow or broad definition. A narrow definition may result in fewer deaths being classified as IDD, leaving more deaths under specific causes of death and creating the appearance of more accuracy to mortality statistics. Nevertheless, a narrow definition may underreport issues by partially excluding specific death that lack clarity such as cardiac arrest with no underlying cause. Conversely, a broad definition might increase the number of IDDs and could highlight the full extent of poor reporting that requires more extensive redistribution, which may risk losing important information in the national context. Using broad definitions can dilute specificity by grouping together certain deaths that are unclear and may inflate the proportions of unclear deaths. The differences in how IDDs are defined, thus, redistributed to specific causes of death (i.e., target codes), can lead to significant variations in the rankings of final causes of death (after redistribution), which can affect trend analyses and health policy decisions.
This commentary has some first, we did not explore the existing differences in IDDs among the three countries, which might be due to different coding practices/systems. Further research is needed to explore that. Second, we did not compare results of this model with other models or independent datasets, to explore the impact of redistribution methods. More research is needed to compare different models used for redistribution of IDDs by using a single dataset. Moreover, a scoping review is needed to compare all the available redistribution methods, highlight the strengths and limitations of each model.
We propose a task force under the supervision of the European Commission/Eurostat could be established, to develop a common/standardized consensus list of IDDs, aligned with the mortality databases in European countries. This task force should include representatives from national statistical offices, mortality data experts and public health institutes from European countries. This task force should agree on a comprehensive definition of IDDs, harmonized list of IDDs validated by the task force, integration of the standardized list into the national mortality database and reporting systems, and to review this list periodically to update the changes. This approach would support comparability, transparency, and data quality in European mortality statistics, especially important for cross-country analysis, burden of disease studies, and health policy planning. This work has important implications to future BoD analyses at national level in terms of transparency of redistribution approach, in-depth understanding of local mortality data set, highlight data gaps and to improve further the quality of data reporting.
Redistribution of IDDs is necessary for the burden of disease analyses to avoid the over- or under-reporting of certain CoDs, which allows for an accurate assessment of national disease burdens and to facilitate the development of targeted interventions. This harmonized approach allows us to understand the in-depth knowledge of various steps applied for the redistribution process transparently and allows comparing these results with other countries. These findings underline the importance of the redistribution method, the definition of IDD and the variations in ICD-coding practices/systems, which can significantly shift the ranking of leading causes of death. We recommend considering creating a common/standardized consensus list of IDDs based on a comprehensive definition and aligned with the context of mortality databases in European countries.
Below is the link to the electronic supplementary material.
Supplementary Material 1: Additional file 1: It is an Excel file, presenting the most frequent IDDs in the three countries corresponding to the redistribution steps applied. Additional file 2: It is an Excel file, presenting the age-standardized mortality rates (ASMRs) after the redistribution process, across three countries.