Authors: Geerthy Thambiraj, Antonis A. Armoundas
Categories: Perspective, Epidemiology
Source: Communications Medicine
Authors: Geerthy Thambiraj, Antonis A. Armoundas
In clinical research, bias is a systematic error that creates a difference between observed and true values. The increasing use of large datasets and artificial intelligence (AI) in medicine necessitates a renewed focus on how such errors can be introduced and propagated.This educational primer provides a consolidated framework of the three primary types of bias selection, information, and confounding/analytical. We synthesize these concepts for an interdisciplinary audience of clinicians and data scientists, using illustrative examples from both traditional clinical trials and modern, data-intensive research. Bias can arise at every stage of the research design, conduct, analysis, reporting, and dissemination. We illustrate how classic issues, such as selection bias (systematic differences between those included and those eligible/targeted, thus distorting effect estimates), manifest in data and how modern analytical methods can introduce novel forms of error if not carefully managed. A shared understanding of bias is essential for effective collaboration between clinical and data science teams. This primer offers a practical conceptual map to help these teams proactively identify, mitigate, and transparently report on potential sources of bias, ultimately fostering more robust and equitable clinical evidence.
In clinical research, bias is defined as a systematic error whereby there is a consistent or proportional difference between the observed and “true values” that the research is trying to detect^1^. In the presence of systematic error, measurements may be precise, but their average is far from the “truth”; consequently, even an infinite number of observations will not converge on the true value^2^. This differs from random error, where measurements lack precision but their average approaches the “truth” as the number of observations increases (Fig. 1). In this educational primer, we define bias as a form of systematic error. We distinguish this from random error, which introduces imprecision into measurements. In statistical terms, this distinction relates to the difference between bias (the difference between the average of a measurement and the true value) and variance (the spread of measurements around their average). For an estimate of a true value, bias refers to systematic error, how far the estimate tends to be, on average, from the truth (average estimate minus true value). Variance reflects random scatter, how much the estimate varies from one sample or measurement to the next. Precision is the opposite of that higher precision means less variability (often described as the inverse of the variance).Fig. 1Examples of systematic and random error.The figure illustrates the difference between systematic error, where measurements may be precise but consistently differ from the true value, and random error, where measurements vary around the true value.
While these concepts are foundational to all experimental and observational studies, the rapid integration of big data and artificial intelligence (AI) into medicine has introduced new complexities and amplified the potential impact of bias^3^, which can affect a health study at all stages (e.g., design, conduct, analysis, reporting, interpretation, and dissemination). Failure to recognize and address these biases can lead to flawed study findings, inequitable algorithmic performance, and compromised clinical outcomes^4,5^.
The objective of this educational primer is therefore to provide a clear and accessible overview of the most common forms of bias for an interdisciplinary audience of clinicians and data scientists. Rather than an exhaustive systematic review, we synthesize the three canonical types of bias, selection, information, and confounding/analytical^2^, and illustrate how they manifest across each stage of the research lifecycle, from design to dissemination. By using examples from both traditional studies and modern data-driven health research, we aim to create a shared conceptual framework to help collaborative teams avoid, detect, and mitigate bias in their work.
This primer is organized along the research lifecycle (design, conduct, analysis, reporting, and interpretation/dissemination). Within each stage, we discuss how similar underlying bias mechanisms can arise across three common research formats, randomized controlled trials, analytical epidemiologic studies, and AI/data science studies, highlighting format-specific examples and practical implications. Figure 1 illustrates the conceptual distinction between systematic error (bias) and random error (imprecision). Table 1 serves as a brief glossary of the three canonical bias categories used throughout (selection, information, and confounding/analytical). Table 2 provides a quick-reference catalog of common sources and manifestations of bias; it is intended to be used as an index and checklist alongside the narrative, helping readers quickly locate where each bias appears in the text and apply the same lens when planning, appraising, or reporting studies. Although we often describe bias as arising from investigator decisions, many safeguards, and many failure modes, also sit outside the investigative team. Ethical review bodies, independent oversight committees, regulators, and journals shape what data are collected, which analyses are considered confirmatory, and what ultimately enters (or fails to enter) the scientific record; accordingly, we briefly summarize these actors’ roles in both mitigating and amplifying bias (Table 3).Table 1Overview of the main types of bias in clinical researchType of biasDefinitionSelectionThis occurs when participants, patients or groups in a study differ systematically from the population of interest which leads to a systematic error in an identified association or outcome.InformationThis occurs when the approach that is utilized to collect or confirm study measurements is sub-optimal. This can lead to misclassification of participants, patients or outcomes and erroneous conclusionsConfounding/AnalyticalThis is a distortion of the true relationship of an exposure with an outcome via the “mixing of effects”. Essentially the effects of the exposure under study on a given outcome are mixed in with the effects of an additional factor (or set of factors) that leads to a spurious relationship.Table 2Sources and manifestations of bias in clinical researchDefinitionSelection biasAscertainmentOccurs when some members of the target population are more likely to be included in the sample than others.Healthy worker effectIndividuals in paid work usually exhibit lower overall death rates than the general population because the severely and chronically ill are usually not fit enough to work.Loss to follow-upIf the loss to follow-up is related to exposures, confounders and outcomes of interest, it may bias association estimates.Non-responseormissing dataCan occur when subjects refuse to take part in a study, or who drop out before the study can be completed (resulting in missing data), form cohorts that are systematically different from those who participate.Information biasRecallOccurs when there are systematic differences in the accuracy or completeness of the recollections retrieved by individuals in the study.Regression dilutionOccurs when random measurement errors in the exposure lead to an underestimation of the observed association with an outcome^111,112^Selective reportingSelective reporting (including selective publication/publication bias): occurs when investigators ‘cherry-pick’ which findings to report or publish in order to suppress negative or undesirable findings.SpinOccurs when the interpretation of study results are distorted (either intentionally or unintentionally.Sources of analytical and confounding biasError in statistical analysis codeA statistical programming error (e.g. incorrectly specified model, missing covariate), can lead to biased results if undetected.Fishing or multiple testingOccurs when numerous statistical tests are conducted in order to ascertain statistically significant findings.Reverse causality biasOccurs when the direction of cause-and-effect association is unclear and can lead to erroneous conclusions.Residual confoundingOccurs when there is incomplete adjustment for confounding (e.g. missing confounder, or confounder poorly measured) that leads to biased associations.Table 3Stakeholders in bias across the research ecosystem (roles in detecting and introducing bias)DescriptionInstitutional Review Boards (IRBs)/Ethics Committees^133^ReasonWell-intended restrictions (e.g., overly conservative eligibility constraints, exclusion of “vulnerable” groups rather than protected inclusion) can reduce representativeness and exacerbate under-inclusion of key populations (selection bias).Detection/MitigationEvaluate whether the selection of subjects is equitable, whether risks are minimized, whether consent/waiver processes are appropriate, and decisions that directly influence selection bias and generalizability.Clinical Endpoint Adjudication CommitteesReasonBias can occur if adjudication is not blinded where feasible, if endpoint definitions drift, if documentation availability differs by site/arm, or if adjudicators are inconsistently calibrated.Detection/MitigationIndependent, standardized adjudication can reduce information bias/misclassification for subjective or complex endpoints (e.g., stroke, MI) and improve cross-site consistency.DSMB/DMCReasonInterim monitoring decisions (e.g., early stopping for benefit) can inflate effect estimates or alter the evidentiary picture; leakage of interim results can distort conduct and analysis.Detection/MitigationProvide independent oversight of safety, efficacy, and data integrity; confidentiality around interim observations helps prevent operational bias (changes in care, enrollment, and ascertainment driven by emerging results).RegulatorsReasonEvidentiary incentives can unintentionally encourage reliance on surrogate endpoints or narrow trial populations, and can shape what gets measured/validated (with downstream implications for observational and AI datasets that reuse those measures).Detection/MitigationSet evidentiary standards for approval/labeling (including acceptable endpoints, pre-specification expectations, and post-market obligations), which can reinforce rigor and transparency.Journal Editors, Peer Reviewers, and PublishersReasonEditorial and reviewer incentives (novelty, “positive” findings), conflicts of interest, space constraints, and inconsistent enforcement can contribute to publication bias, selective reporting, and “spin.”Detection/MitigationRequire trial registration/protocol transparency, enforce reporting standards, commission statistical review, enable corrections/errata, and tools that reduce selective reporting and improve reproducibility.DSMB Data and Safety Monitoring Boards, DMC Data Monitoring Committees.
While our primary definition of bias refers to systematic error, the presented examples also include actions, conditions, or poor practices that can introduce or amplify such errors throughout the clinical research process.
Causal inference often aims to employ the findings of a study to increase understanding of a particular target population (Table 1). The true underlying causal effect typically depends on the target population definition. Thus, when determining a causal effect, it is important to define the target population to which the results are intended to be generalized. Randomized and observational studies (Table 2) have their respective advantages and disadvantages in estimating causal effects, in a target population; in a randomized study, estimates may have internal validity, but often they are not representative of the target population; in observational studies, estimates may better reflect the target population, and thus they are more likely to have external validation, however they are subject to potential bias due to non-measured confounding (a variable that is not included in an experiment, yet affects the relationship between two variables of the experiment)^6^.
Beyond reducing confounding through random allocation, randomized clinical trials (RCTs) often incorporate prospective safeguards that improve interpretability and reduce selective analysis/reporting. These include pre-registration and protocol transparency, explicit pre-specification of primary hypotheses and endpoints, and a statistical analysis plan (SAP) finalized before database lock and unblinding that defines the primary analysis (e.g., intention-to-treat), handling of missing data, and any multiplicity adjustments. Together, these practices limit post hoc ‘data-picking’ and help readers distinguish confirmatory primary analyses from exploratory analyses. Importantly, these safeguards do not eliminate all threats to validity (e.g., restrictive eligibility affecting generalizability, loss to follow-up, or bias from lack of blinding), but they create a clearer, auditable chain from question to protocol, analysis, and interpretation^7–10^.
Across randomized trials, analytical epidemiologic studies, and AI/data science research, a common design vulnerability is the use of surrogate endpoints or proxy outcome labels, which are measures that are expected to reflect clinical benefit but are not themselves direct measures of how patients feel, function, or survive^11^. In drug development and regulatory review, surrogate endpoints can be acceptable, but their use requires careful justification and (when possible) formal evaluation of whether changes in the surrogate reliably predict changes in the true clinical outcome; classic statistical frameworks for surrogate validation and meta-analytic approaches have been proposed for this purpose^11^. This issue can be even more problematic in epidemiologic and AI-based studies that rely on EHR/administrative data, where ‘endpoints’ are often operationally defined surrogates (e.g., billing codes, lab thresholds, medication orders, utilization events) created for clinical, legal, or reimbursement workflows rather than research. Such surrogates may reflect care processes and documentation practices as much as underlying disease biology, and they may not generalize across sites or time periods.
Selection bias, where the study population does not represent the target population, can occur at the design stage of a clinical research study. This is most likely to occur when non-random sampling is utilized, which then results in a non-representative sample, from which the derived effect size estimate (e.g., treatment effect or risk metric) is different from the intended target population (Table 2)^12^.
The choice of method of ascertainment of study participants is extremely important, and selection bias can permeate different types of studies. Ascertainment bias may occur in any kind of study design. It occurs when the cohort of the recruited patients into the study does not sufficiently represent the target population^13^. For example, survivor bias can occur if a significant cognitive pitfall, defined by the logical error of concentrating on the entities that have successfully passed a selection process while ignoring those that did not, often due to the latter’s lack of visibility; the core issue lies in the incompleteness of the data derived solely from survivors^14^. This bias can occur in both cross-sectional and case-control studies. In occupational studies, a sub-optimal definition of the population (that may be unavoidable) often occurs and leads to the “healthy worker effect”. The “healthy worker effect” manifests itself when a lower mortality rate is observed in the selected population, compared with the general population (Table 2). Selection bias can occur in a prospective cohort study, e.g., when individuals, groups, or data are lost to follow-up, resulting in a sample that does not have the same probability of manifesting an outcome of interest when compared to the sample that is no longer under observation, leading to incomplete data and potentially skewed or inaccurate conclusions. Missing data and loss to follow-up are not merely analytic inconveniences; they can create selection bias by changing who remains under observation and, therefore, the effective target population (Table 2). Because the likelihood and patterns of missingness are often influenced by design choices (e.g., participant burden, visit frequency, and how outcomes are ascertained), missing-data mitigation should be planned upfront. At the design stage, investigators should (1) minimize follow-up burden and pre-specify retention procedures, (2) incorporate alternative outcome-ascertainment pathways (e.g., passive follow-up via medical records/registries when appropriate), and (3) pre-specify how missingness will be monitored and handled (including sensitivity analyses) so that the primary inference is robust to plausible missing-data mechanisms^6,15–17^.
Administrative data or other electronic health records (EHR) offer rich information that has been exploited to describe the manifestation of complex disease phenotypes, understand their origins, and assess potential treatments. However, since these data are primarily collected for operative reasons, such as planning and monitoring, or for legal purposes^18^, the processes of generating and linking them may become sources of bias that are not often sufficiently taken into consideration, but can be detected by researchers and clinicians^19^. Decisions involving whose data are collected, which information is recorded, and how it is coded, are primarily motivated by the operational objectives of the data collection processes^20^. While the model developers’ or analysts’ capacity to affect such processes is narrow, these, as well as other data processing steps, have to be considered, if some forms of bias are to be mitigated. Furthermore, depending on the definition of the analysis samples, recording and registration processes may also introduce selection, information, and other forms of bias^21^.
Selection bias can occur in big data analytical analyses and arises when the data employed for AI algorithm development^22^ is based on non-representative (of the entire population) patient cohorts for which the algorithm is intended to be used^23^. In predictive AI models, there is potential to accentuate health disparities through biases (especially selection and information) in the data training of the algorithm, and in the algorithm’s intent^4,5,18,24,25^.
Older people, women, and ethnic minorities have been under-represented in research, and the greater part of contemporary evidence-based medicine may not apply to these target populations (Table 2). This may lead to differences in management practices. In a systematic review of 207 clinical trials, from 2001 to 2018, there has been observed consistent under-reporting of Black and female patients^26^. In a few notable examples, although the same proportion of men and women presented with chest pain, men were 2.5 times more likely than women to be referred to a cardiologist for further management^27^; also, in the emergency room, Black patients were 40% less likely than White patients to receive pain medications^28^. These disparities exist in large-scale EHRs that have been used to train AI algorithms.
Information bias (systematic measurement error or misclassification of exposure, outcome, or covariates), or alternatively as the so-called misclassification, is one of the most common sources of bias that affects the legitimacy and impact of health-related research (Table 1). It stems from the methodology that is employed to acquire or validate the measurements of a study during the conduct of the study. These measurements can be obtained by experimentation (e.g., biological assays) or observation (e.g., questionnaires or surveys). Given that there are no perfect tools to collect data, the majority of the studies must consider a degree of misclassification, and this may lead to bias^29^. There are two main types of misclassification, differential (when the error of the assignment to the wrong category depends on the exposure/outcome), and non-differential (Table 2). Differential misclassification bias takes place when the misclassification is different among the compared groups (e.g., a case-control study, in which the recalled exposure is not the same for controls and cases). As a consequence, the estimated effect may be biased either towards the null or away from the null^29^. Non-differential misclassification bias manifests when the misclassification is the same across the compared groups, e.g., the exposure is equally misclassified in both controls and cases. While for binary variables the estimate is biased towards the null value, for variables with more than two categories, the direction of the bias is unpredictable^29^.
Regression dilution bias (attenuation of associations from random error in continuous exposures) is produced in prospective studies that are interested in the association between a baseline exposure (Table 2), such as a continuous variable like blood pressure (BP), with an incident disease outcome, such as stroke. Baseline BP measurements randomly fluctuate for two variability in the measuring process itself, as well as transient variability in determining the baseline from their usual BP level^30^. These processes result in underestimating the actual exposure-outcome relationship by including more people than they should in extreme categories (the bottom baseline BP categories include more subjects whose BP level is lower than their usual BP, while the top BP categories include more subjects with higher baseline BP than their usual BP level)^31^.
Information bias can manifest in several ways. Given that self-reporting is a common data–collection approach in epidemiologic and medical research, occasionally, study participants may erroneously provide responses that depend on their recollection of past events. Another form of information bias is the interviewer-dependent bias, where knowledge of the study hypothesis, or the disease or the exposure status (including the subject-received intervention), may influence the recorded data^32^. It is also possible that interviewers can introduce error by not adhering to a structured interview or by influencing the subjects in different ways (such as gestures or leading questions)^33^.
Other potential ways that information bias can affect observed measurement are through device inaccuracy, errors in data entry, digit preference by the researcher (e.g., rounding to 5 or 0), or environmental conditions in the laboratory^34^. An example in clinical research involves the inaccurate measurements obtained using frequently employed devices such as thermometers, pulse-oximeters, and sphygmomanometers. A subject’s skin tone is known to affect pulse-oximeters’ measurements, which may result in systematic device overestimation of oxygen saturation levels in non-White patients^35^.
Information bias can also be due to the omission of important data (e.g., in under-represented populations) during the data collection (Table 2), something that can impact AI algorithm development and result in algorithmic bias. For example, Canada and France, which are countries that do not record race and ethnicity in their national health databases, make it difficult to take into account underrepresented populations that may exhibit different outcomes in comparison to the general population^36^. Race-based bias in clinical research impacts a study’s relevancy, validity, and reliability, and there are now processes in place that aim to develop critical appraisal tools that take into consideration racial bias when assessing the quality of a paper^37^.
Bias can be introduced during the data analysis by probing and analyzing the data in ways that favor the conclusions of a particular hypothesis (Table 2). Some poor practices, such as removal of data that do not support the main hypothesis (outliers, or even whole subgroups), use of inappropriate statistical tests, performing multiple testing (“fishing” for associations) by pair-wise comparisons, evaluation of multiple endpoints, and performing data-driven secondary or subgroup analyses in order “to find” a statistically significant difference, can bias the results^38^. Collectively, these practices are often referred to as p-hacking; trying many reasonable analytic variations (e.g., different outcomes, subgroups, covariate sets, or model specifications) and preferentially highlighting the analyses that yield ‘statistically significant’ p values^39^. It has been highlighted that the high rate of non-replication (lack of confirmation) of some research findings is a consequence of claiming conclusive evidence solely based on a single study assessed by formal statistical significance, typically for a p value less than 0.05^40^. Because p values are frequently misinterpreted, especially outside statistical disciplines, the American Statistical Association (ASA) issued a formal statement emphasizing that a p value does not measure the probability that a hypothesis is true (or that results occurred “by chance” alone), and it does not convey the magnitude or clinical importance of an effect. The ASA further cautions against treating p < 0.05 as a bright-line rule for ‘proof’ and recommends interpreting p values in context, alongside study design, data quality, model assumptions, effect sizes, and measures of uncertainty^41^.
In confirmatory RCTs, pre-specifying the primary endpoint/hypothesis and the primary analysis in the protocol/SAP (and finalizing the SAP before database lock/unblinding) is a key countermeasure against selective analyses and helps ensure that multiplicity is handled transparently rather than post hoc^9,42^.
Problems associated with multiple testing may manifest in types of analysis involving the comparison of treatment groups with multiple outcomes, or when patients have several measurements of an outcome over time^43^. In an editorial, Nature^44^ has called for principal investigators to take responsibility for the quality of the data they publish and their reproducibility, as well as for the need to run small-sized laboratories in order to make that possible, addressing these issues from the side of “sloppiness”, rather than being a systemic flaw in the structure of research and publication processes^45^.
Bias due to confounding (Table 2) in observational studies of dietary factors with incident disease, such as cancer^46^ and cardiovascular disease^47^, has led to controversy. These observational associations have been strongly criticized and their credibility questioned due to extensive residual confounding as well as inadequate measures of dietary exposures^48^, which can contribute to biased analyses. For readers seeking more detailed treatment of confounding/residual confounding and causal inference (including practical guidance on design/analysis choices), see Hernán and Robins^49^ and Modern Epidemiology^50^; for perspectives on the specific challenges of nutritional epidemiology, see Ioannidis and Ludwig et al.^51,52^.
Missing data often occur in clinical studies, and there are many different reasons why data may be missing (Table 2)^15^. Understanding and recognizing these patterns of missing data, as well as their associated implications, is important when evaluating whether a dataset is suitable for a particular use case. Failure to appropriately account for missing data in analyses can lead to biased results, particularly if the missingness patterns are complex. For example, subjects who experienced severe adverse events may have dropped out, resulting in a dataset that consists of patients who did not experience severe side effects, which alters the hypothetical target population, resulting in algorithmic bias if not accounted for^6^.
There have been instances of statistical programming errors in the analysis that led to biased results. For example, in a published paper on a widely used cardiovascular risk score in the UK^53^, the authors found that the ratio of total to HDL-cholesterol appeared to be of no relevance to the incidence of cardiovascular disease^53^. This anomaly was pointed out by several other researchers in the letters to the editor of the journal, and it was subsequently found that the authors had imputed data for cholesterol values in the model but had implemented the multiple imputation strategy incorrectly, which had led to the ratio of total to HDL-cholesterol being excluded from the risk score^54^.
Beyond outright programming mistakes, the reliability of the software ecosystem itself can affect results, e.g., bugs in statistical packages, numerical libraries, or unintended changes introduced by software updates and differing default settings. For this reason, good practice increasingly includes sharing the analysis/simulation code, documenting key software versions and dependencies, and using independent re-execution (e.g., code review, rerunning analyses from raw data, or reproducing key results with an alternative implementation when feasible). Reflecting this shift, some statistical/methodological journals now require submission of code (and data when possible) and have implemented reproducibility checks, including re-running code for simulation studies (e.g., Biometrical Journal)^55^.
Reverse causality bias can also commonly occur when there is a bi-directional association between the exposure and the disease outcome. A recent article that investigated a selection of risk factors and behaviors in cardiovascular epidemiological studies concluded that it is more prevalent than was previously thought^56^.
Finally, although uncommon, research misconduct, including data fabrication or falsification (‘data fraud’), can also produce biased or entirely spurious study findings and should be recognized as a potential threat to validity in clinical trials^57^.
Selective reporting of results, whereby a researcher deliberately or partially reports the findings of their study, can lead to bias (Table 2). This can lead to over- or under-estimation of treatment effects or harms in a clinical trial, which can lead to mismanagement of a patient. A recent review article has suggested that biased under-publication and over-publication of research are forms of unscientific and unethical misconduct and lead to avoidable research waste^58^.
A related and longstanding concern is the ‘file drawer problem’, in which studies with null or negative findings are less likely to be written up or accepted for publication. This selective non-publication can skew the published literature toward statistically significant results and lead to overestimation of effects in narrative reviews and meta-analyses^59^. Also, a closely related threat is publication bias, whereby studies with statistically significant or ‘positive’ findings are more likely to be published (and published faster) than studies with null or negative results. This can skew the apparent balance of evidence in the literature and lead to overestimation of effects in narrative reviews and meta-analyses (see Reporting recommendations for approaches).
A high-profile example of selective reporting is the paper on the measles, mumps and rubella (MMR) vaccine published by Wakefield, which has since been retracted but led to many parents being reluctant to vaccinate their children due to fear of potential side effects^60^.
Some researchers may be tempted to interpret the outcomes of a study in ways that confirm their prior assumptions or beliefs, whilst discounting alternative justifications or contradictory indications (Tables 1 and 2). A well-known aspect of this is “spinning” (selective framing/interpretation that overstates the strength or direction of findings) of the results. “Spin” is essentially the misinterpretation of study findings, for example, presenting the outcome of a study in a more positive manner than the actual results reflect, or alternatively, downplaying harms, and is more widespread in the biomedical literature than previously realized^61^. The long-term consequences of “spin” are that it can have a negative, if not harmful, impact on advancing the practice of healthcare and population health, adversely influencing the planning of health policy, contributing to research money waste, lowering the chances of reproducibility of the findings of a particular study, and reducing the return-on-investment from research^61^. For example, a systematic review of “spin” in radiological studies has found that of the 126 articles included, 39 contained some form of over-interpretation, including 29 with an overly optimistic abstract and 10 with discordance between the study aim and conclusion, and a further eight with conclusions based on selected subgroups^62^. A similar type of review of cardiovascular trials found that for reports with non-statistically significant primary outcomes, the investigators often manipulated the language in order to deflect from the neutral primary outcomes^63^.
As research is now widely disseminated not only in journal publications but also in the lay press and social media, there is even more potential for misinterpretation (intentionally or non-intentionally) of the findings^64^. Bias in the interpretation of study results could be due to undisclosed conflicts of interest (either financial^65^ or non-financial^66^). Both of these types of conflicts of interest can affect the peer review process and influence the way the study results are interpreted and disseminated. A recent review noted that consumers of research should be aware that peer review cannot prevent the use of misleading language in published reports^63^.
There are well-established procedures to avoid bias when designing RCTs, such as proper randomization, blinding of both researchers and participants, intention-to-treat analyses, and avoiding overemphasizing data-driven subgroup analysis results^67^.
In the UK, the Innovations in Clinical Trial Design and Delivery for the Under-served (INCLUDE) Ethnicity Framework has been developed in order to improve the design of trials that adequately reflect the communities that the research intends to support and serve^68^.
For analytical epidemiological studies (e.g., case-control and cohort studies), there are more issues to be concerned about at the design stage due to a lack of randomization. The two key criteria that investigators should adhere to when designing case-control studies (1) there must be explicit diagnostic criteria for cases (including eligibility criteria used for case selection); and (2) controls should come from the same “source population” as the cases, and their selection should be independent of the exposure or exposures of interest^32^.
Generally, it is preferred to use incident (new cases) rather than prevalent (existing cases) in a case-control study to minimize the risk of prevalent cases changing their behavior, or the onset of disease itself may have altered the exposures of interest^69^. Case-control studies, although intuitively simple, can be challenging when selecting an appropriate control group, and a clear rationale for the selection of control groups is required^70,71^.
For prospective cohort studies, it is important to minimize loss to follow-up (a situation in which the researcher loses contact with the participant), resulting in missing data. When loss to follow-up of many individuals occurs, the internal validity of the study is reduced, as there may be systematic differences related to the disease or exposure to risk factors for those who drop out versus those who remain in the study. Strategies to minimize loss to follow-up, such as improving participant engagement, education, and establishing trust, are helpful. Other more practical things include having clear study protocols, convenient follow-up procedures (e.g., flexible visits or passive follow-up via hospital or medical records)^16,17^.
For AI algorithms, bias mitigation strategies may involve pre-processing data through sampling before a model is built, in-processing by implementing mathematical approaches to incentivize a model to learn balanced predictions, and post-processing^72^. Furthermore, as experts can be aware of biases specific to datasets, “keeping the human in the loop” can be another important strategy to mitigate bias^73^.
Bias in AI algorithms can also arise from lack of diversity of the people who label data or validate an AI algorithm. Bias may also be introduced during the implementation of AI algorithms, if the diversity of the population due to age, co-morbidities, disability or poverty has not been considered.
It is much better to avoid data issues right at the beginning of a data collection than to search for “statistical solutions” post-hoc. Thus, researchers should closely cooperate on the integrity of a study during its design stage. This encompasses an appropriate, structured setup of metadata before the start of the data collection to identify design flaws.
To avoid information bias in clinical research, it is necessary to utilize good data collection instruments, such as validated measures. It can be useful to pilot test the data collection instruments to ensure consistency and reliability.
If measurement error or misclassification is a concern, it may be useful to collect repeated measures of the exposures of interest or conduct a validation sub-study whereby an exposure measure used in the main study is supplemented by a more accurate or “gold standard” method that can be used to calibrate the main study exposure. This is common in prospective epidemiological studies where measurement error, misclassification, or within-person variation is a potential issue.
A more expensive but alternative approach is to collect objective measurements such as biomarkers or laboratory tests to supplement self-reported information.
Performing regular quality control checks is also useful in order to ensure the data accuracy and potentially identify any anomalies in the data. Clinical trials often perform event adjudication to ensure accuracy of endpoint diagnoses (e.g., stroke and ischemic heart disease).
How the analysis of a clinical research study is conducted is extremely important and should be done in a logical, rigorous, and transparent manner. Initial Data Analysis (IDA) can be considered as the first data analysis step in order to verify if the observed data correspond to expectations about these data^74^.
Key phases of IDA (1) metadata setup; (2) data cleaning; (3) data screening; (4) initial data reporting; (5) refining and updating the research analysis plan; and (6) documenting and reporting IDA. These phases represent essential activities for primary or secondary data analysis by all researchers and clinicians, involving, e.g., designed experiments, observational studies, patient registries, EHRs, biobanks, or biomedical databases^74^.
IDA is a crucial part of the research pipeline, and as such, it should be well documented to promote transparency, utility, and reproducibility. Therefore, keeping records of the changes made to the project data, programs and algorithms (including analysis scripts, libraries, and packages), as well as documentation (including plans and reports) are key factors to IDA good practice^1^.
The key to sensible and reliable statistical analysis is to ensure that the analytic methods are in close alignment with the key scientific research questions. The analyses should be accompanied by a clear explanation of the rationale for the choice of statistical method employed, which should enable the reader to assess whether the analytic technique is appropriately linked to the research questions on interest^75^.
It is always useful to use the principle of parsimony as a guide by starting with simple approaches and only adding complexity on an as-needed basis.
Sensitivity analyses to assess the robustness of the findings are important. Lash and colleagues have suggested performing quantitative bias analyses, which is an overarching term that proposes to use methods to quantify the direction, magnitude, and uncertainty associated with common biases (e.g., selection and information) that can occur in epidemiological studies^46^.
It is useful to remember that a well-designed and implemented study can often allow simple methods of analysis to produce clear and reliable results^76^.
It is always best for the analyst to make sure that their analyses are reproducible. Essentially, when one is given the same set of data and a complete description of the analysis steps, then the tables and figures should be reproducible, and similar statistical inferences should be possible^75^.
Training of new researchers and ongoing continuous professional development of existing researchers can help in ensuring that researchers conducting analyses are aware of the key issues and pitfalls to guard against in statistical methodology in order to improve the reliability and efficiency of clinical or scientific research^77^.
Many journals now increasingly require that study protocols have been registered before initiating or completing the study, in order to mitigate against selective reporting and to increase transparency, and there is some evidence that registration of trials is effective^78^.
In RCTs, this transparency is strengthened when the SAP is prospectively finalized before database lock/unblinding and clearly distinguishes primary confirmatory analyses from secondary/exploratory analyses, including any multiplicity adjustments^7^.
A variety of reporting guidelines have been developed for observational studies^79^, clinical trials^10^ and systematic reviews^80^ to improve the quality of reporting of research and these are endorsed by many of the main journals.
These reporting guidelines have been extended to deal with newly established approaches, i.e., Mendelian Randomisation studies^81^, multi-arm parallel trials^82^, systematic reviews of network meta-analyses^83^ and reporting of trials for interventions that involve AI models^42^.
Open Science has promoted the process of making content and claims as transparent and accessible as possible. Many published articles are not available to people without a personal or institutional subscription, so it means that those readers based in lower- or middle-income countries may have less opportunity to access the latest evidence^84^. This dissemination bias is not only applicable to reports, but it also applies to the data that is used to generate the reports.
The FAIR (Findability, Accessibility, Interoperability, and Reuse) principles provide guidance on the use of digital assets. The principles refer to three types of data (or any digital object), metadata (information about that digital object), and infrastructure. For instance, principle F4 defines that both metadata and data are registered or indexed in a searchable resource (the infrastructure component)^85^.
Beyond formal checklists, investigators often detect bias through simple ‘forensic’ checks that interrogate whether the data behave as expected under the intended target population and sampling mechanism. These include auditing where participants/cases come from (e.g., place of residence versus treatment location), checking whether patterns differ systematically by site, time, or subgroup, and repeating key analyses under plausible restrictions (e.g., limiting to local residents) to see whether conclusions are robust. Such checks are particularly important when using hospital-based or administrative datasets, where referral patterns and documentation practices can strongly shape observed outcomes^86^.
For RCT’s, checking whether an appropriate randomization technique (e.g., constrained randomization), to balance for key prognostic factors, was employed can also be an important gauge of potential bias^87^.
In an epidemiological study, it is useful to know whether there have been any strategies employed at the design stage, such as matching, to avoid bias for key confounding factors (such as age and sex).
Checking whether the instruments used to measure the key exposures and outcomes are appropriate and validated can also be valuable in detecting bias in epidemiological studies.
In big data studies, it is important to try to assess bias during data pre-processing, model development, and model validation stages. To reduce bias, people with diverse backgrounds should be included, as a necessary action in recognizing flaws in the design, functionality, or validation of the AI algorithms.
Code sharing can be used to identify any issues with the implementation of the methodology used to develop, train and validate the AI algorithms^88^.
It is useful to assess whether the randomization was performed correctly by examining the baseline table (Table 1), (not via p-values), to assess for potential imbalances between intervention and control groups^89^. In practice, Table 1 should be treated as a diagnostic instrument, focusing on whether any imbalances are clinically/prognostically meaningful, rather than on null-hypothesis testing of baseline differences (which is discouraged in CONSORT guidance). When an important prognostic imbalance is present, especially in smaller trials, investigators should consider pre-specified covariate adjustment (and/or stratified analyses) in the primary or sensitivity analysis plan to improve precision and to reduce the risk that chance imbalance drives interpretation^90^.
Non-blinded RCTs are known to overestimate the effect sizes, so comparing the findings to blinded studies could enable the detection of potential bias due to a lack of blinding^91^.
Early stopping of trials due to benefit has also been shown to lead to bias, and the play of chance can lead to an overestimation of treatment effect sizes^92^. As part of good clinical practice, trial sponsors are required to monitor the conduct of clinical trials. The aim of monitoring clinical trials is to ensure the patients’ well-being, compliance with the approved protocol and regulatory requirements, and data accuracy and completeness^93,94^.
A useful distinction between the three types of trial monitoring has been suggested by Baigent et al.^93^: oversight by trial committees, on-site monitoring and central statistical monitoring; they argue that the three types of monitoring are useful in their own right to guarantee that the trial has been well conducted, as well as the quality of the trial data and the validity of the trial results.
When interim analyses are planned, preserving trial integrity typically requires confidential handling of unblinded interim results. In many RCTs, interim efficacy/safety data are reviewed by an independent Data Monitoring Committee, with interim analyses prepared by an independent statistician who is operationally ‘firewalled’ from the trial’s day-to-day conduct and decision-making; only high-level recommendations (e.g., continue/stop/modify for safety) are communicated to the study team, rather than unblinded results themselves. This separation helps reduce operational bias (e.g., changes in enrollment, co-interventions, endpoint ascertainment, or selective analytic choices driven by knowledge of emerging results).
These RCT practices also translate to non-trial analytical epidemiology and AI/data science studies can adopt analogous ‘analysis firewalls’, such as pre-specifying analyses, restricting access to outcome labels during feature engineering, using locked holdout/test sets evaluated by an independent analyst/team, and maintaining a clear separation between model development and final performance assessment.
It is sensible to scrutinize the participant inclusion and exclusion criteria, which may give some clues on whether selection bias may be present.
A practical way to probe selection/referral bias in hospital-based outcome analyses is to audit patient origin and residence. For example, if a Chicago hospital system shows unusually high breast-cancer mortality among African American patients, one can examine whether decedents were predominantly local residents or whether a substantial fraction were referred from other regions (potentially presenting with more advanced disease after traveling for care). Stratifying outcomes by residence (local vs. referred), comparing stage/severity at presentation, and performing sensitivity analyses restricted to the intended catchment population can clarify whether an apparent geographic disparity reflects care quality versus case-mix and referral patterns^86^.
The methods to collect the data should be evaluated, and if the data are collected solely based on participant self-reports, that may give some indications of bias.
For big data analytic studies, bias mitigation would involve model developers being transparent about the selected training data and the various pre-processing techniques employed^5,35^.
To increase algorithmic fairness, protected attributes, such as gender or ethnicity, can be included during training in order to ensure that algorithmic predictions are statistically independent from these attributes^95^.
Non-standard analyses for a conventional study design or analyses that are inappropriate for answering the specific research question can be a good indication of bias in the analyses. This could occur if the underlying assumptions of the analyses are not met—e.g., if a cluster randomized trial is analyzed as if the observations are independent^96^.
A lack of intention-to-treat analyses for an RCT can also be a source of bias and an analyst may introduce bias if they perform an on-treatment analysis where patients are only analyzed if they received the treatment they were randomized to^97^. An intention-to-treat analysis preserves the randomization process and thus includes all patients and should be unbiased^97^.
Inappropriate subgroup analyses can also lead to misleading results, particularly if the analyses were not pre-specified and were data-driven, and tests for interaction are not used^98^.
Even in randomized trials, differential loss to follow-up or missing outcomes can bias effect estimates when missingness is related to prognosis, treatment, or outcomes (Table 2). Plan and report the extent/patterns of missingness, retain intention-to-treat as primary, and pre-specify missing-data handling and sensitivity analyses.
In clinical research, missing data can be an important source of bias. In many settings, complete case analysis or simple mean imputation can lead to biased estimates of statistics (e.g., of regression coefficients) and/or confidence intervals that are artificially narrow^99^. More sophisticated approaches, such as Multiple Imputation, to deal with missing data in the analyses, should be utilized^100^.
The emergence of new technologies and standards has increased the ability to detect some errors in statistical analyses. For example, the R package “statcheck” can automatically test whether there are inconsistencies between p values and test statistics^101^.
Other software is available to detect an error of granularity, called granularity-related inconsistency of means (GRIM). This technique evaluates whether the reported means of integer data, such as Likert-type scales, are consistent with the given sample size and number of items^102^.
In the case of administrative data or electronic health records, the use of machine learning or AI algorithms needs to be explainable. Explainable AI models aim to provide clarity on the feature selection as well as making the strengths and weaknesses of a decision-making process more transparent^103^. Explainability of AI models is an imprecise and controversial science. For example, it is not necessary to understand the complex mechanisms of action of a drug to use it according to its label. There is a need for the efficacy of AI models should be “FDA labeled” with precise descriptions of the subject population and intended clinical scenarios for use^5^.
There are several critical appraisal tools to assess the potential for risk of bias for RCT’s^104^ and non-randomized studies^105^. One practical way to detect bias is to ascertain the protocol of a clinical trial or trial registry (via databases such as Clinicaltrials.gov or the WHO clinical trials database) and compare the intended outcomes of interest to the analyzed outcomes published in the final paper^106^.
The reporting of the sampling strategy can enable the detection of the potential for selection bias, a non-representative sample, or a lack of generalizability of the findings.
The reporting of dropout (“lost to follow-up” or “withdrawal from the study”) rates and reasons for attrition can also give some indication whether the results may be affected by bias.
It may also be useful to consider the funding source and potential conflicts of interest both financial and non-financial that could be potential indicators of bias in the reporting of the findings.
For big data analytics, the advent of pre-print servers and open science has made it now possible for researchers to receive rapid feedback on their work from a diverse community instead of waiting several months or more for peer reviewer comments on their work. This may include potential errors being detected earlier in analyses, code or interpretation^84^.
Dissemination bias can be detected if there is a preponderance of positive findings in the literature and it is known that statistically significant results are more likely to be published in a timely manner than studies that have more modest findings^107^.
A recent report offered a two-step process for the identification of “spin bias”, consisting of tracking data and findings and recording of data discrepancies by describing how the “spin bias” was produced in the text^108^.
Check carefully, (1) how results are presented (e.g., non-significant results presented as causal or even as significant in the Discussion section), and (2) for spin bias whereby specific findings are omitted from the Discussion section of a paper^108^.
Scrutinize the logic; poor logic can distort findings, resulting in conclusions that do not follow from the data, analysis, or fundamental premises, such as assuming causality from observational evidence^109^.
Selection bias is best prevented, but when it is suspected or unavoidable, its impact should be made explicit and stress-tested. Practically, mitigation (1) Define the target population and the selection mechanism (the sampling frame, inclusion/exclusion, and any gatekeeping steps) and document deviations; (2) Measure/monitor representativeness by comparing included vs. eligible/non-included individuals on available baseline characteristics (and, where feasible, record minimal data on non-participants and reasons for non-participation/attrition); (3) Adjust using methods appropriate to the mechanism (e.g., inverse-probability weighting for differential selection/attrition when assumptions are plausible, or transportability/generalizability approaches when moving from a trial sample to a target population); and (4) Quantify residual uncertainty using sensitivity analyses (e.g., varying assumptions about selection/attrition, quantitative bias analysis, and transparent reporting of how conclusions change under plausible scenarios).
To mitigate surrogacy-related bias, investigators should (1) prioritize clinically meaningful endpoints when feasible; (2) when a surrogate is necessary, explicitly describe its rationale and limitations, and, where possible, link it to patient-important outcomes using prior evidence or validation studies; and (3) in EHR/AI settings, validate outcome definitions against adjudicated ‘gold-standard’ subsets and report sensitivity analyses across plausible outcome definitions.
Information bias can occur during the conduct of a clinical research study and can usually manifest itself in the form of misclassification or measurement error. Measurement error is one of the key challenges in making valid inferences in clinical research and these types of error can arise due to inaccurate or imprecise measurement instruments, single measurements of non-stationary (time-varying) longitudinal processes, or non-compliance with measurement protocols.
With the increased utilization of EHR or routine care data (or ‘big data’), measurement error is likely to become increasingly relevant in this field^110^. The potential impact of random measurement (non-differential) error may be “corrected for” by collecting replicate measurements on the exposure and relevant confounders to assess their reliability. Random within-person variability and other imprecision in exposure measurement can also attenuate exposure–outcome associations (often termed regression dilution bias)^111,112^.
If systematic error (or differential measurement error) in the exposure is suspected then a validation study should be performed whereby the crude exposure and a “gold-standard” measure of the exposure are measured in a subset of participants in order to assess the accuracy of the exposure measure^113^. Such studies can be very helpful to understand the nature of the measurement error when the exposure of interest is difficult to measure directly such as diet^114^ or physical activity^115^.
Differential weighting can be used when different parts of the population are sampled with unequal selection probabilities. Inverse probability weighting is a technique that can be utilized in different scenarios (e.g., randomized trials with crossing over from one arm to the other)^116^.
While baseline confounding is minimized by randomization, analyses that condition on post-randomization behavior (e.g., as-treated/on-treatment/per-protocol without appropriate methods) can reintroduce bias because adherence is often related to prognosis. Intention-to-treat preserves randomization and should remain the primary analysis.
Differential drop-out or missing outcome data can act like selection mechanisms that distort the randomized comparison (Table 2). This is especially problematic if missingness differs by arm or is related to outcomes; it should be monitored and handled with pre-specified approaches and sensitivity analyses.
In non-randomized studies comparing different treatments, the choice of treatment is likely to be influenced by predictors of outcome known as confounding by indication. Inverse probability of treatment weighting can be used where the weights are based on each individual’s probability of receiving a specific treatment given the confounders, which is known as the propensity score (PS). The weights are 1/PS for the treated participants and 1/(1 − PS) for the untreated participants^117^.
Selection bias due to missing data can be “corrected for” in epidemiological studies using inverse probability weighting, but the appropriateness depends on the missing data mechanism^116^.
When reverse causality bias is suspected, the common ways to deal with this are to exclude participants with a prior history of disease that may have changed their behavior (sick quitters) and those who may have had events early during the follow-up period, as they may have had incipient disease that has gone undetected.
Alternative strategies, such as Mendelian Randomization, may be employed as this technique is more robust to reverse causality^118^.
Depending on the relationship of the confounder with the exposure and the outcome, as well as the type and magnitude of the measurement error on the exposure and/or the confounder, the exposure-outcome relation may be attenuated, exaggerated or remain unaffected due to the measurement error.
There is a range of analytical approaches available that are now routinely available in standard software, such as regression calibration^119^, simulation extrapolation^120^ to address random measurement error in continuous covariates for epidemiological studies^121^.
More sophisticated approaches are also available to correct for measurement error for difficult-to-measure exposures with complex error structures, such as nutritional exposures^122^, or physical activity and misclassification of categorical covariates^123^.
Researchers, clinicians, and data scientists need to be aware that the presence of random measurement error in the exposure or the confounders does not automatically result in an attenuation of the exposure-outcome relationship. In fact, it can be difficult to anticipate the direction of effect of random measurement error on the exposure-outcome relationship when there is at least one confounder measured with error^124^.
To mitigate potential sources of missingness or selection bias, one should collect baseline characteristics and outcome data on study non-participants who are part of the target population.
One should highlight the presence of groups who are at risk of disparate health outcomes caused by structural or societal factors in this dataset, with consideration of both risk factors that are universal and those that are specific to the site of data collection.
Documentation should also describe (1) any known missing groups within the dataset; (2) the proportion, nature, and any reason(s) for their missingness (if known), particularly if there are systematic differences across groups within the dataset; and (3) if and how missing data have been handled^125^.
Perform sensitivity analyses to assess how robust the findings of a clinical research study are to potential biases. Residual confounding is an ever-present issue in observational research.
The concept of the E-value has been introduced and defined as the minimum strength of association, on the risk ratio scale, that an unmeasured confounder would need to have with both the exposure and the outcome to fully explain away a specific exposure-outcome association, conditional on the measured covariates^126^.
A large E-value implies that considerable unmeasured confounding would be needed to explain away a particular effect estimate. A small E-value implies little unmeasured confounding would be needed to explain away an effect estimate^126^.
Another approach is triangulation, this can be used to increase the credibility and validity of research findings. In clinical research studies, this involves integrating results from several different approaches, where each has different underlying assumptions and different potential biases that are orthogonal to each other. In clinical research, if the different approaches are in broad agreement, then this strengthens the confidence in the findings^127^.
In the AI field with applications to healthcare, mathematical de-biasing approaches such as adversarial de-biasing or oversampling have become more prevalent^23^. These approaches allow the model to incorporate underrepresented groups (such as ethnic minorities) more efficiently in the analyses.
A recent report using an adversarial training framework found that adversarial training improved outcome fairness whilst still achieving good performance on established metrics^128^.
Reporting bias can occur in the form of selective reporting or the decision to report depends on the direction or magnitude of the findings. This can be due to several causes, but it can also be due to the motivation of both researchers and journal editors^129^.
It is always useful to put the results of an individual study in context of other previous published work but performing a meta-analysis and graphical displays such as funnel plots can be used to assess whether publication bias may be present^130^.
There are also more sophisticated approaches to model or adjust for the presence of publication bias^131^. These methods aim to give an indication of how many negative unpublished studies are required to modify or “correct” the inferences made from the present study^105^.
Errors in reporting identified by the authors or the readers of a journal, can be corrected by the authors.
The publication of errata that identify errors that may change the inferences or interpretation of a report are relatively frequent^132^.
Although errors can still occur despite meticulous reviewing, greater mitigating strategies are required to reduce the number of errors, and the potential for perpetuation of inaccurate information that may lead to adverse consequences^132^.
Pre-prints can allow the manuscript to be disseminated to a wider audience and feedback and errors can be spotted earlier.
In this primer, we have outlined the three primary categories of bias in clinical research and illustrated how they can manifest throughout a study’s lifecycle (Fig. 2). While we have not provided an exhaustive catalog, we have focused on the most common forms a researcher, clinician, or data scientist is likely to encounter, with examples that bridge traditional epidemiology and modern data science. We refer the reader to an excellent glossary of the different types of bias that can occur under the main headings of selection, information, and analytical biases by Delgado^13^.Fig. 2From trials to bias across modern clinical research.How it enters, spreads, and can be countered across research domains, lifecycle stages, and the wider research ecosystem.
As clinical research becomes increasingly reliant on large-scale data and interdisciplinary collaboration, the need for a shared understanding of bias has never been more critical. Flawed data and biased analyses not only threaten the validity of research findings but can also perpetuate and even amplify health disparities when encoded into clinical algorithms.
Ultimately, the responsibility to produce robust and reliable evidence is a collective one. Investigators should proactively design studies to avoid bias, apply appropriate analytical methods to adjust for it when unavoidable, and transparently discuss the effects of any residual bias on their results (Fig. 2). By fostering a common language and a vigilant approach, collaborative research teams can improve the quality of evidence and ensure that scientific advancements benefit all patients equitably.