Authors: Ko Konno, James Gibbons, Ruth Lewis, Andrew S Pullin
Categories: Methodology, Critical appraisal, Risk-of-bias assessment, Threats to internal validity, Validity of causal inferences
Source: Environmental Evidence
Authors: Ko Konno, James Gibbons, Ruth Lewis, Andrew S Pullin
To inform environmental policy and practice, researchers estimate effects of interventions/exposures by conducting primary research (e.g., impact evaluations) or secondary research (e.g., evidence reviews). If these estimates are derived from poorly conducted/reported research, then they could misinform policy and practice by providing biased estimates. Many types of bias have been described, especially in health and medical sciences. We aimed to map all types of bias from the literature that are relevant to estimating causal effects in the environmental sector. All the types of bias were initially identified by using the Catalogue of Bias (catalogofbias.org) and reviewing key publications (n = 11) that previously collated and described biases. We identified 121 (out of 206) types of bias that were relevant to estimating causal effects in the environmental sector. We provide a general interpretation of every relevant type of bias covered by seven risk-of-bias domains for primary risk of confounding biases; risk of post-intervention/exposure selection biases; risk of misclassified/mismeasured comparison biases; risk of performance biases; risk of detection biases; risk of outcome reporting biases; risk of outcome assessment biases, and four domains for secondary risk of searching biases; risk of screening biases; risk of study appraisal and data coding/extraction biases; risk of data synthesis biases. Our collation should help scientists and decision makers in the environmental sector be better aware of the nature of bias in estimation of causal effects. Future research is needed to formalise the definitions of the collated types of bias such as through decomposition using mathematical formulae.
The online version contains supplementary material available at 10.1186/s13750-024-00324-7.
Estimation of average causal effects of interventions or exposures can provide important evidence to inform environmental policy and practice. There is a risk that if these estimates are derived from poorly conducted (or selective reporting or selectively reported [1–4]) primary (e.g., impact evaluation studies) or secondary research (e.g., evidence reviews), then they will potentially misinform policy and practice. Biased estimates of average causal effects (inaccuracy) are more problematic than imprecise estimates of average causal effects due to random error which may be measured by standard errors or confidence intervals [5] because a bias is a systematic error or ‘deviation from the truth’ [6, 7], and can be a direct threat to the validity of a causal inference [8]. Biases can be quantified only when the true average causal effects are known, or when hypothetical true (or expected) average effects are set as reference values (i.e., B = m − µ, where m is an estimated average causal effect and µ is the true causal effect [7, 9, 10]; it is also defined as \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ E\left( {\widehat{\theta }} \right) - \theta
Many types of bias (specific phenomena resulting in systematic errors described in the literature and on the Internet) have been described, most notably in medical and health sciences where methods for assessing the risk of bias originated (e.g., Catalogue of Bias describes 62 types of bias as of 29 August catalogofbias.org). For example, observer bias is described as ‘*the process of observing and recording information which includes systematic discrepancies from the truth*’ [14], and detection bias is described as ‘*systematic differences between groups in how outcomes are determined*’ [15]. Types of bias, which are considered by individuals or groups as distinctive phenomena that can give rise to biased estimates, are generally described freely, but some can also be defined formally using mathematical languages [e.g., confounding bias and collider bias are often defined using directed acyclic graphs (DAGs)] [16, 17]. In the health sector, there are dedicated tools for assessing the risk of bias or internal validity [18, 19]. For example, Risk of Bias version 2 (RoB 2) [20] and Risk Of Bias In Non-randomised Studies—of Interventions (ROBINS-I) [21] are widely recommended domain-based tools for assessing the risk of bias in primary research (the former is for randomised controlled trials and the latter is for non-randomised intervention studies) when conducting systematic reviews in the health sector [18]. There is another widely recommended domain-based tool called Risk Of Bias In Systematic reviews (ROBIS) which is for specifically assessing the risk of bias arising from the conduct of secondary research in the health sector [18, 22]. These risk-of-bias assessment tools guide users to focus on a few ‘domains’ of bias (broad categories of bias specifically designed for assessing the risk of bias in particular situations; for example, when you plan to synthesise effect sizes from multiple studies, you would need to assess the risk of bias in primary research assessing effects by developing a comprehensive set of domains to address all potential sources of bias occurring in primary research in your subject area) [13]. Risk-of-bias assessors can make overall judgements about the extent of risk of bias for the specified domains (e.g., low risk of bias, some concerns, or high risk of bias [20]). The concept of domains of bias was initially developed in the health sector to assess the risk of bias for groups of types of bias that can occur at specific stages of primary or secondary research [20, 21]. For example, in the ROBINS-I tool, detection bias, recall bias, information bias, misclassification bias, observer bias, and measurement bias are all included in a domain called ‘risk of bias in measurement of outcome’ [21]. Following the developments of the risk-of-bias assessment tools in the health sector, the Collaboration for Environmental Evidence (hereafter ‘CEE’: environmentalevidence.org) is adopting such a domain-based approach and has been developing a domain-based critical appraisal tool called ‘CEE Critical Appraisal Tool’ since 2020. This tool is for assessing the risk of bias in primary research and is available for use in evidence reviews in the environmental sector (environmentalevidence.org/cee-critical-appraisal-tool). Employing domains of bias is now strongly encouraged when conducting evidence reviews in medical and health sciences [20, 21, 23, 24] as well as in the environmental sector [13] because assessing every individual type of bias is laborious and there are no universal definitions of all existing types of bias in estimation of causal effects in science. A recent review on critical appraisal in ecology highlighted this important development and the urgent need to evaluate the validity and reliability of the CEE Critical Appraisal Tool because, remarkably, only 4% of reviews labelled as ‘systematic reviews’ in ecology actually conducted critical appraisal [25]. This finding suggests insufficient education about biases in environmental science courses and a lack of understanding of how to avoid providing potentially biased estimates [26]. We aim to raise awareness about biases in the environmental sector by mapping types of bias that may impact the reliability of the results of environmental research evaluating the effects of interventions or exposures and categorising them into the developed domains of bias. The objectives of our research (1) to identify potential types of bias, that have been previously collated, in estimation of causal effects in either primary or secondary research; (2) to evaluate the relevance of the identified biases to the environmental sector, which incorporates not only non-human biotic (e.g., plants, animals) and abiotic subjects (e.g., water, atmosphere) but also human subjects; and (3) to utilise the resultant list of biases to evaluate the coverage of the domains of bias for a previously developed assessment tool to use in environmental research, namely the CEE Critical Appraisal Tool. To avoid confusion, when we refer to domains of bias, we use plural rather than singular (e.g., confounding biases) as noted below. Also, note that we do not aim to evaluate relatedness among biases (e.g., we do not discuss whether selection bias is related to attrition bias [27]). Rather we aim to provide an overview of the types of bias from the view of managing the risk of bias by employing the developed domains. ## Methods ### Listing the types of bias previously collated and described elsewhere In order to develop a comprehensive list of potential biases we used the Catalogue of Bias (catalogofbias.org), and reviewed 11 key publications in the Sackett [28], Bayliss and Beyer [29], Clarke et al. [30], Smith and Noble [31], Thakur et al. [32], Paradis [33], Warden [34], Delegado-Rodriguez and Llorca [35], Hartman et al. [36], Marchevsky [37], and Pannucci and Wilkins [38]. The Catalogue of Bias is an ongoing collaborative project aiming to map all the biases that affect health and medical evidence. The Catalogue of Bias platform was developed by the Centre for Evidence-Based Medicine (CEBM) at Oxford University through reviewing the literature and regular meetings. We considered the Catalogue of Bias to provide a robust and comprehensive source. To our knowledge, there was no equivalent database of bias in environmental sciences, so the Catalogue of Bias was the only choice, apart from the 11 key publications. The additional 11 key publications used as information sources were selected because they were known to the authors as seminal papers that collated and described biases. We listed biases that were described in these sources in a Microsoft Excel spreadsheet (Redmond, WA, USA). The intention here was to identify the breadth of biases that have been recognised or described and not to identify all studies describing each bias. We therefore did not conduct a systematic literature search. The broad scope of our review and the time and resources available also precluded this. ### Evaluating the relevance of the biases to estimation of effects in the environmental sector Since the vast majority of the types of bias were described in the medical and health research contexts, we evaluated them for relevance to estimation of effects of interventions or exposures in the environmental sector. Relevance was evaluated by a single reviewer (KK) using the algorithm provided in Fig. 1.Fig. 1The algorithm used for evaluating the relevance of biases described in the sources We employed a broad and all-inclusive scope for evaluating relevance to the environmental sector. The following a priori list of relevant topics was developed and used for assessing agriculture, aquaculture, biodiversity conservation, climate change, ecology and evolution, ecosystem services, environment and human wellbeing, environmental education, environmental legislation, fisheries, food security, forestry, invasive species, natural resource management, park and protected area management, pollution, soil management, sustainable energy use, waste management, wastewater management, water security, as well as any research relevant to environmental sustainability. The broad scope is consistent with that of CEE. The scope inherently included not only non-human biotic and abiotic subjects but also human subjects, however, human health outcomes were out of scope. The algorithm was applied by one person (KK) and the resulting decisions were checked by a second person (AP), and any disagreements or uncertainties were discussed among all authors. We highlight that we did not carry out any consistency checking, as we did not aim to aggregate effects, and did not aim to provide databases for aggregating effects. Hence, Cohen’s kappa statistic [39, 40] was not calculated. ### Evaluating the coverage of the developed domains of bias Suzuki et al. [8] provided an organisational schema for systematic error in causal inference in primary research. In primary research, systematic error can be divided into structural error and analytic error. Structural error can be further divided into bias relating to measurement of intervention, exposure or outcome, and bias relating to exchangeability. Exchangeability refers to independence between the counterfactual outcome \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ {Y}^{a} $$\end{document}Ya (the outcome which a subject or area would have experienced if it had received a specified intervention or exposure value [41]) and the observed intervention or exposure \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ A $$\end{document}A, and it is denoted as *Y*^*a*^ ⫫ A, for all treatment values *a* (all values of the treatment variable that may differ among individual subjects or areas) [42]. Alternatively, exchangeability refers to an assumption that the outcome in the comparator group would mirror the outcome in the intervention or exposure group if the subjects or areas in the comparator group had been subjected to the same intervention or exposure as those in the intervention or exposure group [17]. Analytic error is a systematic error that cannot be explained by structural error, and it includes bias relating to reporting [20, 21] and statistical methods [23, 24, 43]. Figure 2 illustrates our assumptions about the organisation of bias occurring in primary and secondary research from the broadest term (systematic error) to the developed domains of bias.Fig. 2Organisation of bias from the broadest term (systematic error) to the developed domains of bias Frampton et al. [13] provided the principles of risk-of-bias assessment in the environmental sector and recommended seven domains (or classes) to be assessed when conducting evidence reviews assessing the effect of an intervention or exposure. These domains for primary research were consistent with the CEE Critical Appraisal Tool version 0.3 (see environmentalevidence.org/cee-critical-appraisal-tool for more details) as including the Confounding Risk of biases due to an uncontrolled (or inappropriately controlled) variable (confounder) that influences both the intervention/exposure and the outcomePost-intervention/exposure selection Risk of biases arising from systematic differences in the selection of subjects or areas into the study or analysis after the intervention or exposureMisclassified/mismeasured comparison Risk of biases arising from misclassification or mismeasurement of the intervention, exposure and/or comparator (applicable to observational studies only)Performance Risk of biases due to altered treatment procedure of interest (applicable to experimental studies only; treatment procedures are series of actions for applying the experimental (intervention or exposure) treatment [44]; note that an experimental treatment can be either an intervention or an exposure, or both, depending on the outcome measure. For example, when a fertilizer is applied to a crop, this is an intervention if the outcome measure is crop yield, but an exposure if the outcome measure is change in soil microfauna [45])Detection Risk of biases arising from systematic differences in measurement of outcomesOutcome reporting Risk of biases in reporting of study findingsOutcome assessment Risk of biases due to error in applied statistical methods Regarding the domains for secondary research, we used the CEE Synthesis Appraisal Tool (CEESAT) version 2 [46, 47]. CEESAT is an eight-criteria checklist consisting of sixteen questions for assessing environmental evidence reviews in terms of risk of bias, repeatability, and transparency, and hence not all the criteria are relevant to risk of bias (e.g., review question setting and provision of limitations). Using only the relevant criteria resulted in four domains of bias which were the same as the ROBIS tool [22], and thus consistent with a widely recommended tool in medical and health sciences [18]:Searching Risk of biases in searches of relevant recordsScreening Risk of biases arising from screening of potentially relevant recordsStudy appraisal and data coding/extraction Risk of biases due to the lack of or inappropriate conduct of study appraisal and data coding/extractionData synthesis Risk of biases due to employing inappropriate synthesis methods Each type of bias was checked by a single reviewer (KK) for relevancy against each domain. Each bias-and-domain combination was data-coded in the database (spreadsheet), by the single reviewer using subjective judgements (‘Yes’ = relevant/‘No’ = not relevant). All decisions were double-checked by the other authors. The evaluator was allowed to take special notes when some types of bias were beyond the provided domains but there were no such cases. We also data-coded relevant levels of research (primary and/or secondary research) for all included types of bias. ## Results ### Types of bias relevant to the environmental sector Figure 3 summarises the results of the bias selection process and provides the numbers of biases identified in each of the sources searched. Some biases were described by multiple sources while some biases were described by one source only. There were 206 types of bias after merging 53 bias descriptors that were specifically described as synonyms; we did not merge bias descriptors unless they were specifically referred to as synonyms to make sure none are missed from the list. Note that we did not reproduce (verbatim) the descriptions of the biases provided in the sources mainly due to copyright restrictions.Fig. 3Identification of biases previously collated and described elsewhere After the evaluation of relevance, we identified 121 types of bias (accompanied by 24 synonyms specified in the sources) that were relevant to estimation of effects in the environmental sector. Of the 85 excluded types of biases, 53 (62%) were excluded due to a lack of relevance to estimation of effects. Thirty-one (37%) were excluded due to a lack of relevance to the environmental sector; the most common reasons for exclusions were specifically referring to human health outcomes, diagnosis, and healthcare and medical interventions. One (1%; culture bias described by Warden 2015 [34]) was excluded because its description was insufficient for judging its relevance. All the excluded types of bias can be found in Additional file 1. The general interpretations and relevant domains of the 121 types of bias relevant to the environmental sector are provided in Table 1. Sixty-eight types of bias (56%) were relevant to primary research only and 18 types of bias (15%) were relevant to secondary research only. Thirty-five types of bias (29%) were relevant to both the primary and secondary research. We provide the database of these biases in Additional file 2 so that readers can use the filter function of the spreadsheet to find types of bias relevant to specific domains. To increase its convenience, we also created a sheet to list types of bias relevant to each domain.Table 1Types of bias relevant to the environmental sector in the alphabetical orderTypes of biasGeneral interpretationsBias occurring in*Bias relating to**Relevant domains of bias***Source****1All’s well literature biasOccurs when review papers deliberately omit or inadequately address studies that present contrasting results compared to studies yielding the results desired by certain individuals or groups of individualsSecondary researchScreeningData synthesisScreening biasesData synthesis biases[28, 34, 36, 48]2Allocation bias (Allocation of intervention bias)Occurs when non-random assignment of participants to study groups (i.e., allocation) creates an imbalance between the groups (i.e., exchangeability does not hold)Primary researchExchangeabilityConfounding biases[30, 35, 37, 49]Allocation of intervention biasSee 2 ‘Allocation bias’3Allocation sequence biasOccurs when the allocation sequence is not concealed to researchers and the selection creates an imbalance between the study groups (i.e., exchangeability does not hold)Primary researchExchangeabilityConfounding biases[34]4Analysis biasOccurs when the knowledge of the results of a study causes inappropriate analyses and selective reportingPrimary researchSecondary researchReportingStatistical methodsData synthesisOutcome reporting biasesOutcome assessment biasesData synthesis biases[30–32]5Ascertainment biasOccurs when the ascertainment of subjects or areas included in a study or the ascertainment of measurements or collected data for analysis differs between the groups after the intervention/exposurePrimary researchExchangeabilityMeasurementPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[33, 35, 50]6Attention bias (Hawthorne effect)Occurs when human or animal subjects change their behaviour due to the awareness of being observed by researchersPrimary researchExchangeabilityMeasurementConfounding biasesPerformance biases[28, 35, 37, 51]7Attrition biasOccurs when missing data between groups are unequal, and exchangeability does not hold (cannot be assumed)Primary researchExchangeabilityPost-intervention/exposure selection biases[34, 52]8Authorization biasOccurs when there is inability to obtain authorization for use of certain data or documents, and that causes deviation from a true effectPrimary researchSecondary researchExchangeabilitySearchingConfounding biasesPost-intervention/exposure selection biasesSearching biases[36]9Availability biasOccurs when the information that is readily available differs systematically from the information obtained by aggregating all pertinent data to answer the review questionSecondary researchSearchingSearching biases[29, 53]10Bigwig biasOccurs when information in documents written by high profile authors differs systematically from the wider evidence base relevant to the review question and is more likely to be cited as evidence in a reviewSecondary researchSearchingScreeningData synthesisSearching biasesScreening biasesData synthesis biases[29]11Bogus control biasOccurs when subjects or areas initially assigned to the intervention/exposure group are taken out from that group or moved to the control (no intervention/exposure) group and the assumption of exchangeability no longer holdsPrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[28, 34, 36]12Case definition biasOccurs when the intervention/exposure is vaguely described or lacks relevant information, and the vagueness or the lack of information leads to mismeasurement of the intervention/exposure by researchers in an observational studyPrimary researchMeasurementMisclassified/mismeasured comparison biases[34]13Channelling biasOccurs when researchers assign subjects or areas, that are more likely to respond to the intervention/exposure, to the intervention/exposure group, creating an imbalance between the groups (i.e., exchangeability does not hold)Primary researchExchangeabilityConfounding biases[33, 34, 38]14Chronological biasOccurs when the timing of allocation of subjects or areas to a group (or groups) is different within or between groups and the time lag influences both the intervention/exposure and the outcome, or the subjects or areas are subject to a deviated intervention/exposure due to the time lagPrimary researchExchangeabilityMeasurementConfounding biasesPerformance biases[33, 38, 54]15Citation biasOccurs when studies with statistically significant results are more likely to be cited as pieces of evidence in a reviewSecondary researchSearchingScreeningData synthesisSearching biasesScreening biasesData synthesis biases[29, 33, 35, 38]Classification biasSee 52 ‘Information bias’16Cognitive dissonance biasOccurs when an investigator or a reviewer’s claim about an effect contradicts their obtained evidence due to the investigator or reviewer’s cognitive biasPrimary researchSecondary researchStatistical methodsData synthesisOutcome assessment biasesData synthesis biases[28, 34]17Collider biasOccurs when conditioning on a variable (i.e., filtering by a certain value of that variable) that is a common effect of the exposure and the outcomePrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[55]18Compliance biasOccurs when subjects compliant with the intervention differ from those not compliant in some characteristic and the characteristic is also associated with the outcome, or when non-compliance causes deviated initiation, implementation or discontinuation of the intervention between groups and the deviation is not taken into accountPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesPerformance biases[32, 34–36, 56]19Confidentiality biasOccurs when unrestricted research data, that systematically differs from restricted research data, are exclusively used due to commercial or privacy considerationsPrimary researchSecondary researchExchangeabilitySearchingConfounding biasesSearching biases[29]20Confirmation biasOccurs when the information to confirm ideas, beliefs or hypotheses is systematically different from the truthPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[34, 57]21Conflict of interestOccurs when a conflict of interest causes a systematic difference in estimation of effectPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[33]22Confounding biasOccurs when an uncontrolled (or inappropriately controlled) variable (confounder) that influences both the intervention/exposure and the outcome existsPrimary researchExchangeabilityConfounding biases[31, 34, 37, 38, 58]23Confounding bias by groupOccurs when there is a variable describing communities (groups) and that influences both the intervention or exposure and the outcomePrimary researchExchangeabilityConfounding biases[35]24Confounding bias by indicationOccurs when there exists an indication for the intervention/exposure (a preceding phenomenon that affects the intervention/exposure) or a variable influenced by this indication that affects both the intervention/exposure and the eventual outcomePrimary researchExchangeabilityConfounding biases[35, 59]25Contamination biasOccurs when subjects or areas in the comparator group (no or alternative intervention/exposure) accidentally receive the intervention/exposure (and vice versa), and this contamination leads to the violation of the assumption of exchangeabilityPrimary researchExchangeabilityConfounding biases[28, 34–36]26Correlation biasOccurs when equating non-causal association with causal associationPrimary researchSecondary researchExchangeabilityStudy appraisal and data coding/extractionData synthesisConfounding biasesStudy appraisal and data coding/extraction biasesData synthesis biases[28, 32, 34, 36, 37]Cost biasSee 100 ‘Resource bias’27Co-treatment biasOccurs when inappropriate or unplanned co-treatment(s) is/are used in an experiment that results in confounding or altered treatment(s)Primary researchExchangeabilityMeasurementConfounding biasesPerformance biases[30, 34]28Data capture bias (Data capture error)Occurs when data capture (data collection and recording) errors favour or disfavour the outcome for the intervention/exposure group or the comparator groupPrimary researchSecondary researchExchangeabilityMeasurementStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesStudy appraisal and data coding/extraction biases[34]Data capture errorSee 28 ‘Data capture bias’29Data collection biasOccurs when data collectors' beliefs influence the way relevant data are collected and that results in inaccurate estimation of an effectPrimary researchSecondary researchExchangeabilityMeasurementStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesStudy appraisal and data coding/extraction biases[31]Data completeness biasSee 68 ‘Missing data bias’Data entering biasSee 30 ‘Data entry bias’30Data entry bias (Data entering bias)Occurs when omitting data entry, inaccurately entering data, or inaccurately converting hand-written data into digital formPrimary researchSecondary researchExchangeabilityMeasurementStatistical methodsStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome assessment biasesStudy appraisal and data coding/extraction biases[34, 37]31Data merging biasOccurs when multiple data sets are merged incorrectly and the merged data produce an intentionally or unintentionally inaccurate estimate of an effectPrimary researchSecondary researchReportingStatistical methodsStudy appraisal and data coding/extractionOutcome reporting biasesOutcome assessment biasesStudy appraisal and data coding/extraction biases[34]32Data-dredging bias (‘Looking for the pony’)Occurs when the results of inappropriate statistical analyses or synthesis methods, that are not pre-specified, are presented, or statistical or synthesis methods are used inappropriately to support claimsPrimary researchSecondary researchReportingStatistical methodsData synthesisOutcome reporting biasesOutcome assessment biasesData synthesis biases[28, 32, 34, 36, 60]33Definition biasOccurs when definitions of important words used in a study or a review are not accurate or precise so that misinterpretation affects estimation of an effectPrimary researchSecondary researchExchangeabilityMeasurementSearchingScreeningStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biases[32]34Design biasOccurs when inferior experiments or treatments, reflecting a desire for certain results, are employedPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesPerformance biasesDetection biases[30, 31, 33, 34]35Detection biasOccurs when there are systematic differences in measurement of outcomes between groupsPrimary researchMeasurementDetection biases[15, 33–36]36Differential maturing biasOccurs when there are uneven secular trends (long-term and sustained changes or patterns) between the groups due to differential maturing, and that creates an imbalance between the groups before the intervention or exposure (time-varying confounding)Primary researchExchangeabilityConfounding biases[35]37Dissemination biasOccurs when retrieval of information does not result in a representative sample of studies due to the way study results are disseminated, or when study results affect reporting or disseminationPrimary researchSecondary researchReportingSearchingOutcome reporting biasesSearching biases[35]38Distribution assumption biasOccurs when a data analyst conducts a statistical test under a false assumption of the data distribution or the approach is not justifiable so that the inference is invalidPrimary researchStatistical methodsOutcome assessment biases[34]Dropout biasSee 120 ‘Withdrawal bias’39Duplication bias (Multiple publication bias)Occurs when duplicated research findings are included and synthesised in a reviewSecondary researchScreeningScreening biases[29, 37]40End-digit preference biasOccurs when there are systematic differences in quantitative measurements due to a preference for a certain ending digit (e.g., 0 by rounding up/down rather than retaining the original digits), typically ending with an unusual frequencyPrimary researchSecondary researchStatistical methodsStudy appraisal and data coding/extractionOutcome assessment biasesStudy appraisal and data coding/extraction biases[28, 32, 34, 36]41Estimator biasOccurs when an inappropriate estimator (a statistic used to estimate the parameter) is used for estimating the true causal effect (parameter)Primary researchSecondary researchStatistical methodsData synthesisOutcome assessment biasesData synthesis biases[34]42Exclusion biasOccurs when the criteria for inclusion of subjects or areas into a study or analysis are applied differently to the intervention/exposure group and the comparator groupPrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[34, 35, 37]43Expectation biasOccurs when observers err in measurement, recording, or data collection due to prior expectationsPrimary researchSecondary researchExchangeabilityMeasurementSearchingScreeningStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biases[28, 34, 36, 37]44Exposure suspicion biasOccurs when a knowledge or awareness of the subject's or area's exposure status influences selection, measurement, reporting, or analysisPrimary researchExchangeabilityMeasurementReportingStatistical methodsConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biases[28, 35, 36]45Familiarity biasOccurs when reviewers limit searching to documents relevant to their own disciplineSecondary researchSearchingSearching biases[29]46Foreign language exclusion biasOccurs when language is restricted in a review and studies with certain direction of results or certain topics are more or less likely to be disseminated in the eligible language(s)Secondary researchSearchingScreeningSearching biasesScreening biases[34]47Friend control biasOccurs when there is an association of the actual intervention/exposure status between human subjects in the intervention/exposure group and their friends in the comparator group due to mismeasurements, misclassifications of or the lack of adherence to the intervention/exposurePrimary researchMeasurementPerformance biases[35]Handling data biasSee 68 ‘Missing data bias’Hawthorne effectSee 6 ‘Attention bias’48Hot stuff biasOccurs when research results on hot topics are published even though they are not validPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[28, 34, 36, 61]49Inclusion biasOccurs when inclusion of subjects or areas into a study creates an imbalance between the groups (i.e., exchangeability does not hold)Primary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[31, 35]50Inclusion control biasOccurs when some subjects or areas of the comparator group also receive the intervention/exposure to some extent (biased towards null when there is an effect)Primary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[34]51Industry sponsorship biasOccurs when research outcomes are intentionally aligned with the interests of the funding commercial manufacturer, undermining the validity of the inferencePrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[36, 62]52Information bias (Observation bias, Classification bias)Occurs when the information used in a study deviates from the truthPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[32–34, 63]53Insensitive measure biasOccurs when changes or differences in the outcome of interest are not detected accurately due to insufficient methods or inappropriate choice of an outcome variablePrimary researchMeasurementDetection biases[28, 32, 34, 36, 64]54Instrument biasOccurs when instruments used for measurement are inaccurate, too imprecise or different between groups (hence not comparable between groups)Primary researchExchangeabilityMeasurementConfounding biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[28, 32, 34, 36]55Interobserver variability biasOccurs when multiple observers' measurements of the intervention/exposure or the outcome are different, and the differences influence the estimate of an effectPrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[34]56Interpretation biasOccurs when researchers misinterpret the obtained data and resultsPrimary researchSecondary researchStatistical methodsStudy appraisal and data coding/extractionData synthesisOutcome assessment biasesStudy appraisal and data coding/extraction biasesData synthesis biases[32, 34]57Intra-observer variability biasOccurs when an observer measures the intervention/exposure or the outcome differently and the differences influence the estimate of an effectPrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[34]58Investigator biasOccurs when study investigators (measurers) are not blinded and the lack of blinding influences estimation of an effectPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[37]59Laboratory data biasOccurs when procedures implemented at a laboratory produce a systematic error in the estimate of an effect (may be confirmed by comparing data from multiple laboratories)Primary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesPerformance biasesDetection biases[34]60Lack of blinding biasOccurs when the lack of concealment of the intervention, exposure or control received by subjects or areas affects selection of subjects, areas or data, measurement of the treatment or the outcome, and/or assessment of the effectiveness or the impactPrimary researchExchangeabilityMeasurementStatistical methodsPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome assessment biases[65]Lack of intention to treat analysis biasSee 117 ‘Treatment analysis bias’61Language biasOccurs when studies with certain results or topics are more likely to be published in a certain languageSecondary researchSearchingScreeningSearching biasesScreening biases[29, 35, 66]62Latency biasOccurs when the outcome is assessed in an inappropriate time interval (i.e., the time interval is too short or long) for the intended assessment of causal effectPrimary researchSecondary researchStatistical methodsData synthesisOutcome assessment biasesData synthesis biases[34]63Literature search biasOccurs when searching of literature does not capture a representative sample of all existing studies relative to the review questionSecondary researchSearchingSearching biases[34]64Location biasOccurs when some journals are not indexed in all databases so that the choice of databases influences results of literature searchingSecondary researchSearchingSearching biases[29]‘Looking for the pony’See 32 ‘Data-dredging bias’Loss to follow-up biasSee 120 ‘Withdrawal bias’65Measurement biasOccurs when measurements of relevant variables are not accurate or precise enoughPrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[31]66Meta-analysis biasOccurs when meta-analyses based on a non-representative sample of studies are conductedPrimary researchSecondary researchSearchingScreeningStudy appraisal and data coding/extractionData synthesisSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[36]67Misclassification biasOccurs when subjects or areas, or the exposure or intervention is misclassified, and that distorts the association between the intervention/exposure and the outcomePrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biases[32, 34–36, 38, 67]68Missing data bias (Handling data bias, Data completeness bias)Occurs when there is an imbalance in missing data between the comparison groupsPrimary researchExchangeabilityPost-intervention/exposure selection biases[34]69Misuse of statisticsOccurs when descriptive or inferential statistical analysis is inappropriately conducted, resulting in unjustified and invalid conclusions about the estimate of an effectPrimary researchSecondary researchStatistical methodsData synthesisOutcome assessment biasesData synthesis biases[37]70Multiple exposure biasOccurs when another exposure acts as a confounder affecting both the exposure of interest and the outcomePrimary researchExchangeabilityConfounding biases[34]Multiple publication biasSee 39 ‘Duplication bias’Neyman biasSee 89 ‘Prevalence-incidence bias’71Non-contemporaneous control bias (Non-simultaneous comparison bias)Occurs when there is a systematic difference in the timing of observation or investigation between the intervention/exposure and the control (or alternative intervention/exposure), and exchangeability between groups cannot be assumed due to the time differencePrimary researchExchangeabilityConfounding biases[28, 34, 36, 37, 68]72Non-random sampling biasOccurs when non-random sampling results in an imbalance between the intervention/exposure and the comparator groups (i.e., exchangeability does not hold)Primary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[34, 35, 37]73Non-respondent biasOccurs when there is unequal loss of human subjects between the groups due to non-responses, or when there is a systematic difference in outcome measurement between the groups due to late and/or early responsesPrimary researchExchangeabilityMeasurementPost-intervention/exposure selection biasesDetection biases[28, 35, 37, 69]74Non-specification biasOccurs when the intended intervention or exposure is not clearly defined, and an unintentional intervention or exposure occurs that influences both the intended intervention or exposure and the outcomePrimary researchExchangeabilityConfounding biases[36]Non-simultaneous comparison biasSee 71 ‘Non-contemporaneous control bias’75Novelty biasOccurs when the interventions or exposures are reported to have a greater effect simply because they are novelPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[70]76Obsequiousness biasOccurs when human subjects alter questionnaire answers due to perception of the desired direction of the outcome held by investigatorsPrimary researchMeasurementDetection biases[28, 34–37]Observation biasSee 52 ‘Information bias’77Observer biasOccurs when there are systematic differences in the process of observing and recording information between the groups in primary researchPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[14, 34, 35]78Omitted-variable biasOccurs when one or more important explanatory variables are not included in a regression model where omission of such variables is not appropriate for estimating an effectPrimary researchStatistical methodsOutcome assessment biases[34, 35]79One-sided reference biasOccurs when citing references to only those supporting one side of available evidence and hence systematically deviating from the truthSecondary researchSearchingScreeningSearching biasesScreening biases[28, 34, 36, 71]80Optimism biasOccurs when researchers' or study subjects' beliefs that new interventions are better than established ones influence a study or reviewPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[33]81Outcome reporting biasOccurs when outcomes are selectively reported in a published documentPrimary researchReportingOutcome reporting biases[72]Outlier handling biasSee 115 ‘Tidying-up bias’82Overmatching biasOccurs when researchers match by (also ‘block on’, meaning filter by a certain value of that variable) a non-confounding variable that is associated with the intervention/exposure but not to the outcomePrimary researchExchangeabilityConfounding biases[34, 35]83Perception biasOccurs when relevant information or subjects, areas, interventions, exposures, controls, alternative interventions, alternative exposures or outcomes are misinterpreted in a study or reviewPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[73]84Performance bias (Procedure bias)Occurs when the treatment procedure of interest is altered without changing the inferential goal in an experimental studyPrimary researchMeasurementPerformance biases[32–34, 38, 74]85Positive results biasOccurs when positive results are more likely to be disseminated or submitted and accepted for publication than non-significant or negative resultsPrimary researchSecondary researchReportingSearchingScreeningOutcome reporting biasesSearching biasesScreening biases[28, 34, 36, 75]Post hoc analysis biasSee 86 ‘Post hoc analysis bias’86Post-hoc significant bias (Post hoc analysis bias)Occurs when researchers choose a critical significance level (α, 0.05 is commonly employed), and/or set a non-directional (two-tailed) or directional (one-tailed) hypothesis after some examination of data in an attempt to yield significant resultsPrimary researchSecondary researchStatistical methodsData synthesisOutcome assessment biasesData synthesis biases[28, 32, 34–37]87Pre-publication biasOccurs when previously published research, that is errant, is used to support a particular outcome in primary researchPrimary researchExchangeabilityMeasurementReportingStatistical methodsConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biases[30]88Prevailing paradigm biasOccurs when studies that are relevant to or support prevailing paradigms are more likely to be published in academic journalsPrimary researchSecondary researchReportingSearchingScreeningOutcome reporting biasesSearching biasesScreening biases[29]89Prevalence-incidence bias (Neyman bias, Selective survival bias)Occurs when uneven exclusion or attrition of subjects or areas with severe or mild responses occurs between the groups, and the assumption of exchangeability no longer holdsPrimary researchExchangeabilityPost-intervention/exposure selection biases[28, 32, 34–37, 76]Procedure biasSee 84 ‘Performance bias’90Proficiency biasOccurs when the intervention/exposure under study is unequally applied to individual subjects or areas due to interpersonal or intrapersonal differencesPrimary researchMeasurementPerformance biases[32, 34, 36]91Publication biasOccurs when the likelihood of a study being accepted for publication is influenced by the results of the studyPrimary researchSecondary researchReportingSearchingScreeningOutcome reporting biasesSearching biasesScreening biases[29, 31, 32, 35, 37, 77]92Recall biasOccurs when recall of relevant events or experiences by human subjects is inaccurate or incompletePrimary researchMeasurementMisclassified/mismeasured comparison biasesDetection biases[32–38, 78]93Reference biasOccurs when studies referenced as evidence in a review are not a representative sample of all existing studies relative to the review questionSecondary researchSearchingScreeningData synthesisSearching biasesScreening biasesData synthesis biases[29]94Regression dilution biasOccurs when imprecise measurements of the intervention/exposure are used to estimate the effect of the intervention/exposure in an observational studyPrimary researchMeasurementMisclassified/mismeasured comparison biases[35]95Relative control biasOccurs when there is an association of the actual intervention/exposure status between the intervention/exposure group and the comparator group due to mismeasurements, misclassifications of or the lack of adherence to the intervention/exposurePrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biases[35]96Repeated peeks biasOccurs when repeated peeks at accumulating data in a study or review leads to inappropriate termination of data collectionPrimary researchSecondary researchExchangeabilityMeasurementSearchingScreeningStudy appraisal and data coding/extractionConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biases[28, 34, 36]97Reporting biasOccurs when certain results of a study or certain studies are more likely to be reported or unreportedPrimary researchSecondary researchReportingSearchingScreeningOutcome reporting biasesSearching biasesScreening biases[30, 34, 79]98Research biasOccurs when studies are carried out and reported on specific organisms or systems, or in particular conditions, because of varying levels of practicality or the anticipation of specific outcomesPrimary researchSecondary researchReportingSearchingScreeningOutcome reporting biasesSearching biasesScreening biases[29]99Researcher bias (Sponsor bias)Occurs when fraud or manipulation of research design, data or results occurs due to vested interests of researchers and organisationsPrimary researchSecondary researchExchangeabilityMeasurementReportingStatistical methodsSearchingScreeningStudy appraisal and data coding/extractionData synthesisConfounding biasesPost-intervention/exposure selection biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biasesOutcome reporting biasesOutcome assessment biasesSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[30]100Resource bias (Cost bias)Occurs when available budget or resources limit searching for evidence, resulting in a systematic difference between the identified and missed studiesSecondary researchSearchingSearching biases[29]101Response fatigue biasOccurs when study subjects suffer from fatigue due to implementation of the intervention/exposure and their responses are not obtained or they did not go through the entire process of the intended intervention/exposurePrimary researchExchangeabilityMeasurementPost-intervention/exposure selection biasesPerformance biases[34]102Review biasOccurs when a researcher's prior knowledge or belief triggers misinterpretation of available data on the intervention/exposure or the outcome in primary researchPrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[34]103Reviewer biasOccurs when the selection of studies from the available evidence is not a representative sample of all existing studies relative to the review questionSecondary researchSearchingScreeningStudy appraisal and data coding/extractionData synthesisSearching biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[30]Rumination biasSee 118 ‘Underlying cause bias’Sample size biasSee 121 ‘Wrong sample size bias’104Sampling biasOccurs when exchangeability between groups does not hold due to employed sampling techniquesPrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[34]105Scale degradation biasOccurs when researchers make outcome measurement scales less precise or accurate to obscure differences between groups under comparisonPrimary researchSecondary researchMeasurementStudy appraisal and data coding/extractionDetection biasesStudy appraisal and data coding/extraction biases[28, 34, 36]106Selection biasOccurs when exchangeability between groups does not hold due to the way that selection of subjects or areas into the study or analysis is carried outPrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[33–35, 37, 38, 80]Selective survival biasSee 89 ‘Prevalence-incidence bias’107Sick quitter biasOccurs when the exposure or intervention that specific subjects receive is inaccurately measured or classified due to changes in the subjects' behaviour caused by sickness (e.g., lack of adherence, reclassification from intervention/exposure to no intervention/exposure by the researchers after the intervention/exposure is applied)Primary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biases[35]108Significance biasOccurs when statistical significance is confused with environmental, ecological, biological or agricultural significanceSecondary researchStudy appraisal and data coding/extractionData synthesisStudy appraisal and data coding/extraction biasesData synthesis biases[28, 32, 34, 36, 37]109Spatial biasOccurs when populations that are spatially distinct are compared, and this spatial difference affects both the intervention/exposure and the outcome being studiedPrimary researchExchangeabilityConfounding biases[34]110Spin biasOccurs when research findings or data are inappropriately used or interpreted to support or oppose effectiveness or impactPrimary researchSecondary researchReportingStatistical methodsScreeningStudy appraisal and data coding/extractionData synthesisOutcome reporting biasesOutcome assessment biasesScreening biasesStudy appraisal and data coding/extraction biasesData synthesis biases[81]Sponsor biasSee 99 ‘Researcher bias’111Starting time biasOccurs when starting time for the intervention/exposure or outcome measurement is different within or between groups and the time lag (or variable influenced by the time lag) influences both the intervention/exposure and the outcome or causes inaccurate measurements of the intervention/exposure or the outcomePrimary researchExchangeabilityMeasurementConfounding biasesMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[28, 34, 36, 82]112Substitution game biasOccurs when substitution of the outcome or the intervention/exposure with a surrogate is imprecise or inaccuratePrimary researchMeasurementMisclassified/mismeasured comparison biasesPerformance biasesDetection biases[28, 34, 36, 83]113Susceptibility bias (Vulnerability bias)Occurs when different study groups have different characteristics and the difference indicates that one group is more or less susceptible to the exposure (i.e., the characteristics are associated with the exposure and thus the difference creates an imbalance between the groups, meaning exchangeability does not hold)Primary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[33–36]114Temporal biasOccurs when studies with smaller *p*-values or larger effects are more likely to be published in academic journals more rapidlySecondary researchSearchingScreeningSearching biasesScreening biases[29]115Tidying-up bias (Outlier handling bias)Occurs when outliers or other untidy measurements or results are inappropriately excluded or non-reportedPrimary researchSecondary researchReportingStatistical methodsStudy appraisal and data coding/extractionData synthesisOutcome reporting biasesOutcome assessment biasesStudy appraisal and data coding/extraction biasesData synthesis biases[28, 32, 34, 36, 37]116Transfer biasOccurs when a certain aspect of a study leads to uneven losses in follow-up between the groupsPrimary researchExchangeabilityPost-intervention/exposure selection biases[32, 33, 38]117Treatment analysis bias (Lack of intention to treat analysis bias)Occurs when researchers fail to keep subjects or areas in the group they are assigned to, when there is deviation from the intended treatment, or when subjects do not adhere to the assigned treatmentPrimary researchExchangeabilityMeasurementConfounding biasesPerformance biases[34, 35]118Underlying cause bias (Rumination bias)Occurs when the intervention or exposure (group) is more carefully planned and implemented than the control (no intervention or exposure), and that results in imbalances between the groups or misclassification or mismeasurement of controlPrimary researchExchangeabilityMeasurementConfounding biasesMisclassified/mismeasured comparison biases[28, 32, 34–36]119Volunteer biasOccurs when human subjects involved in a study differ in some way among those in the intervention/exposure group or between groups due to willingness to participate as volunteersPrimary researchExchangeabilityMeasurementConfounding biasesPost-intervention/exposure selection biasesPerformance biases[28, 32, 34, 36, 37, 84]Vulnerability biasSee 113 ‘Susceptibility bias’120Withdrawal bias (Dropout bias, Loss to follow-up bias)Occurs when subjects who have dropped out of a study differ from remaining subjects, and the differences modify the effectiveness or impactPrimary researchExchangeabilityConfounding biasesPost-intervention/exposure selection biases[28, 32, 34–37]121Wrong sample size bias (Sample size bias)Occurs when used sample size is inappropriate in a study or analysis for assessing the effect of the intervention or exposure due to violation of the assumption that exchangeability holds or the hypothesis can be tested with the collected samplesPrimary researchExchangeabilityStatistical methodsConfounding biasesPost-intervention/exposure selection biasesOutcome assessment biases[28, 34, 36, 85]Synonyms are provided in brackets. Note that these general interpretations assume that no statistical corrections or adjustments are made for providing valid estimates^*^Primary research and/or secondary research. **Selected from the following items which are also provided in Fig. 2: (1) exchangeability; (2) measurement; (3) reporting; (4) statistical methods; (5) searching; (6) screening; (7) study appraisal and data coding/extraction; and (8) data synthesis. ***Domains are (1) confounding biases; (2) post-intervention/exposure selection biases; (3) misclassified/mismeasured comparison biases; (4) performance biases; (5) detection biases; (6) outcome reporting biases; (7) outcome assessment biases; (8) searching biases; (9) screening biases; (10) study appraisal and data coding/extraction biases; and (11) data synthesis biases. ****Documents and specific webpages of Catalogue of Bias that provided descriptions of the biases ### The coverage of the domains of bias There was no type of bias that was not applicable to any domain, and we did not find any types of bias that were beyond the domains, suggesting that domains were sufficient to cover the 121 types of bias. Eighty-six types of bias (71%) were applicable to multiple domains while 35 types of bias (29%) were applicable to only one domain of bias. ## Discussion ### Overview of findings We found 121 (out of 206) types of bias that were relevant to estimation of causal effects in the environmental sector. We provided a general interpretation of every relevant type of bias covered by seven risk-of-bias domains for primary research and four domains for secondary research. ### Strengths and limitations As far as we are aware this is the first study to provide a comprehensive map of the type of bias that may affect the reliability of the results of environmental research studies and reviews. It provides an important summary of the breadth of biases which may impact the findings. There are limitations that should be noted. We used a pragmatic approach to identify potentially relevant biases by utilising the Catalogue of Bias and 11 key publications. We did not conduct a systematic literature search, similar to the approach used by some of the key documents we consulted [28–33, 35–38]. However, the Catalogue of Bias which is developed by CEBM at Oxford University and is continually updated, is considered a robust source for identifying biases that affect health and medical evidence. There may be other relevant types of bias that are collated and described elsewhere, but we are not aware of any published studies specific to the environmental sector that aimed to identify types of bias. We used subjective judgements for evaluating relevance and some aspects of biases were open for interpretation although we conducted independent double checking. We provided only ‘general’ interpretations of the types of bias in this paper. ### Implications for research and practice Since education about biases is not sufficient and the level of knowledge about biases are concerning in the environmental sector [26], we hope the provision of the general interpretations of the 121 types of bias helps environmental scientists, as well as decision-makers, be better aware of the biases. When environmental scientists, practitioners, and policymakers come across certain types of bias when they read research or review papers, or when they communicate about certain types of bias, we suggest they check our dedicated list of biases relevant to estimation of causal effects in the environmental sector. Although some types of bias such as confounding bias, selection bias, and measurement bias have been formally defined [8, 17], future research is needed to formalise definitions of other types of bias such as through decomposition using mathematical formulae and/or by employing directed acyclic graphs (DAGs). For example, Suzuki et al. 2016 [8] decomposed confounding bias by comparing the causal measure (A → Y), *µ*(*E*(*Y*~1~), *E*(*Y*~0~)), and the situation where a confounding variable is present \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$ \mu \left( {E\left( {Y{|}A = 1} \right), E\left( {Y{|}A = 0} \right)} \right) - \mu \left( {E\left( {Y_{1} } \right),E\left( {Y_{0} } \right)} \right) $$\end{document}μEY|A=1,EY|A=0-μEY1,EY0where *Y* refers to outcome and *A* refers to intervention or exposure (*A* = 1 receives the intervention or exposure and *A* = 0 does not receive the intervention or exposure). Although the practicality of defining all biases remains to be evaluated, formalising definitions is particularly important because, to some extent, vagueness remains in the descriptions of the biases. This would also help identify duplication and cross over between biases identified by different researchers. As the types of bias are described theoretically, future research is also needed for collating empirical studies that evaluate the impacts of the identified biases. Such research is desired even in the health sector for better managing the risk of bias [86]. In the environmental sector, for example, Hudson et al. [87] and Takeshita et al. [88] evaluated the potential extent of selection bias and confounding bias, respectively. Konno et al. [89] and Konno and Pullin [90] evaluated the potential extent of language bias and availability bias, respectively. If such studies are collated and summarised, environmental scientists should be better able to manage the risk of bias. We hope this paper guides future research. A lack of and inappropriate conduct of risk-of-bias assessment in the environmental sector has been reported in the literature [13, 25, 91]. Our findings showed that the domains selected for use in CEE Critical Appraisal Tool and CEE Synthesis Assessment Tool are sufficient and cover all the biases identified in primary and secondary research (Fig. 2) and can be used to facilitate communication about the risk of bias. A domain-based tool should incorporate a series of checklist questions within each domain that prompts users to consider relevant types and sources of bias. Developing a series of checklist questions requires relevant expertise [20, 21], and thus, we encourage evidence providers (e.g., impact evaluators, evidence synthesists) as well as evidence consumers (e.g., decision-makers) to make use of the existing tools for assessing the risk of bias (e.g., the CEE Critical Appraisal Tool for primary research and the CEE Synthesis Appraisal Tool for secondary research), rather than developing a tool on their own, as it may result in inappropriate conduct of risk-of-bias assessment [13, 91]. However, at the time of writing, the CEE Critical Appraisal Tool is still under development, and we hope users of this tool will provide their feedback for the developers. The most important aspect for further development of the tool is inter-rater reliability (i.e., consistency between assessors) which can be quantified using a kappa statistic [92] and reported in eventual publications (e.g., Systematic Reviews [93], Systematic Maps [94]). Even in medical and health sciences, inter-rater reliability is still an issue [95, 96]. As part of our future work, we plan to consult researchers and policymakers to reach a consensus on the list of biases for better and active communication within the communities. We highlight the importance of including appropriately qualified and experienced stakeholders in any consultation because it is likely that many people working in environmental research and policy are inadequately familiar with biases [26], and thus there is a need to educate environmental scientists and decision-makers about the types of bias. There is also a need to provide training and develop capacity for use of risk-of-bias assessment tools in the environmental sciences. We are not aware of any institution that teaches this subject in environmental contexts, even at Masters level. In response to this gap, CEE intends to develop a series of online training videos on critical appraisal. If many more scholars, institutions as well as learned societies can help the environmental sector better understand the nature of biases, more active communication of the risk of bias may be achieved. ## Supplementary Information **Additional file 1.** Excluded types of bias.**Additional file 2.** Included types of bias.