Authors: Jacqueline Dinnes, Jonathan J Deeks, Sarah Berhane, Melissa Taylor, Ada Adriano, Clare Davenport, Sabine Dittrich, Devy Emperador, Yemisi Takwoingi, Jane Cunningham, Sophie Beese, Julie Domen, Janine Dretzke, Lavinia Ferrante di Ruffano, Isobel M Harris, Malcolm J Price, Sian Taylor-Phillips, Lotty Hooft, Mariska MG Leeflang, Matthew DF McInnes, René Spijker, Ann Van den Bruel
Categories: Child health, Diagnosis, Infectious disease, COVID-19
Source: The Cochrane Database of Systematic Reviews
Accurate rapid diagnostic tests for SARS‐CoV‐2 infection could contribute to clinical and public health strategies to manage the COVID‐19 pandemic. Point‐of‐care antigen and molecular tests to detect current infection could increase access to testing and early confirmation of cases, and expediate clinical and public health management decisions that may reduce transmission.
To assess the diagnostic accuracy of point‐of‐care antigen and molecular‐based tests for diagnosis of SARS‐CoV‐2 infection. We consider accuracy separately in symptomatic and asymptomatic population groups.
Electronic searches of the Cochrane COVID‐19 Study Register and the COVID‐19 Living Evidence Database from the University of Bern (which includes daily updates from PubMed and Embase and preprints from medRxiv and bioRxiv) were undertaken on 30 Sept 2020. We checked repositories of COVID‐19 publications and included independent evaluations from national reference laboratories, the Foundation for Innovative New Diagnostics and the Diagnostics Global Health website to 16 Nov 2020. We did not apply language restrictions.
We included studies of people with either suspected SARS‐CoV‐2 infection, known SARS‐CoV‐2 infection or known absence of infection, or those who were being screened for infection. We included test accuracy studies of any design that evaluated commercially produced, rapid antigen or molecular tests suitable for a point‐of‐care setting (minimal equipment, sample preparation, and biosafety requirements, with results within two hours of sample collection). We included all reference standards that define the presence or absence of SARS‐CoV‐2 (including reverse transcription polymerase chain reaction (RT‐PCR) tests and established diagnostic criteria).
Studies were screened independently in duplicate with disagreements resolved by discussion with a third author. Study characteristics were extracted by one author and checked by a second; extraction of study results and assessments of risk of bias and applicability (made using the QUADAS‐2 tool) were undertaken independently in duplicate. We present sensitivity and specificity with 95% confidence intervals (CIs) for each test and pooled data using the bivariate model separately for antigen and molecular‐based tests. We tabulated results by test manufacturer and compliance with manufacturer instructions for use and according to symptom status.
Seventy‐eight study cohorts were included (described in 64 study reports, including 20 pre‐prints), reporting results for 24,087 samples (7,415 with confirmed SARS‐CoV‐2). Studies were mainly from Europe (n = 39) or North America (n = 20), and evaluated 16 antigen and five molecular assays.
We considered risk of bias to be high in 29 (37%) studies because of participant selection; in 66 (85%) because of weaknesses in the reference standard for absence of infection; and in 29 (37%) for participant flow and timing. Studies of antigen tests were of a higher methodological quality compared to studies of molecular tests, particularly regarding the risk of bias for participant selection and the index test. Characteristics of participants in 35 (45%) studies differed from those in whom the test was intended to be used and the delivery of the index test in 39 (50%) studies differed from the way in which the test was intended to be used. Nearly all studies (97%) defined the presence or absence of SARS‐CoV‐2 based on a single RT‐PCR result, and none included participants meeting case definitions for probable COVID‐19.
Antigen tests
Forty‐eight studies reported 58 evaluations of antigen tests. Estimates of sensitivity varied considerably between studies. There were differences between symptomatic (72.0%, 95% CI 63.7% to 79.0%; 37 evaluations; 15530 samples, 4410 cases) and asymptomatic participants (58.1%, 95% CI 40.2% to 74.1%; 12 evaluations; 1581 samples, 295 cases). Average sensitivity was higher in the first week after symptom onset (78.3%, 95% CI 71.1% to 84.1%; 26 evaluations; 5769 samples, 2320 cases) than in the second week of symptoms (51.0%, 95% CI 40.8% to 61.0%; 22 evaluations; 935 samples, 692 cases). Sensitivity was high in those with cycle threshold (Ct) values on PCR ≤25 (94.5%, 95% CI 91.0% to 96.7%; 36 evaluations; 2613 cases) compared to those with Ct values >25 (40.7%, 95% CI 31.8% to 50.3%; 36 evaluations; 2632 cases). Sensitivity varied between brands. Using data from instructions for use (IFU) compliant evaluations in symptomatic participants, summary sensitivities ranged from 34.1% (95% CI 29.7% to 38.8%; Coris Bioconcept) to 88.1% (95% CI 84.2% to 91.1%; SD Biosensor STANDARD Q). Average specificities were high in symptomatic and asymptomatic participants, and for most brands (overall summary specificity 99.6%, 95% CI 99.0% to 99.8%).
At 5% prevalence using data for the most sensitive assays in symptomatic people (SD Biosensor STANDARD Q and Abbott Panbio), positive predictive values (PPVs) of 84% to 90% mean that between 1 in 10 and 1 in 6 positive results will be a false positive, and between 1 in 4 and 1 in 8 cases will be missed. At 0.5% prevalence applying the same tests in asymptomatic people would result in PPVs of 11% to 28% meaning that between 7 in 10 and 9 in 10 positive results will be false positives, and between 1 in 2 and 1 in 3 cases will be missed.
No studies assessed the accuracy of repeated lateral flow testing or self‐testing.
Rapid molecular assays
Thirty studies reported 33 evaluations of five different rapid molecular tests. Sensitivities varied according to test brand. Most of the data relate to the ID NOW and Xpert Xpress assays. Using data from evaluations following the manufacturer’s instructions for use, the average sensitivity of ID NOW was 73.0% (95% CI 66.8% to 78.4%) and average specificity 99.7% (95% CI 98.7% to 99.9%; 4 evaluations; 812 samples, 222 cases). For Xpert Xpress, the average sensitivity was 100% (95% CI 88.1% to 100%) and average specificity 97.2% (95% CI 89.4% to 99.3%; 2 evaluations; 100 samples, 29 cases). Insufficient data were available to investigate the effect of symptom status or time after symptom onset.
Antigen tests vary in sensitivity. In people with signs and symptoms of COVID‐19, sensitivities are highest in the first week of illness when viral loads are higher. The assays shown to meet appropriate criteria, such as WHO's priority target product profiles for COVID‐19 diagnostics (‘acceptable’ sensitivity ≥ 80% and specificity ≥ 97%), can be considered as a replacement for laboratory‐based RT‐PCR when immediate decisions about patient care must be made, or where RT‐PCR cannot be delivered in a timely manner. Positive predictive values suggest that confirmatory testing of those with positive results may be considered in low prevalence settings. Due to the variable sensitivity of antigen tests, people who test negative may still be infected.
Evidence for testing in asymptomatic cohorts was limited. Test accuracy studies cannot adequately assess the ability of antigen tests to differentiate those who are infectious and require isolation from those who pose no risk, as there is no reference standard for infectiousness. A small number of molecular tests showed high accuracy and may be suitable alternatives to RT‐PCR. However, further evaluations of the tests in settings as they are intended to be used are required to fully establish performance in practice.
Several important studies in asymptomatic individuals have been reported since the close of our search and will be incorporated at the next update of this review. Comparative studies of antigen tests in their intended use settings and according to test operator (including self‐testing) are required.
Severe acute respiratory syndrome coronavirus 2 (SARS‐CoV‐2) and the resulting COVID‐19 pandemic present important diagnostic evaluation challenges. These range understanding the value of signs and symptoms in predicting possible infection; assessing whether existing biochemical and imaging tests can identify infection or people needing critical care; and evaluating whether in vitro diagnostic tests can accurately identify and rule out current SARS‐CoV‐2 infection, and identify those with past infection, with or without immunity.
We are creating and maintaining a suite of living systematic reviews to cover the roles of tests and patient characteristics in the diagnosis of COVID‐19. This review is the first update of a review summarising evidence of the accuracy of rapid antigen and molecular tests that are suitable for use at the point of care. In some scenarios the tests could potentially be used as alternatives to standard laboratory‐based molecular assays, such as reverse transcription polymerase chain reaction (RT‐PCR) assays, that are relied on for identifying current infection, in others they may be used where no testing is currently done. If sufficiently accurate, point‐of‐care tests have the potential to greatly expand access and speed of testing, In turn, if accurate, they may have greater impact on public health than laboratory‐based molecular methods as they are less expensive, provide results more quickly and do not require the same technical expertise and laboratory capacity. These tests can be undertaken locally, avoiding the need for centralised testing facilities that rarely meet the needs of patients, caregivers, health workers and society as a whole, especially in low‐ and middle‐income countries. As these are rapid tests, their results can be returned within the same clinical encounter, facilitating timely decisions concerning the need for isolation and contract tracing activities.
COVID‐19 is the disease caused by infection with the SARS‐CoV‐2 virus. The key target conditions for this suite of reviews are current SARS‐CoV‐2 infection, current COVID‐19 disease, and past SARS‐CoV‐2 infection. The tests included in this review concern the identification of current infection, as defined by reference standard methods of diagnosis, including molecular assays such as RT‐PCR, or internationally recognised clinical guidelines for diagnosis of SARS‐CoV‐2. In the context of test evaluation, and throughout this review, we use the term 'reference standard' to denote the best available method (test or tests) for diagnosing the target condition, as opposed to other uses of the term in diagnostic virology (such as reference methods or reference materials).
For current infection, the severity of the disease is of ultimate importance for patient outcomes. However, rapid testing does not establish severity of disease, and for this review we consider the role of point‐of‐care tests for detecting SARS‐CoV‐2 infection of any severity, distinguishing only between symptomatic and asymptomatic infection.
COVID‐19 public health interventions focus on reducing disease transmission, thus it is important to identify and isolate people who are infected before or whilst they are infectious. It is reasonably presumed that people with symptoms who meet national criteria for COVID‐19 testing, or who are identified through contact tracing, have a high enough risk of being infectious to ask them to isolate. However, assessing the risk of an individual being infectious in asymptomatic screening is more difficult, as there is no reference standard test for being ‘infectious’. Using RT‐PCR status as a reference standard (as is done for target condition of ‘infection’) will ensure that infectious people are not missed, but as RT‐PCR continues to detect viral RNA days and weeks after the onset of infection will wrongly classify some people as infectious. Alternative reference standards that have been proposed for infectiousness include assessing the viability of the virus using viral culture, or using a value of the cycle threshold (Ct value) from RT‐PCR results to group individuals above or below a particular value (as a proxy for viral load) as more or less likely to be infectious. Converting Ct values (also known as quantification cycle (Cq) or crossing point (Cp) values) into direct quantitative values of viral load (viral copies per cell) is possible but challenging, as the relationship between Ct values and viral load varies between machines and laboratories. Thus comparison at fixed Ct values is unlikely to be comparable across studies. Viral culture is unsuitable as a reference standard because it is technically complex and often unreliable, which leads to it being an insensitive test (the failure to culture virus potentially being a result of the culture technique and not an indicator of non‐infectiousness). The suitability of RT‐PCR is limited as the inverse relationship between viral load (Ct value) and risk of infection is a continuum of risk without there being a meaningful cut‐point (with virus being cultured from samples with Ct values as high as 35 (Singanayagam 2020)). Similarly, those with low viral loads at the onset of infection will be missed. A preferable alternative, of tracking contacts for evidence of secondary infections, requires longitudinal follow‐up and is better considered as a question about risk of transmission, which can be addressed using predictive modelling approaches (taking into account host, agent and environmental factors). This is in contrast to the diagnostic test accuracy paradigm which can only determine if individuals are infected at a single point in time.
For these reasons, this review only focuses on the target condition of 'infection' for both symptomatic and asymptomatic applications of tests. We do report results where they are presented split by an RT‐PCR Ct value to report on accuracy according to groups with higher and lower viral load, but advise caution on their interpretation considering the lack of standardisation of PCR Ct values. Given the current state of the scientific knowledge we do not consider it appropriate to consider these as groups which are defined as 'infectious' and 'not infectious'.
RT‐PCR carries a very small risk of false positive results for infection and a higher risk of false negative results. False positive results may result from failures in sampling or laboratory protocols (e.g. mislabelling), contamination during sampling or processing, or low‐level reactions during PCR (Healy 2020; Mayers 2020). At times when SARS‐CoV‐2 infections have been rare, population prevalence surveys using RT‐PCR have shown test positivity rates of 0.44% (95% credible 0.22% to 0.76%) (August 2020; ONS 2020), and 0.077% (0.065%, 0.092%) (June to July 2020; Riley 2020 React‐1 study). These values can be used to place an upper bound on the possible false positive rate of RT‐PCR of less than 0.077% (as the total numbers testing positive will comprise both true positive and false positive RT‐PCR results). The World Health Organization (WHO) recently issued a notice of concern regarding interpretation of specimens at or near the limit for PCR positivity (i.e. those with high cycle threshold (Ct) values), citing potential difficulties in distinguishing the presence of the target virus from these types of background ‘noise’ (WHO 2020a). False negative rates have been estimated by looking at individuals with symptoms who initially test negative, but positive on a subsequent test. These rates have been estimated to be as high as 20% to 30% in the first week of symptom onset; Arevalo‐Rodriguez 2020; Yang 2020a; Zhao 2020; Kucirka 2020). Including probable COVID‐19 cases within the target condition, as defined by internationally recognised clinical guidelines for diagnosis of SARS‐CoV‐2 will partially mitigate these missed cases.
The primary consideration for the eligibility of tests for inclusion in this review is that they should detect current infection and should have the capacity to be performed at the ‘point of care’ or in a ‘near‐patient’ testing role. There is an ongoing debate around the specific use and definitions of these terms, therefore for the purposes of this review, we consider ‘point‐of‐care’ and ‘near patient’ to be synonymous, but for consistency and avoidance of confusion, we use the term ‘point‐of‐care’ throughout.
We have adapted a definition of point‐of‐care testing, namely that it “refers to decentralized testing that is performed by a minimally trained healthcare professional near a patient and outside of central laboratory testing” (WHO 2018), with the additional caveat that test results must be available within a single clinical encounter (Pai 2012). Our criteria for defining a point‐of‐care test are
Tests for detection of current infection that are currently suitable for use at the point of care include antigen tests and molecular‐based tests. Both types of test use the same respiratory‐tract samples acquired by swabbing, washing or aspiration as for laboratory‐based RT‐PCR. Rapid antigen tests use lateral flow immunoassays, which are disposable devices, usually in the form of plastic cassettes akin to a pregnancy test. Viral antigen is captured by dedicated antibodies that are either colloidal gold‐ or fluorescent‐labelled. Antigen detection is indicated by visible lines appearing on the test strip (colloidal gold‐based immunoassays, or CGIA), or through fluorescence, which can be detected using an immunofluorescence analyser (fluorescence immunoassays or FIA). Molecular‐based tests to detect viral ribonucleic acid (RNA) have historically been laboratory‐based assays using RT‐PCR technology (see Alternative test(s)). In recent years, automated, single‐step RT‐PCR methods have been developed, as well as other nucleic acid amplification methods, such as isothermal amplification, that do not require the sophisticated thermo cycling involved in RT‐PCR (Green 2020). These technological advances have allowed molecular technologies to be developed that are suitable for use in a point‐of‐care context (Kozel 2017), however they still require small portable machines and many take longer to produce results than antigen tests.
Following the emergence of COVID‐19 there has been prolific industry activity to develop accurate tests. The Foundation for Innovative Diagnostics (FIND) and Johns Hopkins Centre for Health Security have maintained online lists of available tests for SARS‐CoV‐2 (FIND 2020). At the time of writing (5 January 2021), FIND listed 129 rapid antigen tests, 118 of which are described as "commercialised" and 92 have been identified as having regulatory approval. These numbers are a substantial increase on the 48 listed, 32 commercialised and 21 with regulatory approval at the time of our original review (19 July 2020). A total of 142 molecular tests were described as automated, including both laboratory‐based assays and assays suitable for use outside of a laboratory setting (i.e. near or at the point of care). Further information from FIND indicates that 53 of the 142 assays were categorised as point‐of‐care or near point‐of‐care tests, including 43 with regulatory approval. This classification was based on the information provided to FIND by the test manufacturers and does not necessarily mean that these tests meet the criteria for point‐of‐care tests that we have specified for this review. The numbers of tests of these types will continue to increase over time.
Given the urgent need to identify the evidence base for tests that are available for purchase, the focus of this first update of the review is on tests that are commercially produced. All commercially produced assays are supplied with a specific product code, product inserts or instructions for use (IFU) sheets that document the intended use of the test; sample storage and preparation and testing procedures; who should deliver the test and in whom; and any restrictions around the type of samples that can be used.
There are many proposals for serial testing with lateral flow tests to detect infection, rather than a single use. In this case it would be appropriate to evaluate the accuracy of the strategy rather than a single test.
Patients may be tested for SARS‐CoV‐2 when they present with symptoms, have had known exposure to a confirmed case, or in a screening context, with no known exposure to SARS‐CoV‐2. The standard approach to diagnosis of SARS‐CoV‐2 infection is through laboratory‐based testing of swab samples taken from the upper respiratory (e.g. nasopharynx, oropharynx) or lower respiratory tract (e.g. bronchoalveolar lavage or sputum) with RT‐PCR. RT‐PCR is the primary method for detecting infection during the acute phase of the illness while the virus is still present. Both the WHO and the China CDC (National Health Commission of the People's Republic of China), have produced case definitions for COVID‐19 that include the presence of convincing clinical evidence (some including positive serology tests) when RT‐PCR is negative (Appendix 1).
Signs and symptoms are used in the initial diagnosis of suspected SARS‐CoV‐2 infection and to help identify those requiring tests. A number of key symptoms have been suggested as indicators of mild to moderate COVID‐19, cough, fever greater than 37.8 °C, headache, breathlessness, muscle pain, fatigue, and loss of sense of smell and taste (Struyf 2021). However, the recently published review of signs and symptoms found good evidence for the accuracy for these symptoms alone or in combination to be lacking (Struyf 2021).
Where people are asymptomatic but are being tested as part of screening (e.g. universal testing of students as part of a risk‐reduction effort) or on the basis of epidemiological risk factors, such as exposure to someone with confirmed SARS‐CoV‐2 or following travel to more highly endemic countries, no prior tests will have been conducted.
For most settings in which testing for acute SARS‐CoV‐2 infection in symptomatic individuals takes place, results of molecular laboratory‐based RT‐PCR tests are unlikely to be available within a single clinical encounter. Point‐of‐care tests potentially have a role either as a replacement for RT‐PCR (if sufficiently accurate), or as a means of triaging and rapid management (quarantine or treatment, or both), with confirmatory RT‐PCR testing for those with negative rapid test results (CDC 2020; WHO 2020b). Obtaining quick results within a healthcare visit will allow faster decisions about isolation and healthcare interventions for those with positive test results, and allow contact tracing to begin in a more timely manner. Modelling studies suggest contact tracing is most effective if it starts within 24 hours of case detection, with delays in testing (e.g. due to laboratory turnaround time for reporting PCR results) leading to reductions in the proportion of onward transmissions per index case that can be prevented by track and trace (Kretzschmar 2020).
If sufficiently accurate, negative rapid test results in symptomatic patients could allow faster return to work or school, therefore conferring important economic and educational implications. Negative results also allow immediate consideration of other causes of symptoms, which may be time‐sensitive, for example bacterial pneumonia or thrombo‐embolism.
For asymptomatic individuals, if accurate, rapid tests may also be considered for screening at‐risk (exposed) populations, for example in hospital workers or in local outbreaks.
Rapid tests, particularly antigen tests which can be more easily delivered at scale, could also be used for mass screening purposes as recently piloted in Slovakia and in Liverpool UK (University of Liverpool 2020), or used in a more targeted fashion such as single test application at airports or for border entry, to allow entry to large public gatherings, or screening students as a risk‐reduction strategy (Ferguson 2020). Preliminary data on the rollout of such a policy in the UK has highlighted the many challenges in such an approach (Deeks 2020a; Nabavi 2021), and the requirement for full and proper field trial evaluations. Frequent repeated use of antigen tests in asymptomatic individuals with no known exposure to identify COVID‐19 cases has also been proposed (Larremore 2020), but field trial evaluations would be required to determine whether promising results from modelling studies can be borne out in practical settings (Crozier 2021).
This review is one of seven that cover the range of tests and clinical characteristics being considered in the management of COVID‐19 (Deeks 2020b; McInnes 2020), five of which have already been published (Deeks 2020c; Salameh 2020; Stegeman 2020; Struyf 2021), including the first iteration of this review (Dinnes 2020). Full details of the alternative tests and evidence of their accuracy is summarised in these reviews. The SARS‐CoV‐2‐specific biomarker tests that might be considered as alternatives to point‐of‐care tests are considered here.
RT‐PCR tests for SARS‐CoV‐2 identify viral ribonucleic acid (RNA). Reagents for RT‐PCR were rapidly produced once the viral RNA sequence was published (Corman 2020). Testing is undertaken in central laboratories and can be very labour‐intensive, with several points along the path of performing a single test where errors may occur, although some automation of parts of the process is possible. The amplification process requires thermal cycling equipment to allow multiple temperature changes within a cycle, with cycles repeated up to 40 times until viral DNA is detected (Carter 2020). Although the amplification process for RT‐PCR can be completed in a relatively short timeframe, the stages of extraction, sample processing and data management (including reporting) mean that test results are typically only available in 24 to 48 hours. Where testing is undertaken in a centralised laboratory, transport times increase this further. The time to result for fully automated RT‐PCR assays is shorter than for manual RT‐PCR, however most assays still require sample preparation steps that make them unsuitable for use at the point of care. Other nucleic acid amplification methods, including loop‐mediated isothermal amplification (LAMP), or CRISPR‐based nucleic acid detection methods, that allow amplification at a constant temperature are now commercially available (Chen 2020). These methods have the potential to reduce the time to produce test results after extraction and sample processing to minutes, but the time for the whole process may still be significant. Laboratory‐based molecular tests are most often applied to upper and lower respiratory samples although they are also being used on faecal and urine samples.
Serology tests to measure antibodies to SARS‐CoV‐2 have been evaluated in people with active infection and in convalescent cases (Deeks 2020c). Antibodies are formed by the body's immune system in response to infections, and can be detected in whole blood, plasma or serum. Antibody tests are available for laboratory use including enzyme‐linked immunosorbent assay (ELISA) methods, or more advanced chemiluminescence immunoassays (CLIA). There are also rapid lateral flow assays (LFA)s for antibody testing that use a minimal amount of whole blood, plasma or serum on a testing strip as opposed to the respiratory specimens that are used for rapid antigen tests; all assays for antibody detection are considered in Deeks 2020c.
It is essential to understand the clinical accuracy of tests and clinical features to identify the best way they can be used in different settings to develop effective diagnostic and management pathways for SARS‐CoV‐2 infection and disease. The suite of Cochrane living systematic reviews summarises evidence on the clinical accuracy of different tests and diagnostic features. Estimates of accuracy from these reviews will help inform diagnosis, screening, isolation, and patient‐management decisions.
The first iteration of this review (Dinnes 2020), included 22 publications reporting on a total of 18 study cohorts with 3198 unique samples, 1775 of which had confirmed SARS‐CoV‐2 infection. We identified data for eight commercial tests (four antigen and four molecular) and one in‐house antigen test.
We did not find any studies at low risk of bias and had concerns about applicability of results across all studies. We judged patient selection to be at high risk of bias in 50% of the studies because of deliberate oversampling of samples with confirmed SARS‐CoV‐2 infection (sample enrichment) and unclear in 38% (7/18) because of poor reporting. Sixteen (89%) studies used only a single, negative RT‐PCR to confirm the absence of SARS‐CoV‐2 infection, risking missing infection. There was a lack of information on blinding of index test (n = 11), and about participant exclusions from analyses (n = 10). We did not observe differences in methodological quality between antigen and molecular test evaluations.
The eight evaluations of antigen tests reported considerable variation in sensitivity across studies (from 0% to 94%) with less variation in specificities (from 90% to 100%). The average sensitivity was 56.2% (95% CI 29.5 to 79.8%) and average specificity was 99.5% (95% CI 98.1% to 99.9%) (based on 943 samples, 596 with confirmed SARS‐CoV‐2). Data for individual antigen tests were limited with no more than two studies for any test.
We observed less variation in sensitivities across 13 evaluations of rapid molecular assays (range 68% to 100%) with similar variation in specificities (range 92% to 100%). Average sensitivity was 95.2% (95% CI 86.7% to 98.3%) and specificity 98.9% (95% CI 97.3% to 99.5%) based on a total of 2255 samples.
We were able to calculate pooled results for only two molecular ID NOW (Abbott Laboratories; 5 evaluations) and Xpert Xpress (Cepheid Inc; 6 evaluations). Summary sensitivity for the Xpert Xpress assay (99.4%, 95% CI 98.0% to 99.8%) was 22.6 (95% CI 18.8 to 26.3) percentage points higher than that of ID NOW (76.8%, (95% CI 72.9% to 80.3%), whilst the specificity of Xpert Xpress (96.8%, 95% CI 90.6% to 99.0%) was marginally lower than ID NOW (99.6%, 95% CI 98.4% to 99.9%; a difference of −2.8 percentage points (95% CI from 6.4 percentage points lower to 0.8 higher).
There has been a considerable increase in the number of evaluations available of antigen tests, and a lesser rise in the number of evaluations of molecular tests. More studies report key population features such as setting, and symptom status, and there has been an increase in direct swab testing as would occur in a point‐of‐care setting. However, due to the nature of sampling and the use of direct swab testing, few comparative studies are available. This review considers the available evidence in relevant population groups and settings according to test brand and compliance with manufacturer IFUs. We used the WHO's priority target product profiles for COVID‐19 diagnostics (i.e. acceptable performance criterion of sensitivity ≥ 80% and specificity ≥ 97%, or desirable criterion of ≥ 80% sensitivity and ≥ 99% specificity; WHO 2020c) as a benchmark against which to consider test performance.
We will update this review as often as is feasible to ensure that it provides current evidence about the accuracy of point‐of‐care tests.
This review follows a generic protocol that covers six of the seven Cochrane COVID‐19 diagnostic test accuracy reviews (Deeks 2020b). The Background and Methods sections of this review therefore use some text that was originally published in the protocol (Deeks 2020b), and text that overlaps some of our other reviews (Deeks 2020c; Struyf 2021).
To assess the diagnostic accuracy of rapid point‐of‐care antigen and molecular‐based tests to determine if a person presenting in the community or in primary or secondary care has current SARS‐CoV‐2 infection, and to consider accuracy separately in symptomatic and asymptomatic population groups.
We estimated accuracy overall and separately according to symptom status (symptomatic and asymptomatic). Although we might expect to see differences in accuracy for testing of asymptomatic individuals with an epidemiological exposure to SARS‐CoV‐2 (targeted screening) compared to testing of asymptomatic individuals in a population screening setting, we did not anticipate finding sufficient numbers of studies for each testing application to allow any such difference to be explored. We will revisit this decision in subsequent iterations of this review.
Where data are available, we will investigate potential sources of heterogeneity that may influence diagnostic accuracy (either by stratified analysis or meta‐regression) according to test method and index test, participant or sample characteristics (duration of symptoms and viral load), study setting, study design and reference standard used.
We investigated adherence to manufacturers' IFUs in sensitivity analyses.
We applied broad eligibility criteria to include all patient groups (that is, if patient population was unclear, we included the study) and all variations of a test.
We included studies of all designs that produce estimates of test accuracy or provide data from which we can compute estimates, including the following.
We excluded studies from which we could not extract data to compute either sensitivity or specificity.
We carefully considered the limitations of different study designs in the quality assessment and analyses.
We included studies reported in published journal papers, as preprints, and publicly available reports from independent bodies.
We included studies recruiting people presenting with suspicion of current SARS‐CoV‐2 infection or those recruiting populations where tests were used to screen for disease (for example, contact tracing or community screening).
We also included studies that recruited people known to have SARS‐CoV‐2 infection and known not to have SARS‐CoV‐2 infection (i.e. cases only or multi‐group studies).
We excluded small studies with fewer than 10 samples or participants. Although the size threshold of 10 is arbitrary, such small studies are likely to give unreliable estimates of sensitivity or specificity and may be biased.
We included studies evaluating any rapid antigen or molecular‐based test for diagnosis of SARS‐CoV‐2, if it met the criteria outlined in the Background, that
All sample types (respiratory or non‐respiratory) were eligible. Strategies based on multiple applications of a test were also eligible for inclusion.
The target condition was current SARS‐CoV‐2 infection (either symptomatic or asymptomatic). We also refer to SARS‐CoV‐2 infection as ‘COVID‐19 infection’, particularly in the Plain Language Summary and Table 1.
We anticipated that studies would use a range of reference standards to define both the presence and absence of SARS‐CoV‐2 infection. For the QUADAS‐2 (Quality Assessment tool for Diagnostic Accuracy Studies; Whiting 2011), assessment we categorised each method of defining the presence of SARS‐CoV‐2 according to the risk of bias (the chances that it would misclassify the presence or absence of infection) and whether it defined COVID‐19 in an appropriate way that reflected cases encountered in practice. Likewise, we considered the risk of bias in definitions of the absence of SARS‐CoV‐2, and whether the definition captured all those who might be tested in practice.
Evaluations of molecular tests generally consider agreement between molecular assays, for example, agreement of a new rapid test against a more standard RT‐PCR test. For the purposes of this review, we considered RT‐PCR to be the ‘reference standard’ for SARS‐CoV‐2 infection, and present results as ‘sensitivity’ and ’specificity’ as opposed to percentage agreement. The result of further RT‐PCR analysis of discrepant cells (samples with results disagreeing on the rapid test and the RT‐PCR) were also considered in sensitivity analyses. As discrepant analysis involves retesting only a sub‐sample of patients selected according to index and reference standard results, it can introduce bias (Hadgu 1999). Retesting of all samples with a second test in a composite reference standard would be preferable when there are concerns over the accuracy of the first reference test.
We used two main sources for our electronic searches through 30 September 2020, which were devised with the help of an experienced Cochrane Information Specialist with diagnostic test accuracy review expertise (RSp). These searches aimed to identify all articles related to COVID‐19 and SARS‐CoV‐2 and were not restricted to those evaluating a particular type of test. Thus, the searches used no terms that specifically focused on an index test, diagnostic accuracy or study methodology.
We used the Cochrane COVID‐19 Study Register (covid-19.cochrane.org/), for searches conducted from inception of the Register to 28 March 2020. At that time, the register was populated by searches of PubMed, as well as trials registers at US National Institutes of Health Ongoing Trials Register ClinicalTrials.gov (clinicaltrials.gov) and the WHO International Clinical Trials Registry Platform (apps.who.int/trialsearch).
Search strategies were designed for maximum sensitivity, to retrieve all human studies on COVID‐19 and with no language limits. See Appendix 2.
From 28 March 2020, we used the COVID‐19 Living Evidence database from the Institute of Social and Preventive Medicine (ISPM) at the University of Bern (www.ispm.unibe.ch), as the primary source of records for the Cochrane COVID‐19 diagnostic test accuracy reviews. This search includes PubMed, Embase, and preprints indexed in bioRxiv and medRxiv databases. The strategies as described on the ISPM website are described here (ispmbern.github.io/covid-19/). See Appendix 3. To ensure comprehensive coverage we also downloaded records from the ‘Bern feed’ from 1 January to 28 March 2020 and de‐duplicated them against those obtained via the Cochrane COVID‐19 Study Register.
Due to the increased volume of published and preprint articles, from 25 May 2020 onwards we used artificial intelligence text analysis to conduct an initial classification of documents, based on their title and abstract information, for relevant and irrelevant documents (Appendix 4).
The decision to focus primarily on the Bern feed was because of the exceptionally large numbers of COVID‐19 studies available only as preprints. We are continuing to monitor the coverage of the Cochrane COVID‐19 Study Register and may move back to it as the primary source of records for subsequent review updates.
Prior to 28 March 2020 (when we began using the ‘Bern feed’), we identified Embase records through the Centers for Disease Control and Prevention (CDC), Stephen B Thacker CDC Library, COVID‐19 Research Articles Downloadable Database (cdc.gov/library/researchguides/2019novelcoronavirus/researcharticles.html), and de‐duplicated them against results from the Cochrane COVID‐19 Study Register. See Appendix 5.
We also checked our search results against two additional repositories of COVID‐19 publications up to 30 September
Both repositories allow their contents to be filtered according to studies potentially relating to diagnosis, and both have agreed to provide us with updates of new diagnosis studies added.
We have also contacted or accessed the websites of independent research groups undertaking test evaluations (for example, UK Public Health England (PHE), the Société Française Microbiologie (SFM), the Dutch National Institute for Public Health and the Environment (RIVM)) and studies co‐ordinated by FIND (finddx.org/covid-19/sarscov2-eval) and accessed the Diagnostics Global Health listing of manufacturer independent evaluations of antigen detecting rapid diagnostic tests (Ag‐RDTs) for SARS‐CoV‐2 (diagnosticsglobalhealth.org). We last accessed these additional resources on 16 November 2020.
We appeal to researchers to supply details of additional published or unpublished studies at the following email address, which we will consider for inclusion in future updates (coviddta@contacts.bham.ac.uk).
A team of experienced systematic review authors from the University of Birmingham screened the titles and abstracts of all records retrieved from the literature searches following the application of artificial intelligence text analysis (described in Electronic searches). Two review authors independently screened studies in Covidence. A third, senior review author resolved any disagreements. We tagged all records selected as potentially eligible according to the Cochrane COVID‐19 diagnostic test accuracy review(s) for which they might be eligible and we then exported them to separate Covidence reviews for each review title.
We obtained the full texts for all studies flagged as potentially eligible. Two review authors independently screened the full texts for one of the COVID‐19 biomarker reviews (molecular, antigen or antibody tests). We resolved any disagreements on study inclusion through discussion with a third review author.
One review author extracted the characteristics of each study, which a second review author checked. Items that we extracted are listed in Appendix 6.
Both review authors independently performed data extraction of 2x2 contingency tables of the number of true positives, false positives, false negatives and true negatives. They resolved disagreements by discussion. Where possible, we separately extracted data according to symptom status (symptomatic, asymptomatic, mixed symptom status or not reported), viral load (high or low, according to Ct cut‐offs defined within each study), and time post‐symptom onset (week one versus week two) and for molecular assays, before and after re‐analysis of samples in discrepant cells. For categorisation by symptom status, we classed studies reporting at least 75% of participants as symptomatic as ‘mainly symptomatic', we considered studies with less than 75% symptomatic participants to report ‘mixed’ groups along with those that reported recruiting both symptomatic and asymptomatic participants but did not provide the percentages in each group. We considered studies that provided no information as to the symptom status of included participants ‘not reported’. We also coded evaluations according to compliance with manufacturer IFUs. We based coding on three aspects of
We encourage study authors to contact us regarding missing details on the included studies (coviddta@contacts.bham.ac.uk).
Two review authors independently assessed risk of bias and applicability concerns using the QUADAS‐2 checklist tailored to this review (Appendix 7; Whiting 2011). The two review authors resolved any disagreements by discussion.
Ideally, studies examining the use of tests in symptomatic people should prospectively recruit a representative sample of participants presenting with signs and symptoms of COVID‐19, either in community or primary care settings or to a hospital setting, and they should clearly record the time of testing after the onset of symptoms. Studies in asymptomatic people at risk of infection should document time from exposure. Studies applying tests in a screening setting should document eligibility criteria for screening, particularly if a targeted approach is used and should take care to record any previous confirmed or suspected SARS‐CoV‐2 infection or any relevant epidemiological exposures. Studies should perform tests in their intended use setting, using appropriate samples with or without viral transport medium and within the time period following specimen collection as indicated in the IFU document. Tests should be performed by relevant personnel (e.g. healthcare workers), and should be interpreted blinded to the final diagnosis (presence or absence of SARS‐CoV‐2). The reference standard diagnosis should be blinded to the result of the rapid test, and should not incorporate the result of the rapid test. If the reference standard includes clinical diagnosis of COVID‐19 for RT‐PCR‐negative patients, then established criteria should be used. Studies including samples from participants known not to have COVID‐19 should use pre‐pandemic sources or if contemporaneous samples then at least two RT‐PCR‐negative tests were required to confirm the absence of infection. Data should be reported for all study participants, including those where the result of the rapid test was inconclusive, or participants in whom the final diagnosis of COVID‐19 was uncertain. Studies should report whether results relate to participants (one sample per participant), or samples (multiple samples per participant).
We analysed rapid antigen and molecular tests separately. Studies often referred to ‘samples’ rather than ‘patients’, especially for the rapid molecular tests, however for many studies we do not suspect that inclusion of multiple samples per study participant was a significant issue. For consistency of terminology throughout the review, we refer to results on a per‐sample basis. If studies evaluated multiple tests in the same samples, we included them multiple times. We present estimates of sensitivity and specificity per study for each test brand using paired forest plots, and summarise results using average sensitivity and specificity in tables as appropriate. As heterogeneity is apparent in many analyses, these point estimates must be interpreted as the average of a distribution of values.
We did not make any formal comparisons between antigen assay brands because of the large number of different assays and small study numbers for many of them. We did however carry out a formal comparison (based on between‐study comparisons) for studies using two brands of molecular tests (ID NOW (Abbott Laboratories) and Xpert Xpress (Cepheid Inc)).
We estimated summary sensitivities and specificities with 95% confidence intervals (CI) using the bivariate model (Reitsma 2005), via the meqrlogit command of Stata/SE 16.0. When few studies were available, we simplified models by first assuming no correlation between sensitivity and specificity estimates and secondly by setting near‐zero variance estimates of the random effects to zero (Takwoingi 2017). In cases where there was only one study per test, we reported individual sensitivities and specificities with 95% CI constructed using the binomial exact method.
Where studies presented only estimates of sensitivity or of specificity, we fitted univariate, random‐effects, logistic regression models. In a number of instances where there was 100% sensitivity or specificity for all evaluations, we computed estimates and 95% CIs by summing the counts of TP, FP, FN and TN across 2x2 tables. These analyses are clearly marked in the tables. We present all estimates with 95% confidence intervals.
We examined heterogeneity between studies by visually inspecting the forest plots of sensitivity and specificity. Where adequate data were available, we investigated heterogeneity related to symptom status, time post‐symptom onset, viral load, test brand, and test method by including indicator variables in the random‐effects logistic regression models. Absolute differences between the sensitivity or specificity and the P values were reported from the model. In instances where only one study was available per test or when tests were being directly compared following summing of counts of the 2x2 tables, we performed test comparison using the two‐sample test of proportions. Few studies reported specificity estimates by time after symptom onset, therefore for this variable and for analyses by viral load, we considered only effects on sensitivity.
We performed four sensitivity analyses.
We made no formal assessment of reporting bias but have indicated where we were aware that study results were available but unpublished.
We summarised key findings in a 'Summary of findings' table indicating the strength of evidence for each test and findings, and highlighted important gaps in the evidence.
We are aware of additional studies published since the electronic searches were conducted on 30 September 2020 and plan to update this review. We have already conducted the next search to 1 January 2021.
We screened 34,742 unique records (published or preprints) for inclusion in the complete suite of reviews to assist in the diagnosis of COVID‐19 (Deeks 2020b; McInnes 2020). Of 1749 records selected for further assessment for inclusion in any of the four molecular, antigen or antibody test reviews, we identified 199 full‐text reports requiring assessment for inclusion in this review; 90 for the first iteration of the review and 109 for this review update. See Figure 1 for the PRISMA flow diagram of search and eligibility results (McInnes 2018; Moher 2009).
1 Study flow diagram
We included 64 reports in this review, and we excluded 135 publications that did not meet our inclusion criteria. Exclusions were mainly based on index test (n = 85) or ineligible study designs (n = 26), for example, designs that did not allow estimation of test accuracy. The reasons for exclusion of all 135 publications are provided in Characteristics of excluded studies. Appendix 8 provides a list of studies evaluating eligible tests but excluded for other reasons (n = 5), and studies evaluating technologies not yet suitable for use at the point of care (n = 41).
Of the 64 study reports, 18 were available only as preprints, 38 were published papers and eight were publicly available reports either from independent reference laboratories (one from Public Health England and two identified via the SMF) or were independent evaluations co‐ordinated by FIND (n = 5).
We contacted the authors of 10 study reports for further information (Blairon 2020; Courtellemont 2020; Diao 2020; Gibani 2020; Gremmels 2020(a); Linares 2020; Nash 2020; Porte 2020a; Schildgen 2020 [A]; Weitzel 2020 [A]), and received replies and the requested information with one exception (Linares 2020). We also contacted the evaluation teams at FIND and Public Health England and received additional information about study methods from FIND and some additional data from Public Health England.
The 64 included study reports relate to 78 separate studies. Please note when naming studies, we use the letters [A], [B], [C] etc. in square brackets to indicate data on different tests evaluated in the same study and (a), (b), (c) to indicate data from different participant cohorts from the same study report. For example, the five included reports from FIND correspond to eight ‘studies’ because three reports separately provided data from more than one evaluation centre.
Of the 78 studies, 77 reported data for respiratory samples and one (Szymczak 2020), reported data for faecal samples. The main results, Tables and Figures focus on the respiratory samples, with Szymczak 2020 reported separately.
The 77 studies using respiratory samples included a total of 24,418 unique samples, with 7484 samples with RT‐PCR‐confirmed SARS‐CoV‐2 (some samples were analysed by more than one index test). Forty‐eight studies evaluated antigen tests (Albert 2020; Alemany 2020; Billaud 2020; Blairon 2020; Cerutti 2020; Courtellemont 2020; Diao 2020; Fenollar 2020(a); Fenollar 2020(b); FIND 2020a; FIND 2020b; FIND 2020c (BR); FIND 2020c (CH); FIND 2020d (BR); FIND 2020d (DE); FIND 2020e (BR); FIND 2020e (DE); Fourati 2020 [A]; Gremmels 2020(a); Gremmels 2020(b); Gupta 2020; Kruger 2020(a); Kruger 2020(b); Kruger 2020(c); Lambert‐Niclot 2020; Linares 2020; Liotti 2020; Mak 2020; Mertens 2020; Nagura‐Ikeda 2020; Nash 2020; PHE 2020(a); PHE 2020(b); PHE 2020(c) [non‐HCW tested]; PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]; PHE 2020(e); Porte 2020a; Porte 2020b [A]; Schildgen 2020 [A]; Scohy 2020; Shrestha 2020; Takeda 2020; Van der Moeren 2020(a); Van der Moeren 2020(b); Veyrenche 2020; Weitzel 2020 [A]; Young 2020) and 29 studies evaluated molecular tests (Assennato 2020; Broder 2020; Chen 2020a; Collier 2020; Cradic 2020(a); Cradic 2020(b); Dust 2020; Ghofrani 2020; Gibani 2020; Goldenberger 2020; Harrington 2020; Hogan 2020; Hou 2020; Jin 2020; Jokela 2020; Lephart 2020 [A]; Lieberman 2020; Loeffelholz 2020; Mitchell 2020; Moore 2020; Moran 2020; Rhoads 2020; Smithgall 2020 [A]; SoRelle 2020; Stevens 2020; Thwe 2020; Wolters 2020; Wong 2020; Zhen 2020 [A]). Summary study characteristics are presented in Table 2 with further details of study design and index test details in Appendix 9 and Appendix 10 for antigen assays and Appendix 11 and Appendix 12 for molecular assays. Full details are provided in the Characteristics of included studies table.
The median sample size of the included studies is 182 (interquartile range (IQR) 104 to 400) and median number of SARS‐CoV‐2 confirmed samples included is 63 (IQR 38 to 119). Sample sizes for antigen test evaluations were larger than those for molecular test evaluations (median 291.5 (IQR 155 to 502.5) compared to 104 (IQR 75 to 172)). Half of the studies (39/77, 51%) were conducted in Europe, 20 in North America, seven in South America, seven in Asia, one study included samples from more than one country and in one, the country of sample origin was unclear.
Over half of the antigen test studies included samples from participants presenting in the community for COVID‐19 testing community test centres (22/48, 46%); emergency departments (3, 6%); or as part of contact tracing or outbreak investigations (4, 8%) (Table 2). Eleven antigen test studies (23%) selected samples from those submitted to laboratories for routine RT‐PCR testing with limited detail of the participants providing the samples (‘laboratory‐based’ studies), or included multiple (8%) or unclear (2%) settings. Over half of antigen test studies were conducted in symptomatic (16, 33%) or mainly symptomatic (11, 23%) populations, with only three (6%) exclusively in asymptomatic populations (two in asymptomatic contacts of confirmed cases (Fenollar 2020(b); Shrestha 2020), and one involved staff screening, all of whom were RT‐PCR‐negative (PHE 2020(e)). The remaining antigen studies included samples from populations with mixed symptom status (8, 17%) or provided no information regarding symptom status (10, 21%). Of the 10 that provided no information, seven were laboratory‐based studies providing no details of the settings from which the tested samples had been obtained, one included samples from a COVID‐19 test centre, one was an outbreak investigation and in one the study setting could not be derived. There were no studies evaluating strategies of multiple tests.
A total of 13 studies provided accuracy data for people with no symptoms at the time of testing (3 studies exclusively in asymptomatic populations, and 10 studies providing subgroup data for people with no reported symptoms); one study provided only specificity data. Of the 12 datasets reporting both sensitivity and specificity, one (Alemany 2020), purportedly described preventive screening of the general population (although the reported prevalence of 24% is very high for such a scenario), one (Cerutti 2020), described targeted traveller screening, four (Billaud 2020; Fenollar 2020(b); Gupta 2020; Shrestha 2020), tested contacts of confirmed cases (one as part of an outbreak investigation) and the remaining six datasets were subgroups of samples from people presenting for routine testing. We identified one additional asymptomatic dataset in a report of several substudies but we did not include it as participants underwent antigen testing up to five days after a positive PCR test and it was not possible to determine the time point at which symptom status was recorded; it was also not possible to determine which 'substudy’ the data related to (PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]).
Thirty‐one of the 48 studies evaluating antigen tests reported results for SARS‐CoV‐2‐confirmed samples above and below a Ct value from the reference standard RT‐PCR. The median proportion of participants with 'high' viral load was 52% (IQR 35% to 60%). The most commonly used threshold was 24 or 25 Ct or less (n = 29 studies (or 36/58 test evaluations); 11 studies (15/58 test evaluations) reported results with at a threshold of between 31 and 33 Ct or less ; and 13 studies (13 evaluations) reported other thresholds including less 28 Ct (n = 3), 30 Ct (n = 5), 31 Ct (n = 3), or 35 Ct (n = 2)
In contrast, studies evaluating molecular tests were mainly laboratory‐based (20, 69%), with three (10%) including samples from participants presenting to emergency department or urgent care settings, two in hospital inpatients (7%), and four (14%) including samples from participants presenting in multiple settings. Twelve of the 29 studies (41%) reported included only samples from symptomatic patients, four reported mixed symptom status (10%) and 14 (48%) provided no information regarding symptom status. Of the 14 that provided no information, one was based in a hospital Accident and Emergency department, and the remaining 13 were laboratory‐based studies, only three of which gave any details of the settings from which the tested samples had been obtained (three reported inclusion of samples from either inpatients and outpatients (n = 1), inpatients and ambulatory patients (n = 1) or inpatients and emergency department patients (n = 1) but did not provide the number of samples from each source). There were no studies evaluating strategies of multiple tests.
Five studies evaluating molecular assays, reported proportions with high viral load ranging from 33% to 80%, median 46%. All five studies reported results above and below a Ct value of 30.
Table 2 shows a similar distribution of study designs between those evaluating antigen and molecular tests. Overall, 60% of studies (n = 46) used a ‘single group’ design to estimate both sensitivity and specificity and 22% (n = 17) used a ‘two group’ design with separate selection of RT‐PCR‐positive and RT‐PCR‐negative samples. In four studies (5%), the design could not be fully determined but probably deliberate separate sampling of RT‐PCR‐positive and RT‐PCR‐negative samples had been used.
Nine studies included only samples with confirmed SARS‐CoV‐2, thus only allowing estimation of sensitivity (six antigen and three molecular assay studies), and one study included only SARS‐CoV‐2‐negative samples allowing estimation of specificity only. All studies defined the presence or absence of SARS‐CoV‐2 infection based on RT‐PCR. Of the 68 studies that included SARS‐CoV‐2‐negative samples, 63 (93%) required a single, negative PCR to confirm absence of infection and two (3%) required two negative PCR results. The remaining three studies used pre‐pandemic samples (n = 2) or contemporaneous samples with other respiratory infections.
Thirty‐three studies (43%), obtained paired swabs for index and reference standard, 39 (51%) used the same swab for point‐of‐care and RT‐PCR (18 antigen and 21 molecular studies) and five studies used a mix of paired and same swabs (n = 1) or it was not possible to determine this information from the study report.
Fifteen studies evaluated only one test, seven compared two or more tests in the same participants (four with two tests each, one with three tests and one each with four or five tests). In total the 77 studies that used respiratory samples reported on a total of 90 test evaluations. Appendix 13 provides details extracted from the manufacturer’s instructions for use documents for all included tests.
Forty‐eight studies reported 58 evaluations of antigen tests; 41 of CGIAs, nine FIA, two alternative type of LFA using alkaline phosphatase‐labelled antibodies, and six where assay type could not be determined. Studies evaluated 16 different commercially produced assays, as documented, with full assay identification details, in Appendix 13. One study reported the development of the Shenzhen Bioeasy assay (Diao 2020), but it is not clear whether the commercially available assay is identical to the one reported in the study or whether it has undergone further refinement. One study reported evaluating a Roche SARS‐CoV‐2 assay, which appears to be the SD Biosensor STANDARD Q (Schildgen 2020 [A]). Only 12 studies provided product codes for the tests evaluated (FIND 2020a; FIND 2020b; FIND 2020c (BR); FIND 2020c (CH); FIND 2020d (BR); FIND 2020d (DE); FIND 2020e (BR); FIND 2020e (DE); Gremmels 2020(a); Gremmels 2020(b); Porte 2020a; Weitzel 2020 [A]). The study reports or manufacturer IFUs for 11 assays reported targeting the nucleocapsid protein; this information was not reported for the Beijing Savant, Bionote, Biosynex, Liming Bio‐Products, or RapiGEN Inc assays (Appendix 13). We were unable to identify any information for Beijing Savant, E25Bio or Liming Bio‐Products assays online.
Multiple combinations of sample types and use of direct swab testing or swabs in viral transport medium or saline were reported across the studies (Table 2). Forty‐one of 58 evaluations used nasopharyngeal (n = 30), oropharyngeal (n = 1) or nasal (n = 2) samples (type of nasal sample was not reported), or combinations of nasopharyngeal, nasal or oropharyngeal samples (n = 8; nasopharyngeal or nasal mid‐turbinate in one, nasopharyngeal or combined naso‐ and oropharyngeal in two, naso‐ or oropharyngeal in two, and naso‐ or oropharyngeal or combined naso‐ and oropharyngeal samples in three. Thirteen evaluations used combined naso‐ and oropharyngeal samples for all participants, one used saliva samples and three evaluations (from one study) used bronchoalveolar lavage or throat wash samples. Of the six studies using nasal samples either alone (n = 2) or for at least some participants (n = 4), one reported that these were nares swabs, and the remaining five did not specify the type of nasal sample. Almost half of studies used direct swab testing (n = 28, 48%), 22 (38%) tested samples in viral transport medium, saline or other medium, and in 8 (14%) this information was not provided.
IFUs for five assays explicitly recommend against using any transport medium for swab testing (assays from Becton Dickinson, Bionote, Quidel and SD Biosensor; Appendix 13), one (Coris BioConcept) states that viral transport medium may be used, and the other nine do not mention use of transport medium, although two of the nine (from AAZ and Biosynex) imply that viral transport medium should not be used (using statements such as "use within one hour, stored in clean unused plastic tube"). We considered 29 of 58 antigen evaluations (50%) to be compliant with manufacturer IFUs in terms of sample type, use of viral transport medium and time interval between collection and testing. Sixteen evaluations were not compliant with IFUs; nine used viral transport medium, four used freezing, four tested samples not listed on the IFUs, and in two testing was not always conducted within the one‐hour time period specified in the IFU. For the remaining 13 evaluations either no IFU was available (n = 4), viral transport medium or saline was used but the IFU did not specifically address whether viral transport medium was recommended or not (n = 7), or insufficient detail was provided in the study.
Samples were collected by healthcare workers in 15 (26%) evaluations, by trained non‐healthcare workers, such as firefighters or Ministry of Health employees in three (5%) evaluations, self‐collected in six (10%) and the collection was not described in 34 evaluations (59%). Sample testing was conducted ‘on‐site’ immediately or within one hour of collection in 21 (36%) evaluations by the same healthcare workers (n = 13), trained non‐healthcare workers (n = 3) who collected the samples, or this information was not provided (n = 5). In the remaining 27 evaluations (47%), testing was conducted by laboratory staff (n = 12) or was inferred to be by laboratory staff (n = 15). For the latter group, the time interval between sample collection and testing was on receipt at the laboratory, some reporting delays of up to six hours.
Twenty‐nine studies reported 32 evaluations of five different commercially available rapid molecular 13 evaluating ID NOW (Abbott Laboratories), 15 evaluating Xpert Xpress (Cepheid Inc), two of SAMBA II (Diagnostics for the Real World), and one evaluation each of Accula (Mesa Biotech Inc.) and COVID Nudge (DNANudge). None of the studies reported product codes for the tests evaluated. One study of Xpert Xpress used the 'research use only' (RUO) version of the test but reported that the RUO version contains the same reagents as the 'emergency use authorisation' (EUA) version. The RUO test allows the user to view the amplification curves for the RdRp gene as well as for the E‐gene and N2 targets whereas the EUA version restricts the amplification curves to E and N2 only. ID NOW and SAMBA‐II use isothermal techniques, Xpert Xpress and COVID Nudge are based on RT‐PCR, and Accula is described as a PCR plus lateral flow assay.
Multiple combinations of sample types and use of direct swab testing or swabs in viral transport medium or saline were reported across the studies (Table 2). The sample types used included combined naso‐ and oropharyngeal samples (n = 2), nasopharyngeal samples alone (n = 16), nasal alone (n = 2), oropharyngeal samples alone (n = 1), or a combination of two or more of either nasopharyngeal or nasal or oropharyngeal samples (n = 8). One evaluation used throat saliva or lower respiratory tract specimens, one used saliva samples alone and one did not specify the sample type used. Of the six studies using nasal samples either alone (n = 2) or for at least some participants (n = 4), one reported using nares swabs, and the remaining five did not specify the type of nasal sample used.
Eight evaluations (25%) reported direct swab testing in some (n = 1) or all (n = 7) samples, 18 (59%) used swabs in viral transport medium only (n = 12) or in viral transport medium or some other transport medium (n = 6), and six did not report whether they used any transport medium.
Sample collection was described in only three evaluations (9%) (Gibani 2020; Harrington 2020; Rhoads 2020; Table 2); the remaining studies did not describe sample collection but it is likely that samples were collected as part of routine care by healthcare workers. Sample testing was clearly described as conducted on‐site by medical personnel or by laboratory personnel at local laboratories in one of the studies reporting sample collection (Harrington 2020), while a second implied testing as soon as possible after collection, possibly by the same healthcare worker (Gibani 2020). Four (12.5%) evaluations stated that laboratory staff carried out the tests. In 16 of the remaining 26 studies, testing by laboratory staff was inferred, based on delays between collection and testing of 18 hours to seven days (n = 10), or reported use of archived or frozen samples (n = 6). The remaining eight evaluations provided no useful information regarding who carried out the test (Assennato 2020; Dust 2020; Ghofrani 2020; Jin 2020; Jokela 2020; Moran 2020; Rhoads 2020; SoRelle 2020).
Two of the five manufacturers document IFU for samples stored in transport medium (Xpert Xpress and SAMBA II assays); two explicitly recommend against the use of viral transport medium (ID NOW and Accula), although at the time of the test evaluations some viral transport media were documented as acceptable for ID NOW; and one IFU does not mention the use of viral transport medium (COVID Nudge). Although immediate sample testing is preferred, all manufacturers document an acceptable period of refrigerated storage of between eight hours (COVID Nudge), and seven days with refrigeration (Xpert Xpress). See Appendix 13.
We considered only nine of 32 (28%) evaluations to be compliant with manufacturer IFUs in regard to sample type, use of viral transport medium and time interval between collection and testing. Sixteen evaluations were not compliant with IFUs; eight used viral transport medium, six used frozen samples, and two tested samples not listed on the IFUs. For the remaining seven evaluations, either the testing interval from sample collection was unclear (n = 5) or saline was used but the IFU did not specifically address whether this was recommended or not (n = 2).
We report the overall methodological quality assessed using the QUADAS‐2 tool for all included studies (n = 78) in Figure 2 (Whiting 2011). See Appendix 14 for separate summary plots by test method and for a plot of study‐level ratings by quality domain. We explain how we reached these judgements in the Characteristics of included studies table.
2 Risk of bias and applicability concerns review authors' judgements about each domain presented as percentages across included studies. Numbers in the bars indicate the number of studies
We considered whether the findings of individual studies were at risk of bias, and whether there were concerns that results might not apply to standard use of the tests. We did not judge any study at low risk of bias, although in 11 of 78 studies the only concern was that a single negative RT‐PCR was used to confirm absence of COVID infection rather than the preferred two negative tests. All studies raised concerns regarding the applicability of their results, but in 13 of 78 studies the only concern was the reliance on only PCR to identify SARS‐CoV‐2 cases (and nine of these 13 are in common with the 11 using a single negative RT‐PCR).
We judged 22 studies (28%) to be at low risk of bias, and 29 (37%) at high risk of bias because of deliberate sampling of participants based on the reference standard result (n = 25; 16 two‐group studies and nine that only included samples with confirmed SARS‐CoV‐2 infection or absence of infection) or use of convenience sampling (n = 4). In 27 studies (35%) the risk of bias was unclear because of poor reporting of recruitment procedures or inclusion criteria (Figure 2).
A third (27/78) of studies were likely to have selected an appropriate patient group, recruiting participants from COVID‐19 test centres, urgent care or emergency departments or identified through contact tracing. We had high concerns about the applicability of the selected participants in almost half of studies (35/78). Recruited participants were unlikely to be similar to those in whom the test would be used in clinical practice because of deliberate sampling (n = 25) or sample inclusion based on the availability of residual and sometimes frozen samples, or both (n = 22).
Poor reporting meant we could not clearly assess whether there was a risk of bias through performance of the index test in 41 (53%) studies. In general, antigen test studies were of a higher methodological standard for the index test domain compared to studies of molecular tests (Figure 2).
For antigen tests, we observed low risk of bias in 60% of studies (29/48). Risk of bias was unclear in the remaining studies because we could not judge whether interpretation of the index test was undertaken with knowledge of the reference standard result. For molecular tests, risk of bias was low in only 17% of studies (5/30). We observed high risk of bias in three studies (Moran 2020; Smithgall 2020 [A]; Wolters 2020) because they did not follow the manufacturer’s prespecified threshold for the Xpert Xpress test (re‐testing of samples with presumptive positive results). Risk of bias was unclear in 73% (22/30) of studies because they did not report blinding to the reference standard (n = 22), six of these studies also did not report how they handled presumptive positive results on Xpert Xpress.
Fourteen studies (18%), including 13 antigen and one molecular test study, conducted testing as would be expected in practice (low concern regarding applicability). We had high concerns about applicability in half of all studies (39/78); 48% (23/48) of antigen and 57% (16/30) of molecular studies. Twenty‐seven (11 antigen and 16 molecular) did not comply with manufacturers’ IFU and a further 10 (all antigen studies), did not carry out tests as would occur in practice (i.e. trained, centralised laboratory staff carried out testing). In another two antigen studies concerns for applicability were high because tests were not available for purchase (Diao 2020; Nash 2020). Of the remaining 25 studies (12 antigen and 13 molecular) 16 conducted the test within the manufacturer IFU but none clearly described the setting for testing or personnel conducting the test.
Six studies were at low risk of bias for the reference standard. Although 12 used an appropriate reference standard, half (6/12) did not clearly implement blinding of the reference standard to the index test. High risk of bias (66/78) was present because studies did not use an adequate reference standard (Figure 2); they used either a single negative RT‐PCR to define absence of SARS‐CoV‐2 infection (n = 64) or the index test formed part of a composite reference standard (n = 2).
A total of 36 studies reported blinded RT‐PCR interpretation, two (with composite reference standard) did not implement blinding, and 40 (51%) provided insufficient information about blinding of the reference standard to the index test to judge risk of bias.
We judged 76 of the 78 studies to raise concerns about applicability (97%) because of defining the presence of SARS‐CoV‐2 infection based on a single RT‐PCR‐positive result. These studies will have excluded individuals who are RT‐PCR‐negative but have exposure and clinical features that meet the case definitions for COVID‐19.
Only 13 (17%) studies (all of antigen tests) were at low risk of bias for participant flow and timing (Figure 2). Twenty‐nine (37%) were at high risk of bias (19 antigen and 10 molecular) because of exclusion of samples following invalid index test results (n = 23); delays between ‘paired’ swabs of up to three days (n = 4), different reference standards used (n = 3), or because they provided results on a per sample instead of per patient basis (n = 2). These categories are not mutually exclusive.
We judged risk of bias unclear for 36 (46%) studies, primarily because of lack of clarity about participant inclusion and exclusion from analyses (n = 34), with no missing data or indeterminate test results reported and no Standards for Reporting Diagnostic Accuracy Studies (STARD)‐style participant flow diagram and checklist (Bossuyt 2015), to fully report outcomes for all samples.
In 27 studies all authors declared no conflicts of interest, although one study that reported the validation of a new test included a co‐author affiliated to the test manufacturing company. Of these 27 studies, 19 were independent evaluations published by FIND or were from national reference laboratories. Twenty studies did not provide a conflict of interest statement, including 13 published studies and one study that reported affiliations to the test manufacturer. In the 12 remaining studies at least one author declared potential conflicts of interest in relation to the test.
Twenty‐six studies provided no funding statement, 12 reported no funding sources to declare, and the remainder (n = 40) reported one or more funding sources.
Of the 78 included studies, eight reported evaluations of more than one test using the same samples and one reported evaluations of three tests using different samples (Table 2). To include all results from all tests in these analyses we have treated results from different tests of the same samples within a study as separate data points, such that data are available on 91 test evaluations (58 evaluations of antigen tests in 48 studies and 33 evaluations of rapid molecular tests in 30 studies).
As previously stated, 77 of the 78 studies reported data for respiratory samples and one (Szymczak 2020), reported data for non‐respiratory (faecal) samples. The main results, Tables and Figures focus on the respiratory samples, with Szymczak 2020 reported separately.
The results tables identify where estimates are based on multiple assessments of the same samples by including both the number of test evaluations and the number of studies. Nine datasets are from ‘cases only’ studies reporting only sensitivity estimates (six for antigen tests and three for molecular assays), and one antigen test evaluation is for ‘non‐COVID‐19’ cases reporting only specificity. Summary results are presented for studies providing both sensitivity and specificity data and then adding in the data from sensitivity‐ or specificity‐only evaluations. The numbers of true positives, false positives, and total samples with and without confirmed SARS‐CoV‐2 infection are based on test result counts.
We present results for antigen tests overall and by subgroup in Table 3. Table 4 and Table 5 present results by test brand overall and by symptom status, and give results of sensitivity analyses restricting by compliance with manufacturer IFU. Forest plots of study data for the primary analysis are in Figure 3 and for subgroup analyses by symptom status and time after symptom onset are in Figure 4 and Figure 5. Appendix 15 provides forest plots for study data according to Ct value and study design. Individual plots by test brand are provided in Figure 6 for test brands with three or more evaluations and Figure 7 for test brands with one or two evaluations. Figure 8 shows data from studies comparing the accuracy of two or more antigen assays. Full identification details for studies of antigen‐based assays are provided in Appendix 9 and Appendix 10.
3 Forest plot of studies evaluating antigen tests. BR: Brazil; CH: Switzerland; DE: Germany; HCW: healthcare worker; Lab: laboratory
4 Forest plot of data for antigen tests according to symptom status. A&E: accident and emergency; BR: Brazil; CH: Switzerland; DE: Germany; HCW: healthcare worker; Lab: laboratory
5 Forest plot of antigen test evaluations by week post symptom onset (pso). A&E: accident and emergency; Ag: antigen; BR: Brazil; CH: Switzerland; DE: Germany
6 Forest plot by test brand for assays with ≥ 3 evaluations. BR: Brazil; CGIA: colloidal‐gold immunoassay; CH: Switzerland; DE: Germany; FIA: fluorescent immunoassay; HCW: healthcare worker; IFU: instructions for use; Lab: laboratory; LFA: lateral flow assay
7 Forest plot by test brand for assays with < 3 evaluations; CGIA: colloidal‐gold immunoassay; FIA: fluorescent immunoassay; IFU: instructions for use; LFA: lateral flow assay
8 Forest plot of studies reporting comparative data. CGIA: colloidal‐gold immunoassay; FIA: fluorescent immunoassay; LFA: lateral flow assay; nos: not otherwise specified
Results for molecular tests overall and by subgroup are reported in Table 6. Forest plots of study data for the primary analysis is in Figure 9 and for subgroup analyses by Ct value, study design and sensitivity analyses by pre‐ and post‐discrepant analysis in Appendix 16. Individual plots by test brand are provided in Figure 10. Full identification details for studies of molecular‐based assays are provided in Appendix 11 and Appendix 12. Appendix 17 provides forest plots for study data according to Ct value and discrepant analysis.
9 Forest plot of studies evaluating rapid molecular tests. A&E: accident and emergency
10 Forest plot by test brand for molecular assays. A&E: accident and emergency; IFU: instructions for use
Results showed high levels of heterogeneity in sensitivity. Average sensitivity was 68.9% (95% CI 61.8% to 75.1%) and average specificity was 99.6% (95% CI 99.0% to 99.8%) across the 51 evaluations of antigen tests reporting both sensitivity and specificity (based on 21,614 samples, including 6136 samples with confirmed SARS‐CoV‐2; Table 3; Figure 3). Adding the six ‘sensitivity only’ datasets and single ‘specificity only’ datasets had a negligible impact on results (Table 3). In the sections below we show that there are substantial differences between subgroups of studies according to symptom status, timing, test method and brand, therefore this average value is unlikely to accurately predict the performance of the test in a given setting and should not be used for this purpose.
Subgroup analysis by symptom status suggests that average test sensitivity to detect infection is 13.8 percentage points lower in asymptomatic (58.1%, 95% CI 40.2% to 74.1%; based on 12 evaluations, 1581 samples and 295 cases) compared to symptomatic (72.0%, 95% CI 63.7% to 79.0%; based on 37 evaluations, 15,530 samples and 4410 cases) participants (95% CI for the difference in 33.1 percentage points lower to 5.4 percentage points higher; Table 3; Figure 4). Restricting the comparison by symptom status to the nine evaluations reporting data for both symptomatic and asymptomatic subgroups (thus ensuring the comparison is made between the same tests used in the same way) showed a similar difference in sensitivity (14.4 percentage points lower in asymptomatic participants, 95% CI 38.8 lower to 10.0 percentage points higher; Table 3). Average results for the 19 evaluations in participants with mixed symptom status (n = 10) or symptom status not reported (n = 9) were between those observed for the symptomatic and asymptomatic sensitivity 63.0% (95% CI 52.2% to 72.6%) and specificity 98.4% (95% CI 98.0% to 98.8%) (6220 samples; 2392 cases).
We did not observe any important differences in specificity according to symptom status (Table 3).
We pooled data by time from symptom onset separately for sensitivity and specificity because the majority of evaluations did not report these data for people without SARS‐CoV‐2 (Table 3; Figure 5). Sensitivity was 78.3% (95% CI 71.1% to 84.1%) (26 evaluations; 5769 samples, 2320 cases) in the first seven days after symptom onset compared to 51.0% (40.8% to 61.0%) (22 evaluations; 935 samples, 692 cases) in the second week of symptoms (a decrease of 27.3 percentage points, 95% CI −32.8 to −21.9 percentage points decrease). This difference remained on restriction to the 22 evaluations reporting data for people in both week one and week two of symptoms (removing other between‐study differences; Table 3).
We did not observe any differences in specificity according to time after symptom onset (Table 3).
A total of 36 evaluations reported sensitivity according to Ct value using a threshold of 24 (n = 18) or 25 (n = 18) Ct or less to define higher viral load (Table 3; Appendix 15). Summary sensitivity in those with higher viral load was 94.5% (95% CI 91.0% to 96.7%) (based on 2613 cases), compared to 40.7% in those with lower viral load (95% CI 31.8% to 50.3%) (based on 2632 cases) (i.e. sensitivity was 53.8 percentage points lower for those with lower viral load; 95% CI 63.6 to 44.1 percentage points lower)). Applying a Ct threshold of ≤ 33 (n = 13) or < 32 (n = 2) led to a bigger difference in sensitivity although the number of samples in the lower viral load subgroup was considerably sensitivity associated with higher viral load was 82.5% (95% CI 74.0% to 88.6%) (based on 2127 samples) and for lower viral load was 8.9% (3.3% to 21.7%) (based on 346 samples), a difference of 73.5 percentage points (95% CI 84.7 to 62.4 percentage points lower).
We did not observe any clear differences in average sensitivity or specificity when studies were grouped by study design (15,336 samples and 3536 cases in 29 single group studies and 5729 samples and 2396 cases in 20 two‐group studies; Table 3; Appendix 15). Average sensitivity was lower in two‐group studies (64.1%, 95% CI 48.5% to 77.2%) compared to single‐group studies (72.1%, 95% CI 64.8% to 78.3%), however confidence intervals overlapped and the difference was within that which may be expected by chance (8.0 percentage points lower, 95% CI from 24.2 percentage points lower to 8.2 higher). Average specificities were 2.3 percentage points lower in the two‐group studies (95% CI from 2.9 to 1.6 percentage points lower), at 97.3% (95% CI 96.7% to 97.8%) compared to 99.6% (95% CI 99.1% to 99.8%) in single‐group studies.
We observed differences in accuracy according to test method (Table 3). The majority of evaluations (n = 36; 17,448 samples, 5085 cases) reported using a CGIA, average sensitivity was lower (64.0%, 95% CI 55.7% to 71.6%) than for FIAs (79.6%, 95% CI 67.5% to 88.0%; n = 9; 2820 samples, 712 cases; absolute difference of 15.6 percentage points, 95% CI 2.6 to 28.5 percentage points). We also observed marginal differences in specificity, with estimates of 99.0% (95% CI 98.8% to 99.2%) for CGIA and 97.7% (95% CI 95.3% to 98.8%) for FIA, a difference of 1.3 percentage points (95% from 3.0 percentage points lower to 0.3 higher). Results for lateral flow assays where the method could not be determined (n = 5) and for the single evaluation of an alkaline phosphatase (ALP)‐labelled assay were heterogeneous but largely in the realms of those observed for the other assay types (Table 3).
Results by test brand overall and sensitivity analyses by IFU compliance (based on sample type, use of viral transport medium, and time period between sample collection and test procedure) are reported in Table 4. Results by test brand for symptomatic and asymptomatic subgroups overall and by IFU compliance are in Table 5. Given the mixed settings in which asymptomatic individuals were tested (Results of the search), the data for asymptomatic subgroups cannot be considered applicable to any particular scenario for asymptomatic testing. Only three studies reported direct comparisons of tests, two using nasopharyngeal or oropharyngeal samples (Fourati 2020 [A]; Weitzel 2020 [A]).
We observed considerable heterogeneity in sensitivities for all assays.
Two evaluations of the COVID‐VIRO assay included 880 samples and 396 SARS‐CoV2‐positive samples (Figure 7). We did not pool the studies due to the heterogeneity in both sensitivity and specificity, although both were conducted in symptomatic or mainly symptomatic participants using nasopharyngeal samples.
In one study that compared antigen assays using nasopharyngeal samples in viral transport medium, sensitivity was 61.7% (95% CI 55.9% to 67.3%) and specificity (in pre‐pandemic samples) 100% (95% CI 98.9% to 100%; 632 samples, 295 cases;'Fourati 2020 [E]).
The second study used direct swab testing in compliance with the manufacturer’s IFU. Twenty participants in the study who previously tested positive on PCR retested negative with PCR at the time of the antigen test. All twenty samples showed weak lines on antigen testing. We considered these as false positives in the review (based on the negative result of the concurrent PCR test) whereas the study authors considered them to be true positives. With our re‐calculation, the test demonstrated sensitivity of 96.0% (95% CI 90.2% to 98.9%) and specificity of 86.4% (95% CI 79.8% to 91.5%; Courtellemont 2020). Sensitivity in this study may have been inflated by the inclusion of hospitalised, confirmed SARS‐CoV‐2‐positive participants.
We identified 11 evaluations of the Panbio assay, including 5691 unique samples, with 2031 SARS‐CoV‐2‐positive cases (Figure 6). One of the 11 evaluations included only SARS‐CoV‐2‐positive cases (n = 182 samples). Studies were conducted in community COVID‐19 test centres or emergency departments (n = 6), in contacts of confirmed cases (n = 2), and laboratory‐based evaluations (n = 2). The setting was not clear in one study. Participants were reportedly symptomatic (n = 5), asymptomatic (n = 1), with mixed symptom status (n = 4), or symptom status was not reported (n = 1). Nine evaluations used nasopharyngeal samples (Albert 2020; Billaud 2020; Fenollar 2020(b); FIND 2020b; Fourati 2020 [C]; Gremmels 2020(a); Gremmels 2020(b); Linares 2020), one (Alemany 2020), tested nasopharyngeal or nasal samples and one (Schildgen 2020 [A]), used bronchoalveolar lavage or throat wash samples. Only three of the 11 evaluations reported product codes for the assays used, one of which was for the assay for use with nasopharyngeal swabs (41FK10) and two (from the same study report) were for the assay for use with nasal swabs (41FK11), although the study reports using nasopharyngeal samples (Gremmels 2020(a); Gremmels 2020(b)).
Five of the 11 evaluations complied with manufacturer IFU for the test. Reasons for non‐compliance included use of viral transport medium, frozen storage, type of swab tested, or lack of clear reporting of test procedures used.
The average sensitivity and specificity of the Panbio assay
Restricting to IFU‐compliant evaluations, average sensitivities and specificities
The addition of one evaluation that reported sensitivity only in symptomatic participants led to only marginal differences in average sensitivity (Fenollar 2020(a); Table 5).
We identified three evaluations of the BD Veritor assay, including 727 unique samples, with 180 SARS‐CoV‐2‐positive cases (Figure 6). One of the three evaluations included only SARS‐CoV‐2‐positive cases (n = 125 samples). Studies were conducted in community COVID‐19 test centres (n = 2), or in multiple settings (n = 1). All participants were symptomatic. Two evaluations used combined naso‐ and oropharyngeal samples and one tested nasal samples.
None of the evaluations complied with manufacturer IFU for the test because the interval between sample collection and testing was greater than the maximum of one hour.
Average sensitivity and specificity of the BD Veritor assay
Adding the ‘cases only’ evaluation reduced average sensitivity to 79.4% (95% CI 72.9% to 84.7%) (n = 3; 180 cases; Van der Moeren 2020(b)).
The BD Veritor assay requires interpretation using a Veritor analyzer device, but Van der Moeren 2020(a) found that visual inspection of the test device resulted in the same sensitivity as with the Analyzer device, and similar specificity (100% compared to 99% using the Analyzer device).
We identified a single IFU‐compliant evaluation of the NowCheck assay in symptomatic participants (FIND 2020a;Figure 7). The study included 400 samples with 102 SARS‐CoV‐2‐positive cases, from participants presenting at a community‐based COVID‐19 test centre.
The sensitivity and specificity in this study were 89.2% (95% CI 81.5% to 94.5%) and 97.3% (95% CI 94.8% to 98.8%; Table 4; Table 5).
We identified a single evaluation of the Biosynex assay in symptomatic participants (Fourati 2020 [D]), including 634 samples with 297 with confirmed SARS‐CoV‐2 (Figure 7). The evaluation was not in compliance with the manufacturer’s IFU because samples were stored in viral transport medium and frozen prior to testing. The setting in which participants presented for testing was not reported.
Observed sensitivity was 59.6% (95% CI 53.8% to 65.2%) and specificity 100% (95% CI 98.9% to 100%; Table 4; Table 5).
The seven evaluations of the Coris Bioconcept assay included 1781 samples, with 707 SARS‐CoV‐2‐positive cases (Blairon 2020; Fourati 2020 [A]; Kruger 2020(b); Lambert‐Niclot 2020; Mertens 2020; Scohy 2020; Veyrenche 2020; Figure 6). Five of the seven were laboratory‐based evaluations with limited detail regarding study participants. One study recruited from community‐based COVID‐19 test centres and one included samples from hospital inpatients. Three studies included only or mainly symptomatic participants, one was in a mixed group and three did not report symptom status.
All evaluations tested naso‐ or oropharyngeal swabs and were compliant with the manufacturer IFU, however, it may be worth noting that the IFU for this assay permits the use of viral transport medium and freezing of samples, although immediate testing is recommended.
The average sensitivity and specificity of the COVID‐19 Ag Respi‐Strip
We identified a single evaluation of the E25Bio DART assay that included 190 samples, 100 with SARS‐CoV‐2 (Nash 2020; Figure 7). The symptom status of included participants was not reported and the manufacturer IFU is not yet available as the assay has been submitted for Emergency Use Authorisation (EUA) approval with the US Food and Drug Administration (FDA).
Sensitivity was 80.0% (95% CI 70.8% to 87.3%) and specificity 91.1% (95% CI 83.2% to 96.1%; Table 4).
We included two eligible evaluations were included, with a total of 265 samples, 165 were SARS‐COV‐2‐positive (Nagura‐Ikeda 2020; Takeda 2020; Figure 7). One study reported only sensitivity data (Nagura‐Ikeda 2020).
Takeda 2020 reported sensitivity of 80.6% (95% CI 68.6% to 89.6%) and specificity of 100% (95% CI 96.4% to 100%) in nasopharyngeal samples (162 samples, 62 cases; Table 4). They did not report symptom status of participants and provided insufficient detail to allow us to judge IFU compliance.
Nagura‐Ikeda 2020 evaluated the assay using saliva samples in symptomatic participants (not within IFU specifications), the ESPLINE assay correctly identified 12 of 103 PCR‐positive samples (sensitivity 11.6%, 95% CI 6.2% to 19.5%; Table 4; Table 5).
We included one report that evaluated the Innova study as six separate substudies; three reported both sensitivity and specificity (PHE 2020(a); PHE 2020(b); PHE 2020(c) [non‐HCW tested]), two reported sensitivity alone (PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]), and one reported specificity alone (PHE 2020(e); Figure 6). The studies reported a total of 3904 participants, including 1017 SARS‐CoV‐2‐positive cases. Detail regarding symptom status, was limited, however the study populations were coded symptomatic (samples from hospital inpatients in PHE 2020(a)), mainly symptomatic for samples from COVID‐19 testing centres (PHE 2020(c) [non‐HCW tested]; PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]), although data on symptom status were reported for only two of these studies (PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]), not reported for the outbreak investigation in PHE 2020(b) and asymptomatic staff screening for PHE 2020(e). The study authors for the outbreak evaluation study did not report the sensitivity value of 28.3% (95% CI 16.0% to 43.5%) in the publications but provided it to us on request.
All evaluations used naso‐ or oropharyngeal samples, two in viral transport medium (PHE 2020(a); PHE 2020(b)), and four using direct swab testing in compliance with manufacturer IFU (PHE 2020(c) [non‐HCW tested]; PHE 2020(d) [HCW tested]; PHE 2020(d) [Lab tested]; PHE 2020(e)).
For studies reporting both sensitivity and specificity, average sensitivity and specificity
Only one of the three studies that reported both sensitivity and specificity was compliant with manufacturer IFU, the sensitivity and specificity
Summary results from the four IFU‐compliant evaluations were calculated as
Adding data from single‐group evaluations in either RT‐PCR‐positive or RT‐PCR‐negative
Results for each of the three IFU‐compliant evaluations by test operator were (Figure 6):
We identified a single evaluation of the StrongStep assay in 19 symptomatic participants with nine SARS‐CoV‐2 positive samples ((Weitzel 2020 [B]; Figure 7). We could not identify the manufacturer’s IFU for this assay. The study authors terminated the evaluation early following poor early results for this assay.
Sensitivity was 0% (95% CI 0% to 33.6%) and specificity 90.0% (95% CI 55.5% to 99.7%; 19 samples, 9 cases; Table 4; Table 5).
We identified a single evaluation of the SOFIA assay in symptomatic participants, including 64 samples with 32 SARS‐CoV‐2‐positive cases (Porte 2020b [A]; Figure 7). The study used combined naso‐ and oropharyngeal swab samples in viral transport medium, therefore the evaluation was not compliant with the manufacturer IFU.
Sensitivity was 93.8% (95% CI 79.2% to 99.2%) and specificity was 96.9% (95% CI 83.8% to 99.9%; Table 4; Table 5).
We identified six evaluations of the RapiGen BIOCREDIT assay; these reported data for 2170 samples, with 470 confirmed SARS‐COV‐2‐positive cases (FIND 2020e (BR); FIND 2020e (DE); Mak 2020;Schildgen 2020 [A]; Shrestha 2020; Weitzel 2020 [A]; Figure 6). One laboratory‐based study included cases only (n = 160). The other evaluations included participants from community‐based COVID‐19 test centres (n = 2), emergency departments (n = 1), contact tracing (n = 1) or did not clearly report the setting (n = 1). Two studies included only symptomatic participants, two reported including both symptomatic and asymptomatic participants (mixed group) and one did not report symptom status. All evaluations apart from one (Schildgen 2020 [A]), tested nasopharyngeal or combined naso‐ or oropharyngeal samples.
Only three of the six evaluations complied with manufacturer IFU, with non‐compliance because of the use of viral transport medium, or the type of swab tested.
The average sensitivity and specificity of the BIOCREDIT assay
Restricting to IFU‐compliant evaluations, average sensitivities and specificities
The addition of one evaluation that reported sensitivity only led to a decrease in overall average sensitivity of 5.6 percentage points (Mak 2020; Table 5).
According to the manufacturer IFU, the Roche SARS‐CoV‐2 assay is available under a partnership with SD Biosensor.
There was a single evaluation of the Roche assay using 73 bronchoalveolar lavage or throat wash samples (not covered by the IFU) in participants with mixed symptom status (Figure 7); 42 of the 73 samples were RT‐PCR‐positive (Schildgen 2020 [A]).
Overall, using bronchoalveolar lavage or throat wash samples, the sensitivity and specificity were 88.1% (95% CI 74.4% to 96.0%) and 19.4% (95% CI 7.5% to 37.5%) (73 samples, 42 cases; Table 4). Only the results for the subgroup of 50 throat wash samples could be separated by symptom
We identified a single evaluation of the Huaketai assay in 109 symptomatic participants, using combined naso‐ or oropharyngeal swabs in viral transport medium (Weitzel 2020 [C]; Figure 7). We could not obtain the manufacturer IFU.
Sensitivity was 16.7% (95% CI 9.2% to 26.8%) and specificity was 100% (95% CI 88.8% to 100%; 109 samples, 78 cases; Table 4; Table 5).
We identified four evaluations of the STANDARD F assay; these reported data for 1552 samples, with 295 confirmed SARS‐COV‐2‐positive cases (FIND 2020d (BR); FIND 2020d (DE); Liotti 2020; Porte 2020b [B]; Figure 6). Three evaluations included all or mainly symptomatic participants from community‐based COVID‐19 test centres and one was a laboratory‐based study that did not provide details regarding symptom status.
All evaluations tested nasopharyngeal or combined naso‐ or oropharyngeal samples, however only two complied with manufacturer IFU. Reasons for non‐compliance were the use of viral transport medium, or lack of information concerning viral transport medium.
The average sensitivity and specificity of the STANDARD F COVID‐19 Ag assay
No data for asymptomatic people were available.
Restricting to IFU‐compliant evaluations, average sensitivity and specificity
We identified six evaluations of the STANDARD Q assay; these reported data for 3480 samples, with 821 confirmed SARS‐CoV‐2‐positive cases (Figure 6). Four evaluations included participants from community‐based COVID‐19 test centres, one was a laboratory‐based study, and one included multiple settings. Four evaluations included symptomatic or mainly symptomatic participants, and two included mixed symptomatic and asymptomatic participants.
All evaluations tested nasopharyngeal or combined naso‐ or oropharyngeal samples, four of which were compliant with manufacturer’s IFUs, the other two used samples in viral transport medium.
The average sensitivity and specificity of the STANDARD Q COVID‐19 Ag assay
Restricting to IFU‐compliant evaluations, average sensitivities and specificities
We included three evaluations of the Bioeasy FIA; these included 965 samples with 177 SARS‐CoV‐2‐positive cases ((Kruger 2020(a); Porte 2020a; Weitzel 2020 [D]; Figure 6). Studies were conducted in hospital emergency departments (n = 2) or a community COVID‐19 test centre (n = 1). Participants in studies were all symptomatic or mainly symptomatic.
Two evaluations used combined naso‐ or oropharyngeal swabs and one tested either nasopharyngeal or oropharyngeal swabs. Two evaluations used swabs in viral transport medium, which was not documented as suitable for use on the manufacturer IFU.
The average sensitivity and specificity of the Shenzhen Bioeasy assay were :
The single IFU‐compliant evaluation Kruger 2020(a) reported sensitivity of 66.7% (95% CI 38.4% to 88.2%) and specificity of 93.1% (95% CI 91.0% to 94.9%; 727 samples, 15 cases).
We also included an additional study that reported the development of this assay but we did not pool data with the other evaluations as it was a development and not a validation study (Diao 2020; Figure 7). Sensitivity was 67.8% (95% CI 61.0% to 74.1%) and specificity was 100% (95% CI 88.8% to 100%; 239 samples, 208 cases).
Three studies reported direct comparisons of different antigen assays in naso‐ or oropharyngeal samples; however none of the studies had any assay comparisons in common. All three studies utilised swabs in viral transport medium and all were conducted in symptomatic participants. We cannot derive any clear conclusions about comparative performance of tests from these studies.
Figure 8 shows variable diagnostic performance between and to some extent within studies. Four of the five assays in Fourati 2020 [A] demonstrated sensitivities in the range of 55% to 62% (SD Biosensor STANDARD Q, Abbott Panbio Covid‐19 Ag, Biosynex COVID‐19 Ag, AAZ – COVID‐VIRO), with one outlier (Coris Bioconcept – Covid‐19 Ag) at 35% (maximum of 297 cases). Specificity was 100% for all assays apart from SD Biosensor SDQ (specificity 93%; 337 pre‐pandemic samples).
In Porte 2020b [A] (32 cases) both assays had sensitivities over 90% (SD Biosensor STANDARD F and Quidel Sofia SARS Antigen), with specificities 97% (32 non‐COVID‐19 samples)
Weitzel 2020 [A] observed a range in assay sensitivities from 0% for the Liming Bio‐Products assay (based on only nine cases), to 17% (for Savant Biotech – Huaketai SARS‐CoV‐2 N), 62% (RapiGEN – BIOCREDIT COVID‐19 Ag) and 85% for Shenzhen Bioeasy Biotech – 2019 nCov Ag (78 to 80 cases for the latter three assays). Specificities were 100% for all assays (based on 30 to 31 samples) apart from the one from Liming Bio‐Products (specificity 90% based on 10 samples).
Average sensitivity and specificity for the 29 rapid molecular test evaluations that included samples with and without SARS‐CoV‐2, were 95.1% (95% CI 90.5% to 97.6%) and 98.8% (95% CI 98.3% to 99.2%; 4351 samples, 1781 with confirmed SARS‐CoV‐2; Table 6). Adding the three 'cases only' studies made little difference to the average sensitivity (95.5%, 95% CI 91.5% to 97.7%; 1973 cases).
Figure 9 demonstrates heterogeneity in sensitivity estimates (ranging from 57% to 100%), with consistently high specificities (92% to 100%, but with upper limits of 95% CIs of 99% or 100% in every study).
We extracted sensitivity data according to viral load from 10 evaluations of molecular tests, six of which reported data at a Ct threshold for higher viral load of 30 or less (Jokela 2020; Lieberman 2020; Mitchell 2020; Smithgall 2020 [A]; Smithgall 2020 [B]; Wolters 2020), four using Xpert Xpress and two using ID NOW. (Appendix 16)
All sensitivity estimates for the higher viral load subgroups were 100% (based on 204 samples with confirmed SARS‐CoV‐2), with a 95% CI for the average of 98.2% to 100%. For the lower viral load group, average sensitivity was 95.6% (95% CI 55.7% to 99.7%) (149 samples with confirmed SARS‐CoV‐2; Table 6).
We observed a similar pattern for the studies using alternative Ct thresholds to define higher and lower viral load (Appendix 17).
We did not observe any clear differences in average sensitivity or specificity when studies were separated by study design (2899 samples and 976 cases in 18 single‐group studies and 1265 samples and 718 cases in nine two‐group studies; Table 6; Appendix 17). Average sensitivity was higher in two‐group studies (97.2%, 95% CI 90.7% to 99.2%) compared to single‐group studies (93.2%, 95% CI 85.5% to 97.0%); a difference of 4.0 percentage points (95% CI from 2.2 percentage points lower to 10.1 higher). Average specificities had almost identical point estimates at 99.4% (95% CI 98.4 to 99.8%) and 99.3% (95% CI 96.5% to 99.8%) respectively (Table 6).
Thirteen studies evaluated the ID NOW assay, with 1949 samples and 730 confirmed SARS‐CoV‐2 cases; one study included only SARS‐CoV‐2‐positive cases (n = 36; Figure 10). Seven evaluations were laboratory‐based, three recruited participants from emergency department settings and three were conducted in multiple settings. Seven studies included only symptomatic participants, two included both symptomatic and asymptomatic people, and four did not report symptom status.
Eleven evaluations used nasopharyngeal or nasal swab samples, one was conducted using saliva samples and one did not specify the sample type. Only four evaluations were compliant with manufacturer IFUs; lack of compliance was based on the use of viral transport medium, sample type, and interval between sample collection and testing.
Pooled analyses demonstrated average sensitivity and specificity
Average sensitivity increased to 81.5% (95% CI 75.2% to 86.5%), with the addition of the cases only study (730 cases; Rhoads 2020).
The Xpert Xpress assay was evaluated in 15 studies using respiratory specimens, with 1781 samples and 1001 confirmed SARS‐CoV‐2 cases; two of the studies included only SARS‐CoV‐2‐positive cases (n = 90; Figure 10). Thirteen evaluations were laboratory‐based, one recruited participants from emergency department settings and one included samples from hospital inpatients. Three studies included only symptomatic participants, one included both symptomatic and asymptomatic people (mixed symptom status), and 11 did not report symptom status.
Fourteen evaluations used nasopharyngeal, oropharyngeal or nasal swab samples, and one was conducted using throat saliva or lower respiratory samples. Only three evaluations were compliant with manufacturer IFUs. Lack of compliance with the IFU was because of the use of frozen samples (n = 8), or sample type (n = 1) or concerns about the timing between sample collection and testing (n = 3).
Pooled analyses demonstrated average sensitivity and specificity
Average sensitivity did not change with addition of two cases‐only studies (99.1%, 95% CI 97.8% to 99.6%; n = 15; 730 cases; Broder 2020; Chen 2020a).
One additional study considered accuracy in non‐respiratory samples using Xpert Xpress (Szymczak 2020). Sensitivity in stool samples obtained up to 33 days after symptom onset was 93.1% (95% CI 77.2% to 99.1%) and specificity was 96.0% (95% CI 86.3% to 99.5%; 79 samples, 29 cases).
####### Comparison of ID NOW with Xpert Xpress
Comparing the overall pooled results between ID NOW and Xpert Xpress, the average sensitivity of Xpert Xpress was 19.8 (95% CI 14.9 to 24.7) percentage points higher than that of ID NOW (P < 0.0001; Table 6).
The average specificity of Xpert Xpress was marginally lower than that of ID NOW, a difference of −1.9 percentage points (95% CI −3.8 to −0.1).
We included one evaluation of COVID Nudge with a total of 386 participants and 71 SARS‐CoV‐2‐positive cases (Gibani 2020; Figure 10). Participants were recruited from multiple settings including hospital inpatients (n = 88), accident and emergency (n = 15) and healthcare workers and their families (n = 280). All participants were symptomatic and direct testing of nasopharyngeal samples was used (within manufacturer IFU).
The sensitivity of the COVID Nudge assay was 94.4% (95% CI 86.2 to 98.4%) and specificity was 100% (95% CI 98.8% to 100%; 386 samples and 71 cases; Table 6).
We included two evaluations of SAMBA II with 321 samples (121 with confirmed SARS‐CoV‐2; Figure 10). All participants were symptomatic. One study conducted direct testing of combined naso‐ or oropharyngeal samples from hospital inpatients and the other obtained combined naso‐ or oropharyngeal samples in viral transport medium from Public Health England. It was not reported whether the PHE samples were stored or frozen prior to testing so we could not determine whether they complied with the IFU for the assay.
The average sensitivity and specificity of SAMBA‐II were 96.0% (95% CI 81.1% to 99.3%) and 97.0% (95% CI 93.5% to 98.6%; 2 studies; 321 samples, 121 with confirmed SARS‐CoV‐2; Table 6).
In the IFU‐compliant evaluation, sensitivity was 87.9% (95% CI 71.8% to 96.6%) and specificity was 97.4% (95% CI 92.6% to 99.5%; 149 samples, 33 cases; Collier 2020; Table 6).
We included one evaluation of the Accula assay with a total of 100 samples (50 SARS‐CoV‐2 positive; Hogan 2020; Figure 10). The study was laboratory‐based and symptom status was not reported.
The study used nasopharyngeal samples in viral transport medium or saline, therefore the evaluation was not compliant with IFU requirements.
The sensitivity and specificity of the Accula test were 68.0% (95% CI 53.3% to 80.5%) and 100% (95% CI 92.9% to 100%; 100 samples, 50 cases; Table 6).
Six evaluations of molecular tests (in 1533 samples) reported results before and after discrepant analysis where selected samples were re‐tested with either the same (Collier 2020; Harrington 2020; Moran 2020; Stevens 2020), or an alternative RT‐PCR assay (Assennato 2020; Loeffelholz 2020). Four studies also reported re‐testing of samples with the index test (Assennato 2020; Collier 2020; Harrington 2020; Moran 2020; Appendix 16; Appendix 17).
Discrepant analysis reduces the number of samples deemed to be false negative or false positive errors. Discrepant analysis reduced the false negative proportion (1‐sensitivity) from 2.1% to 0.8% and the false positive rate (1‐specificity) from 2.2% to 0.4%. Three of the five studies reporting initially false positive results reported zero false positives after sample re‐testing and one reported a drop in false positives from 11 to 3 (Loeffelholz 2020; Appendix 16). Three of the four studies that reported re‐testing of initially false negative results reported reclassification as true negative on re‐testing, and in the other the single false negative remained as a false negative. Given the bias inherent in choosing the reference test dependent on the observed results, we caution against these findings.
An additional study tested all samples with two different RT‐PCR assays, and hence used a more accurate reference standard in all samples, not just samples with discrepant results (Moore 2020). Six initial true negatives were reclassified as false negatives after the second RT‐PCR. Had discrepant analysis been undertaken these misclassifications would have been missed, further underlining the methodological flaws inherent to discrepant analysis.
We also planned to evaluate the effect of sample type and reference standard.
For sample type, the use of variable combinations of sample types with or without viral transport media created numerous sparse subgroups by sample type (Appendix 18). Instead we considered study compliance with manufacturer IFU requirements which is a more pragmatic classification.
All studies used RT‐PCR alone as the reference standard for diagnosing SARS‐CoV‐2 infection.
We did not formally test for publication bias evident in the pattern of results, but did note that the identity of tests not meeting the PHE assessment criteria were not reported due to confidentiality agreements (PHE 2020(a)).
This is the second iteration of a Cochrane living review summarising the accuracy of point‐of‐care antigen and molecular tests for detecting current SARS‐CoV‐2 infection. This version of the review is based on published journal articles or studies available as preprints from 1 January 2020 up until 30 September 2020. In addition, we also included evaluations of antigen assays that were available as independent national reference laboratory publications or that were co‐ordinated and published by FIND, and journal articles that were listed on the Diagnostics Global Health website to 16 November 2020.
We included data from 77 studies using respiratory specimens, including 24,418 samples (7484 samples with confirmed SARS‐CoV‐2), and one study of faecal specimens (79 samples, 29 with confirmed SARS‐CoV‐2). Forty‐eight studies (reporting 58 test evaluations) considered antigen tests; 30 studies (reporting 33 test evaluations) considered rapid molecular tests, including the single study (evaluation) in faecal samples. Key findings are presented in the Table 1.
We summarise six key findings from this
Despite a considerable increase in the number of studies evaluating point‐of‐care tests, particularly antigen tests, there are still no published or preprint reports of accuracy for a significant number of commercially produced point‐of‐care tests. This review located evaluations for 16 antigen tests (three of which we could not identify as available for purchase) and five molecular assays. These represent a small proportion of assays currently on the market (118 commercialised antigen tests and 53 molecular assays).
The new studies have more robust and appropriate study designs compared to those in the first version of this review. Particularly for antigen tests where there are now studies recruiting participants from community‐based COVID‐19 testing clinics. Reporting of key details, such as settings and symptom status have improved, and studies are now evaluating direct swab testing as would occur in a point‐of‐care setting. However, concerns about risk of bias and applicability of results remain, and further improvements in study methods and reporting are needed before strong conclusions can be drawn about the accuracy of many antigen and molecular tests reviewed here. As it is not known whether these limitations will lead to over‐ or underestimates of test accuracy, estimates should be cautiously interpreted in context of their methodological limitations and the settings in which they were conducted. More direct comparisons of test brands are needed, with evaluations undertaken in the intended use settings for these tests.
Particular methodological concerns include the use of deliberate sampling according to known presence or absence of SARS‐CoV‐2 infection; use of anonymised samples submitted to laboratories for routine RT‐PCR testing (with no setting or participant details); and no information on symptoms or time from symptom onset. Differences in case‐mix related to symptomatic status, time post‐symptom onset and distribution of viral load are likely to have contributed to the observed variation in accuracy.
RT‐PCR was the reference standard in all studies ‐ no study defined the presence of COVID‐19 using clinical or radiological features in the absence of a negative RT‐PCR result.
Studies frequently did not follow the manufacturer’s instructions or did not use the test at the point of care. Fewer than half conducted the tests according to the manufacturers' IFU (41% (37/91); 29/58 antigen test evaluations and 8/33 molecular test evaluations). Reasons for non‐compliance included use of frozen samples, use of viral transport media, or lengthy intervals between sample collection and testing. Almost a third of studies (23/78) undertook on‐site, direct swab testing immediately or within an hour of sample collection; trained laboratory staff conducted tests in 16 (21%) studies, and 31 (40%) studies did not clearly describe the test operator and setting for the test procedure but we inferred that tests were carried out in a centralised laboratory setting, for example based on reported delays between collection and testing or reported use of archived or frozen samples.
For antigen test evaluations in symptomatic participants, we observed considerable heterogeneity in sensitivities (and to a lesser extent the specificities). Whilst the average sensitivity was 72.0% (95% CI 63.7% to 79.0%) and specificity was 99.5% (95% CI 98.5% to 99.8%), average sensitivity decreased with time since onset of symptoms, being higher in the first week (78.3%, 95% CI 71.1% to 84.1%) than when done later (51.0 95% CI 40.8% to 61.0%). Sensitivity was high in those with higher viral loads defined by Ct values ≤ 25 (94.5% 95% CI 91.0% to 96.7%) compared to those with lower viral loads (40.7%, 95% CI 31.8% to 50.3%). Focusing on studies that used the test in accordance with the manufacturer’s instructions, sensitivities for different brands varied from 34% to 96% (either based on pooled results or single studies). WHO have set a minimum 'acceptable' sensitivity requirement of 80%, and acceptable and ideal (or 'desirable') specificity requirements of 97% and 99% respectively (WHO 2020c). Only one assay (SD Biosensor STANDARD Q) met the WHO acceptable criterion for sensitivity based on pooled results of several studies. One further test (BIONOTE NowCheck) also met the acceptable sensitivity criterion, but only one study evaluated it. Abbott Panbio met the sensitivity criterion in individual studies but not overall. The acceptable performance criterion of 97% specificity was also met for all three tests, and two tests met the desirable criterion of more than 99% specificity (Abbott Panbio and SD Biosensor STANDARD Q).
Considerable heterogeneity in sensitivities remained after restricting analyses by test brand and symptom status, suggesting an effect not only from participant characteristics but from setting, sample type and collection method, sample storage and preparation, and testing procedures that cannot be easily unpicked. The PHE studies included in this review allow some consideration of the effect of test operator experience on the accuracy of the Innova test although different samples were tested by each test operator such that only an indirect comparison of sensitivity can be made. Sensitivity increased from 57.5% (95% CI 52.3%, 62.6%; 372 samples) when testing was conducted on‐site by trained non‐healthcare workers (PHE 2020(c) [non‐HCW tested]), to 70.0% (95% CI 63.5% to 75.9%; 223 samples) in samples tested on‐site by healthcare workers ((PHE 2020(d) [HCW tested]), to 78.8% (95% CI 72.4% to 84.3%; 198 samples) for those tested by laboratory scientists (PHE 2020(d) [Lab tested]). The effect of test operator on accuracy has been observed for rapid diagnostic tests for other infectious diseases such as malaria (Boyce 2018; Landier 2018), and is worthy of further investigation for diagnosis of SARS‐CoV‐2.
Twelve studies evaluated the accuracy of antigen tests in asymptomatic people for detection of SARS‐CoV‐2 infection defined by PCR status. As discussed, this does not address the issue of whether the test is identifying those who are infectious (as there is no reference standard that can be used). The average sensitivity for detecting infection in asymptomatic participants was 58.1% (95% CI 40.2% to 74.1%) with specificity of 98.9% (95% CI 93.6% to 99.8%), both lower than in symptomatic people. Only half of studies reported clearly defined asymptomatic cohorts (e.g. preventive screening in the general population (n = 1), in returning travellers (n = 1), or in contacts of confirmed cases (n = 4)), the other six reported asymptomatic subgroups from mixed symptom cohorts. Only one of the 12 studies provided data by viral load (Fenollar 2020(b)); 5% (1/22) of RT‐PCR‐positive samples had a Ct value of 25 or less, but 50% (11/22) had Ct values of 30 or less. No information on time after exposure to infection was reported.
For rapid molecular assays there were differences between test brands. Most data were for ID NOW and Xpert Xpress assays; average sensitivity for ID NOW was 78.6% (95% CI 73.7% to 82.8%) and Xpert Xpress 99.1% (95% CI 97.7% to 99.7%). Specificity for ID NOW was 99.8% (95% CI 99.23%, 99.9%) and Xpert Xpress 97.9% (95% CI 94.6% to 99.2%). These differences are beyond those expected by chance (P < 0.0001).
We were not able to investigate the effects of symptomatic status, or time from symptom 12/29 were from symptomatic populations, three from ‘mixed’ symptomatic and asymptomatic populations (percentage from each group not reported), and the remaining 14 evaluations provided no information on symptom status (2/14 recruited from A&E and 12 were laboratory‐based). These and other methodological limitations in the studies mean that we do not know how the assays would perform in any specific clinical setting when used in people suspected of having SARS‐CoV‐2 infection on the basis of symptoms, or of exposure to a confirmed case in the absence of symptoms. It is likely however that some difference in sensitivity between ID NOW and Xpert Xpress would be maintained in the absence of bias. The difference in specificity between the tests is small (ID NOW being 1.9% more specific compared to Xpert Xpress), but potentially important especially if used in a low‐prevalence setting. However, this difference in specificity would not be an issue should test‐positives be confirmed by a laboratory‐based RT‐PCR assay.
We did not formally compare antigen with molecular assays because there were no head‐to‐head comparisons of the two test types. Instead, we illustrate predicted numbers of true positives, false positives, false negatives and true negatives, applying summary estimates of test accuracy to a hypothetical cohort of people suspected of SARS‐CoV‐2 infection across a range in prevalence of SARS‐CoV‐2 infection (Table 1). For both antigen and molecular assays, we only use summary data from evaluations conducted in accordance with manufacturers’ IFUs, and for antigen tests we used separate results from symptomatic and asymptomatic participants.
For antigen test evaluations in symptomatic people, we selected three assays representing the range in observed average Coris Bioconcept COVID‐19 Ag Respi‐Strip (34.1% to 95% CI 29.7% to 38.8%), Abbott ‐ Panbio Covid‐19 Ag (75.1% to 95% CI 57.3% to 87.1%); and SD Biosensor ‐ STANDARD Q COVID‐19 Ag (88.1% to 95% CI 84.2% to 91.1%). Average specificities for the same three assays were 100% (95% CI 99.0% to 100%) to 99.5% (95% CI 98.7% to 99.8%) and 99.1% (95% CI 97.8% to 99.6%) respectively. Applied to a cohort of 1000 people with signs and symptoms of COVID‐19, in whom 50 people had confirmed infection (prevalence of 5%), for the three assays above we predicted
Increasing the prevalence to 10% or 20%, increases PPV and decreases NPV. As there is considerable heterogeneity in the estimates of sensitivity, the values observed in practice could vary considerably from these figures as shown by the estimates derived from the confidence intervals (Table 1).
For antigen test evaluations in asymptomatic participants there was considerably less available data from IFU‐compliant evaluations. We selected the same three exemplars, average sensitivities for identification of any infection (whether infectious or not) were lower than for symptomatic 28.6% (95% CI 8.4% to 58.1%) for the Coris Bioconcept assay; 48.9% (95% CI 35.1% to 62.9%) for the Abbott assay; and 69.2% (95% CI 38.6% to 90.9%) for the SD Biosensor assay. Average specificities for the same three assays 100% (95% CI 88.8% to 100%), 98.1% (95% CI 96.3% to 99.1%), and 99.1% (95% CI 95.2% to 100%).
Applying the average values to a larger cohort of 10,000 people asymptomatic for COVID‐19 and with a lower prevalence of 0.5% in whom 50 people had confirmed infection (infectious or not):
We derived the summary estimates used in these calculations from asymptomatic participants identified for testing in a number of scenarios and they cannot be directly translated to a particular setting, such as mass screening, for example. The confidence intervals for the average estimates used in these calculations are also extremely wide for both sensitivities and specificities, such that the numbers of false positives and false negatives observed in practice could differ substantially from these figures. Increasing the prevalence of confirmed SARS‐CoV‐2 infection to 1% or 2% makes little difference to the absolute number of false positive results for these assays, but has a large relative effect when considered in relation to the number of positive test results (PPVs for the Abbott and SD Biosensor assays increasing to 40% and 61% at 2% prevalence).
For molecular assays, data from IFU‐compliant evaluations were available for four of the five ID NOW (Abbott Laboratories), Xpert Xpress (Cepheid Inc), SAMBA II (Diagnostics for the Real World) and COVID Nudge (DNAnudge). Average sensitivities were derived as 73.0% (95% CI 66.8% to 78.4%), 100% (95% CI 88.1% to 100%), 87.9% (95% CI 71.8% to 96.6%) and 94.4% (95% CI 86.2% to 98.4%). Average specificities were 99.7% (95% CI 98.7% to 99.9%), 97.2% (95% CI 89.4% to 99.3%), 97.4% (95% CI 92.6% to 99.5%) and 100% (95% CI 98.8% to 100%), respectively (Table 1).
Data by symptom status for these assays were very limited, therefore we assumed that the intended use is most likely to be for diagnosis of acute infection in symptomatic individuals and have applied the average estimates of accuracy to a hypothetical cohort of 1000 people, at prevalences of 5%, 10% and 20% (Table 1). If 50 of 1000 people had confirmed infection (5% prevalence):
Increasing the prevalence of confirmed SARS‐CoV‐2 infection to 10% or 20% has a large relative effect when considered in relation to the number of positive test results for both Xpert Xpress and SAMBA II (PPVs were 64.9% and 63.8% at 5% prevalence compared to 90.1% and 89.3% at 20% prevalence). Less variation in PPV was observed for ID NOW and COVID‐Nudge because of the higher observed specificities. The NPV for the molecular assays is not affected to the same degree by these prevalence changes because of their relatively high sensitivities and the relatively low‐prevalence scenarios being considered.
Across all exemplar assays in the Table 1, we observed the widest variation in NPV for the Coris Bioconcept antigen assay in symptomatic participants (86% to 97%), demonstrating that even in a low‐prevalence setting, tests with poor sensitivity can have a considerable impact on the level of confidence that can be had in a negative test result.
Our review used a broad search screening all articles concerning COVID‐19 or SARS‐CoV‐2. We undertook all screening and eligibility assessments, QUADAS‐2 assessments (Whiting 2011), and data extraction of study findings independently and in duplicate. Although it is possible that the use of artificial intelligence text analysis to identify studies most relevant to diagnostic questions may have led to some eligible studies being missed, we believe that the multi‐stranded search strategy used will have identified most if not all relevant literature. Whilst we have reasonable confidence in the completeness and accuracy of the findings up until the search date, should errors be noted please inform us at coviddta@contacts.bham.ac.uk so that we can verify and correct in our next update.
We undertook a careful assessment of sample preparation and biosafety requirements as well as time to test result, to ensure that included tests were suitable for use at the point of care. The application of these index test criteria led to the exclusion of 39 of the 85 studies that we excluded on the basis of the index tests evaluated. Evaluations of alternative laboratory‐based molecular technologies are under consideration for inclusion in another review in our series of Cochrane COVID‐19 diagnostic test accuracy reviews. Furthermore, for this iteration of the review, we explicitly considered whether the test evaluations were conducted in accordance with the manufacturer IFU, regarding the sample types used, the use of viral transport medium and the permitted time between sample collection and testing.
We did not consider any manufacturer statements on the intended use of the tests by population, but we are aware that some IFUs recommend testing only in symptomatic people and within certain time frames after symptom onset (e.g. the Innova assay). Where possible, however, we did provide data separately for symptomatic and asymptomatic participants and identified clear trends towards lower sensitivities in asymptomatic individuals for detection of infection. We were unable to assess the accuracy of antigen tests for identification of infectious individuals, as there is no established reference standard for infectiousness (and it seems unlikely that one will ever be established). We have presented results by Ct value where it has been reported by the individual studies. We recognise the limitations from this approach, and given the extent to which RT‐PCR Ct values vary between assays (Vogels 2020), and between laboratories, we strongly caution against the direct application of our results in high and low Ct value subgroups to any particular clinical context. There is no 'step change' in 'infectiousness' according to any fixed Ct value; increasing numbers of studies demonstrate successful viral culture in individuals considered to have 'low' viral load (Jaafar 2020; Singanayagam 2020), and, more importantly, that transmission of infection does occur from index cases with low RT‐PCR Ct values (Lee 2021; Marks 2021). Ultimately, viral load on its own is only one factor influencing an individual's ability to transmit infection, 'infectiousness' being modified by host factors such as the health of an individual’s immune system or presence of comorbidities, and environmental risk factors including closeness and length of contact with others.
Weaknesses of the review primarily reflect the weaknesses in the primary studies and their reporting. Although study quality improved in comparison to the first iteration of this review, many studies continue to omit descriptions of participants, and key aspects of study design and execution. In order to include data for all tests in pooled analyses we had to include some samples multiple times. We have been explicit about these issues where they arose. It is possible that eligible studies have been missed by our search strategy however we believe the risk to be very low considering our broad approach to identification of literature. Despite our best efforts to be as comprehensive as possible, new evaluations are continuously becoming available and it is impossible for any published and peer‐reviewed systematic review to be fully up to date.
Around a quarter (18/78) of the studies we have included are currently only available as preprints, and as yet, have not undergone peer review. As published versions of these studies are identified in the future, we will double‐check study descriptions, methods and findings, and update the review as required.
There are an increasing number of roles and testing strategies for which antigen and rapid molecular assays are considered, and it is likely that the performance of these tests needs to be considered separately for each of the use cases.
Our review shows that antigen tests do not appear to perform as well in asymptomatic populations compared to symptomatic populations for detecting infection. The amount of available data for asymptomatic populations is less than that from symptomatic populations and is also based on asymptomatic individuals tested in a range of scenarios, from preventive or targeted screening, to contact tracing or testing at dedicated COVID‐19 test centres, which may explain some of the observed variability. It is also not clear whether individuals in these studies were truly cases of asymptomatic infection as opposed to pre‐ or post‐symptomatic, or were even mildly symptomatic and mislabelled as asymptomatic. Incomplete symptom assessment and lack of adequate follow‐up to identify subsequent development of symptoms or previous history of symptoms can all contribute to inappropriate classification of individuals as asymptomatic infection (Meyerowitz 2020). As the studies in our review did not systematically attempt to identify pre‐ or post‐symptomatic individuals, it may be more appropriate to consider the estimates for test accuracy for asymptomatic populations as primarily representing accuracy in those without clearly defined symptoms at the time of testing.
We are aware that several important studies in asymptomatic individuals have been reported since the close of our search. In mass screening in Liverpool, Innova was positive in 28 of 70 PCR‐detected cases (sensitivity for infection 40.0%, 95% CI 28.5% to 52.4%) and 26 of 39 with Ct values less than 25 (sensitivity 66.7%, 95% CI 49.8% to 80.9%). Screening University of Birmingham students found 2 of 7185 students positive with Innova, and estimated sensitivity of 3.2% (95% CI 0.6% to 15.6%) for detecting any infection, 9.1% (95% CI 1.0% to 49.1%) for Ct values less than 30 and 100% (95% CI 15.8% to 100%) for Ct less than 25 (Ferguson 2020). BinaxNOW (which uses the same test strip as PanBio) has been tested in asymptomatic in San Francisco the test detected 7 of 11 PCR‐positive cases (sensitivity 63.6%, 95% CI 30.8% to 89.1%), and 6 of 6 with Ct values less than 30 (100%, 95% CI 54.1% to 100%; Pilarowski 2021); in a drive‐through centre in Massachusetts it detected the virus in 70 of 107 in adults (sensitivity 65.4%, 95% CI 55.6 to 74.4) and 40 of 57 in children (70.2%, 95% CI 56.6% to 81.6%)); no breakdown by viral load is available (Pollock 2020). The specificity of the tests in all studies has remained high (above 99%). This selection of results is not based on a systematic search (this will occur in the next update) but these results suggest that emerging evidence is illustrating a range of sensitivity values for the ability of the tests to detect infection, with high detection rates only in groups with very high viral loads.
Given the superior test performance characteristics for symptomatic populations in the first week of symptoms and in those with higher viral loads, the observed poorer performance in those without symptoms is perhaps not surprising. Evidence suggests that higher viral loads are observed in the first week of illness, beginning two days prior to the development of symptoms (Cevik 2021). Viral load patterns in asymptomatic people are less clear but similarly high titers of SARS‐CoV‐2 have been observed at the onset of infection with a suggestion of faster clearance (Cevik 2021). However, variation in viral trajectories means that even if an asymptomatic person can identify a clear contact with a confirmed case of SARS‐CoV‐2 infection, it is not possible to pinpoint when (or even if) that individual will have a sufficient viral load to be detected on antigen testing. A serial testing policy would be likely to identify at least some infected asymptomatic contacts, but comes at the cost of increased numbers of false positives, especially in low‐prevalence settings. There were no evaluations of serial testing in any of the studies.
For molecular tests, we observed a lack of studies undertaken in intended use settings, with most data being from laboratory testing. Although more evidence is available for accuracy in symptomatic people, applicability issues regarding the way in which the tests are carried out and in how cases of SARS‐CoV‐2 infection are defined remain, and it is not yet possible to determine how tests will perform in practice.
We recommend caution in applying the results outside of the individual study (or closely related) contexts and use case scenarios.
Review first Issue 8, 2020
Members of the Cochrane COVID‐19 Diagnostic Test Accuracy Review Group
The editorial process for this review was managed by Cochrane's Editorial and Methods Department Central Editorial Service in collaboration with Cochrane Infectious Diseases. We thank Helen Wakeford, Anne‐Marie Stephani and Deirdre Walshe for their comments and editorial management. We thank Liz Bickerdike for comments on the Abstract. We thank Robin Featherstone and Douglas M Salzwedel for comments on the search and Mike Brown and Paul Garner for sign‐off comments. We thank Denise Mitchell for her efforts in copy‐editing this review.
Thank you also to peer referees Kristien Verdonck, David Sinclair and Jim Hugget, consumer referees Brian Duncan and Ceri Dare, methodological referees Mia Schmidt‐Hansen and Jo Leonardi‐Bee, for their insights.
The editorial base of Cochrane Infectious Diseases is funded by UK aid from the UK Government for the benefit of low‐ and middle‐income countries (project number 300342‐104). The views expressed do not necessarily reflect the UK Government’s official policies.
The authors thank Dr Mia Schmidt‐Hansen who was the Cochrane Diagnostic Test Accuracy (DTA) Contact Editor for this review; the clinical and methodological referees; the Cochrane DTA Editorial Team; and Anne Lawson who copy‐edited the protocol. We would also like to thank all corresponding authors who provided additional information regarding their studies, and colleagues at both the Norwegian Insititute of Public Health and the EPPI‐Centre who provided updates from their COVID‐19 living evidence maps.
Jonathan Deeks is a UK National Institute for Health Research (NIHR) Senior Investigator Emeritus. Yemisi Takwoingi is supported by a NIHR Postdoctoral Fellowship. Jonathan Deeks, Jacqueline Dinnes, Yemisi Takwoingi, Clare Davenport and Malcolm Price are supported by the NIHR Birmingham Biomedical Research Centre. Sian Taylor‐Phillips is supported by an NIHR Career Development Fellowship. This paper presents independent research supported by the NIHR Birmingham Biomedical Research Centre at the University Hospitals Birmingham NHS Foundation Trust and the University of Birmingham. The views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health and Social Care.
Includes laboratory testing guidelines and global surveillance guidelines
^a^Source data from Laboratory testing of 2019 novel coronavirus (2019‐nCoV) in suspected human interim guidance, World Health Organization. 10 January, 17 January, 2 March, 19 March, 21 March 2020 (WHO 2020d), and Global surveillance for COVID‐19 caused by human infection with COVID‐19 virus, interim guidance, 31 January, 27 February, and 20 March 2020 (WHO 2020e).
The following information is taken from the university of Bern website (see: ispmbern.github.io/covid-19/living-review/collectingdata.html).
The register is updated daily and CSV file downloads are made available.
From 1 April 2020, we will retrieve the curated BioRxiv/MedRxiv dataset (connect.medrxiv.org/relate/content/181).
MEDLINE: ("Wuhan coronavirus" [Supplementary Concept] OR "COVID‐19" OR "2019 ncov"[tiab] OR (("novel coronavirus"[tiab] OR "new coronavirus"[tiab]) AND (wuhan[tiab] OR 2019[tiab])) OR 2019‐nCoV[All Fields] OR (wuhan[tiab] AND coronavirus[tiab])))))
Embase: (nCoV or 2019‐nCoV or ((new or novel or wuhan) adj3 coronavirus) or covid19 or covid‐19 or SARS‐CoV‐2).mp.
BioRxiv/MedRxiv: ncov or corona or wuhan or COVID or SARS‐CoV‐2
With the kind support of the Public Health & Primary Care Library PHC (www.unibe.ch/university/services/university_library/faculty_libraries/medicine/public_health_amp_primary_care_library_phc/index_eng.html), and following guidance of the Medical Library Association (www.mlanet.org/p/cm/ld/fid=1713).
MEDLINE: ("Wuhan coronavirus" [Supplementary Concept] OR "COVID‐19" OR "2019 ncov"[tiab] OR (("novel coronavirus"[tiab] OR "new coronavirus"[tiab]) AND (wuhan[tiab] OR 2019[tiab])) OR 2019‐nCoV[All Fields] OR (wuhan[tiab] AND coronavirus[tiab])))))
Embase: ncov OR (wuhan AND corona) OR COVID
BioRxiv/MedRxiv: ncov or corona or wuhan or COVID
We needed a more efficient approach to keep up with the rapidly increasing volume of COVID‐19 literature. A classification model for COVID‐19 diagnostic studies was built with the model building function within Eppi Reviewer, which uses the standard SGCClassifier in Scikit‐learn on word trigrams. As outputs, new documents receive a percentage (from the predict_proba function) where scores close to 100 indicate a high probability of belonging to the class ‘relevant document’ and scores close to 0 indicate a low probability of belonging to the class ‘relevant document’. We used three iterations of manual screening (title and abstract screening, followed by full‐text review) to build and test classifiers. The final included studies were used as relevant documents, while the remainder of the COVID‐19 studies were used as irrelevant documents. The classifier was trained on the first round of selected articles, and tested and retrained on the second round of selected articles. Testing on the second round of selected articles revealed poor positive predictive value but 100% sensitivity at a cut‐off of 10. The poor positive predictive value is mainly due to the broad scope of our topic (all diagnostic studies in COVID‐19), poor reporting in abstracts, and a small set of included documents. The model was retrained using the articles selected of the second and third rounds of screening, which added a considerable number of additional documents. This led to a large increase in positive predictive value, at the cost of a lower sensitivity, which led us to reduce the cut‐off to 5. The largest proportion of documents had a score between 0‐5. This set did not contain any of the relevant documents. This version of the classifier with a cut‐off 5 was used in subsequent rounds and accounted for approximately 80% of the screening burden.
Embase records from the Stephen B. Thacker CDC Library, COVID‐19 Research articles Downloadable database
Records were obtained by the CDC library by searching Embase through Ovid using the following search strategy.
Figure 61
11 Risk of bias and applicability concerns review authors' judgements about each domain presented as percentages across included studies
Figure 62
12 Risk of bias and applicability concerns review authors' judgements about each domain presented as percentages across included studies
Figure 63
13 Risk of bias and applicability concerns review authors' judgements about each domain for each included study
Figure 64
14 Forest plot of antigen test evaluations by study design. BR: Brazil; CH: Switzerland; DE: Germany; HCW: healthcare worker
Figure 65
15 Forest plot of studies evaluating antigen higher versus lower viral load (< or > 25 Ct). BR: Brazil; CH: Switzerland; Ct: cycle threshold; DE: Germany; HCW: healthcare worker
Figure 66
16 Forest plot of studies evaluating antigen higher versus lower viral load (< or > 32/33 Ct threshold). BR: Brazil; CH: Switzerland; ; Ct: cycle threshold; DE: Germany
Figure 67
17 Forest plot of studies evaluating antigen higher versus lower viral load (other Ct thresholds). Ct: cycle threshold; HCW: healthcare worker
Figure 68
18 Forest plot of molecular test evaluations by study design
Figure 69
19 Forest plot of studies evaluating rapid molecular high versus low viral load (30 Ct threshold). Ct: cycle threshold
Figure 70
20 Forest plot of studies evaluating rapid molecular high versus low viral load (other Ct thresholds). Ct: cycle threshold
Figure 71
21 Rapid molecular assays before and after discrepant analysis
Presented below are all the data for all of the tests entered into the review.
1 TestAntigen tests ‐ All
2 TestAntigen tests ‐ symptomatic
3 TestAntigen tests ‐ asymptomatic
4 TestAntigen tests ‐ mixed symptoms or not reported
5 TestAntigen tests ‐ Ct values < or <=25
6 TestAntigen tests ‐ Ct values >25
7 TestAntigen tests ‐ Ct values < or <=32/33
8 TestAntigen tests ‐ Ct values >32/33
9 TestAntigen tests ‐ other Ct thresholds for 'higher' viral load
10 TestAntigen tests ‐ other Ct thresholds for 'lower' viral load
11 TestAntigen tests ‐ week 1 after symptom onset
12 TestAntigen tests ‐ week 2 after symptom onset
13 TestMolecular tests ‐ all
14 TestMolecular tests ‐ all (before discrepant analysis)
15 TestMolecular tests ‐ all (after discrepant analysis)
16 TestMolecular tests ‐ Ct values < or <=30
17 TestMolecular tests ‐ Ct values >30
18 TestMolecular tests ‐ other Ct thresholds for 'higher' viral load
19 TestMolecular tests ‐ other Ct thresholds for 'lower' viral load
20 TestMolecular tests ‐ other sites
21 TestAntigen tests ‐ direct comparisons
22 TestAAZ ‐ COVID‐VIRO (CGIA)
23 TestAbbott ‐ Panbio Covid‐19 Ag (CGIA)
24 TestBecton Dickinson ‐ BD Veritor (LFA – method not specified)
25 TestBIONOTE ‐ NowCheck COVID‐19 Ag (LFA – method not specified)
26 TestBiosynex ‐ Biosynex COVID‐19 Ag BSS (CGIA)
27 TestCoris Bioconcept ‐ COVID‐19 Ag Respi‐Strip (CGIA)
28 TestE25Bio ‐ DART (NP) (CGIA)
29 TestFujirebio ‐ ESPLINE SARS‐CoV‐2 [LFA(ALP)]
30 TestInhouse (Bioeasy co‐author) ‐ n/a (FIA)
31 TestInnova Medical Group ‐ Innova SARS‐CoV‐2 Ag (CGIA)
32 TestLiming Bio‐Products ‐ StrongStep® COVID‐19 Ag (CGIA)
33 TestQuidel Corporation ‐ SOFIA SARS Antigen (FIA)
34 TestRapiGEN ‐ BIOCREDIT COVID‐19 Ag (CGIA)
35 TestRoche ‐ SARS‐CoV‐2 (LFA – method not specified)
36 TestSavant Biotech ‐ Huaketai SARS‐CoV‐2 N Protein (LFA – method not specified)
37 TestSD Biosensor ‐ STANDARD F COVID‐19 Ag (FIA)
38 TestSD Biosensor ‐ STANDARD Q COVID‐19 Ag (CGIA)
39 TestShenzhen Bioeasy Biotech ‐ 2019‐nCoV Ag (FIA)
40 TestAbbott ‐ ID NOW (Isothermal PCR)
41 TestCepheid ‐ Xpert Xpress (Automated RT‐PCR)
42 TestDNANudge – COVID Nudge (Automated RT‐PCR)
43 TestDRW ‐ SAMBA II (Automated RT‐PCR)
44 TestMesa Biotech ‐ Accula (other molecular)
45 TestAntigen test evaluations ‐ Single group design
46 TestAntigen test evaluations ‐ Two group design
47 TestAntigen test evaluations ‐ Unclear design
48 TestMolecular test evaluations ‐ Single group design
49 TestMolecular test evaluations ‐ Two group design
50 TestMolecular test evaluations ‐ Unclear design
We planned to check the following websites for eligible index tests, however these did not prove to be very accessible or easy to use and, after initial review, were not further
We planned to check the following evidence repository for additional eligible studies however, the EPPI‐Centre and Norwegian Institute of Public Health resources proved to be more accessible therefore we decided to prioritise our other sources of evidence.
We intended for two authors to independently perform data extraction, however one review author extracted study characteristics, and a second author checked them. Contingency table data were extracted independently by two review authors as planned.
We planned to evaluate the effect of additional sources of heterogeneity, including reference standard and sample type. However, additional formal investigations using meta‐regression were not possible because of lack of variability across the studies in these features.
We planned to conduct a sensitivity analysis excluding studies that are solely published as preprints. We have inadequate study numbers to allow this at present but will reconsider for the next update.
JD was the contact person with the editorial base. JDI co‐ordinated contributions from the co‐authors and wrote the final draft of the review. JJD, JDi, YT, CD, STP, IH, AA, LFR, MP, MT, JDr, SB screened papers against eligibility criteria. RS conducted the literature searches. JDi, MT and AA appraised the quality of papers. JDi, MT and AA extracted data for the review and sought additional information about papers. JDi entered data into Review Manager 2020. JDi, JJD, YT and SB, analysed and interpreted data. JJD, JDi, YT, CD, STP, RS, ML, LH, AVB, DE, SD, JC worked on the methods sections and commented on the draft review. JJD and JDi responded to the comments of the referees. JJD is the guarantor of the update.
Jonathan J Deeks: JD has published or been quoted in opinion pieces in scientific publications, and in the mainstream and social media related to diagnostic testing. JD was the statistician on the Birmingham evaluation of the Innova test which is mentioned in the discussion of the paper. There was no funding for this evaluation of the Innova test. JD is a member of the Royal Statistical Society (RSS) COVID‐19 taskforce steering group, and co‐chair of the RSS Diagnostic Test Advisory Group. He is a consultant adviser to the WHO Essential Diagnostic List. JD receives payment from the BMJ as their Chief Statistical advisor.
Jacqueline Dinnes: none known
Yemisi Takwoingi: none known
Clare Davenport: none known
Mariska MG Leeflang: none known
René Spijker: none known
Lotty Hooft: none known
Ann Van den Bruel: none known
Devy Emperador: is employed by FIND with funding from DFID and KFW. FIND is a global non‐for profit product development partnership and WHO Diagnostic Collaboration Centre. It is FIND’s role to accelerate access to high‐quality diagnostic tools for low‐resource settings and this is achieved by supporting both R&D and access activities for a wide range of diseases, including COVID‐19. FIND has several clinical research projects to evaluate multiple new diagnostic tests against published Target Product Profiles that have been defined through consensus processes. These studies are for diagnostic products developed by private sector companies who provide access to know‐how, equipment/reagents, and contribute through unrestricted donations as per FIND policy and external SAC review.
Sabine Dittrich: is employed by FIND with funding from DFID and Australian Aid. FIND is a global non‐for profit product development partnership and WHO Diagnostic Collaboration Centre. It is FIND’s role to accelerate access to high‐quality diagnostic tools for low‐resource settings and this is achieved by supporting both R&D and access activities for a wide range of diseases, including COVID‐19. FIND has several clinical research projects to evaluate multiple new diagnostic tests against published Target Product Profiles that have been defined through consensus processes. These studies are for diagnostic products developed by private sector companies who provide access to know‐how, equipment/reagents, and contribute through unrestricted donations as per FIND policy and external SAC review.
Ada Adriano: none known
Sophie Beese: none known
Janine Dretzke: none known
Lavinia Ferrante di Ruffano: none known
Isobel Harris: none known
Malcolm Price: none known
Sian Taylor‐Phillips: none known
Sarah Berhane: none known
Jane Cunningham: none known
Edited (no change to conclusions)