Authors: Elisabet Rondung, Pamela Massoudi, Katri Nieminen, Birgitta Wickberg, Nathalie Peira, Rebecca Silverstein, Klas Moberg, Martina Lundqvist, Åke Grundberg, Monica Hultcrantz
Categories: Systematic Reviews, anxiety, depression, identification, meta‐analysis, pregnancy, screening, systematic review, Systematic Review
Source: Acta Obstetricia et Gynecologica Scandinavica
Doi: 10.1111/aogs.14734
Authors: Elisabet Rondung, Pamela Massoudi, Katri Nieminen, Birgitta Wickberg, Nathalie Peira, Rebecca Silverstein, Klas Moberg, Martina Lundqvist, Åke Grundberg, Monica Hultcrantz
Depression and anxiety are significant contributors to maternal perinatal morbidity and a range of negative child outcomes. This systematic review and meta‐analysis aimed to review and assess the diagnostic test accuracy of selected screening tools (Edinburgh Postnatal Depression Scale [EPDS], EPDS‐3A, Patient Health Questionnaire [PHQ‐9]‐, PHQ‐2, Matthey Generic Mood Question [MGMQ], Generalized Anxiety Disorder scale [GAD‐7], GAD‐2, and the Whooley questions) used to identify women with antenatal depression or anxiety in Western countries.
On January 16, 2023, we searched 10 databases (CINAHL, Cochrane Library, CRD Database, Embase, Epistemonikos, International HTA Database, KSR Evidence, Ovid MEDLINE, PROSPERO and PsycINFO); the references of included studies were also screened. We included studies of any design that compared case‐identification with a relevant screening tool to the outcome of a diagnostic interview based on the Diagnostic and Statistical Manual of Mental Disorders, fourth or fifth edition (DSM‐IV or DSM‐5), or the International Statistical Classification of Diseases and Related Health Problems, 10th revision (ICD‐10). Diagnoses of interest were major depressive disorder and anxiety disorders. Two authors independently screened abstracts and full‐texts for relevance and evaluated the risk of bias using QUADAS‐2. Data extraction was performed by one person and checked by another team member for accuracy. For synthesis, a bivariate model was used. The certainty of evidence was assessed using Grading of Recommendations Assessment, Development and Evaluation (GRADE). Registration: PROSPERO CRD42021236333.
We screened 8276 records for eligibility and included 16 original articles reporting on diagnostic test 12 for the EPDS, one article each for the GAD‐2, MGMQ, PHQ‐9, PHQ‐2, and Whooley questions, and no articles for the EPDS‐3A or GAD‐7. Most of the studies had moderate to high risk of bias. Ten of the EPDS articles provided data for synthesis at cutoffs ≥10 to ≥14 for diagnosing major depressive disorder. Cutoff ≥10 gave the optimal combined sensitivity (0.84, 95% confidence interval [CI]: 0.75–0.90) and specificity (0.87, 95% CI: 0.79–0.92).
Findings from the meta‐analysis suggest that the EPDS alone is not perfectly suitable for detection of major depressive disorder during pregnancy. Few studies have evaluated the other instruments, therefore, their usefulness for identification of women with depression and anxiety during pregnancy remains very uncertain. At present, case‐identification with any tool may best serve as a complement to a broader dialogue between healthcare professionals and their patients.
Key messageIn this systematic review, the Edinburgh Postnatal Depression Scale was the most frequently validated tool for identification of women with possible antenatal depression or anxiety. Our meta‐analyses indicated a low pooled sensitivity, emphasizing the importance of combining questionnaire screening with other identification methods.
Antenatal depression and anxiety are significant contributors to maternal morbidity in the perinatal period, ^1^ with prevalence rates ranging between 9% and 19%. ^2^ , ^3^ Antenatal depression and anxiety do not only affect women, but have also been associated with a broad range of negative child outcomes, such as emotional problems, externalizing difficulties, attachment issues and a less positive cognitive development. ^4^ , ^5^ Despite this, most research has focused on the identification and treatment of mothers' postnatal emotional problems, most frequently on postnatal depression. ^1^ For example, Sambrook Smith et al. ^6^ identified 30 systematic reviews synthesizing validation studies of screening tools for perinatal mental disorders. Of these, 27 evaluated tools for depression, two of which focused on antenatal depression. Only four reviews validated tools for identification on anxiety. Despite the possibility that antenatal anxiety disorders may be more prevalent than antenatal depression, ^7^ it appears to have been less studied.
In 2020 we performed systematic searches (registered in PROSPERO), ^8^ indicating that no RCTs assessing the effects of programs designed to identify women with depression or anxiety during pregnancy had been published. Both previously published ^9^ , ^10^ and later ^11^ systematic reviews have reported similar findings, revealing a paucity of primary studies evaluating potential positive and negative effects of structured case‐identification of depression and anxiety during pregnancy.
In the absence of clear effect evaluations, the question of screening for antenatal mental health problems has been debated, and national guidelines have reached different conclusions regarding how structured identification of women with antenatal depression and anxiety should be conducted. ^12^ , ^13^ , ^14^ , ^15^ There is, however, a general consensus among professionals that identification of women with possible antenatal depression or anxiety by trained health care staff is good clinical practice ^5^ and that identification can be facilitated by the use of structured approaches, such as standard identification questions ^12^ or specific assessment tools. ^5^
With regard to the diagnostic test accuracy (DTA) of screening instruments aimed at identifying cases of antenatal depression or anxiety, earlier reviews have typically validated one scale at the time, most commonly in a postnatal or perinatal population, ^16^ , ^17^ , ^28^ and have shown sparse data for several instruments, ^12^ indicating a need for further DTA validations. Earlier reviews have also suggested that cultural differences may affect test accuracy. ^19^ Since the most recent systematic reviews have evaluated the psychometric properties of instruments regardless of cultural context, ^16^ , ^17^ , ^20^ results might be difficult to generalize to specific contexts. Hence, systematic reviews of how instruments for identification of antenatal depression and anxiety perform in different cultural contexts are needed.
This study was planned to inform national policy documents in Sweden and include instruments that are feasible to use in maternity care settings within primary care. From this perspective we decided to focus on the Edinburgh Postnatal Depression Scale (EPDS) ^21^ and its anxiety subscale EPDS‐3A, the Patient Health Questionnaire (PHQ‐9) ^22^ and its short version PHQ‐2, ^23^ the Generalized Anxiety Disorder scale (GAD‐7) ^24^ and its subscale GAD‐2, ^25^ the Whooley questions, ^26^ and the Matthey Generic Mood Questionnaire (MGMQ). ^27^ Of these, only the MGMQ was developed with the antenatal population in mind, while the others were developed either for the general population ^22^ , ^23^ , ^24^ , ^25^ , ^26^ or for postpartum use. ^21^ For our findings to be culturally relevant to our context, we decided to include studies conducted in Western countries. Hence, the specific research question that guided this systematic review What is the diagnostic accuracy or comparative accuracy of the following self‐report instruments, alone or in combination, for detecting depression and/or anxiety during the EPDS, EPDS‐3A, PHQ‐9, PHQ‐2, MGMQ, GAD‐7, GAD‐2, and Whooley questions?
This study is an updated part of a broader review at The Swedish Agency for Health Technology Assessment and Assessment of Social Services (SBU). The full report in Swedish was submitted to the Swedish National Board for Health and Welfare in June 2021 to support policy development. A peer‐reviewed protocol was registered in PROSPERO (CRD42021236333) on March 15, 2021. This systematic review follows the general concepts covered by PRISMA. ^28^
A detailed description of our inclusion and exclusion criteria is available in our study protocol (CRD42021236333).
Studies eligible for inclusion had a population of pregnant women without an established clinical diagnosis of depression or anxiety, regardless of risk‐ or other predisposing factors. We excluded studies based on selected populations; that is, populations with specific somatic or psychiatric diagnoses, or specific age groups.
The following screening instruments were EPDS, EPDS‐3A, PHQ‐9, PHQ‐2, MGMQ, GAD‐7, GAD‐2, and the Whooley questions. Any combination of these index tests were also included.
The reference standard was a diagnosis according to the Diagnostic and Statistical Manual of Mental Disorders, fourth or fifth edition (DSM‐IV and DSM‐5) or the International Statistical Classification of Diseases and Related Health Problems 10th revision (ICD‐10). The diagnosis needed to be based on a structured clinical interview, corresponding to any of the above diagnostic systems. Diagnoses of interest were major depressive disorder (MDD), and anxiety disorders (primarily generalized anxiety disorder, panic disorder, social anxiety disorder, and specific phobias). Minor depression, and combinations of major and minor depressions were excluded to reflect the diagnostic criteria set out in the DSM‐5. ^29^ Optimally, the reference test should have been conducted by a mental health professional within 2 weeks of the index test, but this was not a requirement for inclusion.
The outcomes of interest were the sensitivity and specificity of the instruments of interest for diagnosing depression or anxiety disorders, or comparisons of the sensitivity and specificity between instruments.
DTA studies could have any study design, provided sensitivity and specificity were presented or could be calculated. Systematic reviews were included so that their reference lists could be screened for additional relevant studies but were not included in analyses.
To minimize cultural differences that might potentially affect the psychometric properties of the tools, only studies conducted in Europe, USA, Canada, Australia or New Zealand were included.
The systematic literature searches included the databases CINAHL with full text (EBSCO), Cochrane Library (Wiley), CRD Database (including DARE, HTA and NHS EED), Embase (Elsevier), Epistemonikos, International HTA Database, KSR Evidence, Ovid MEDLINE(R) ALL, PROSPERO and PsycINFO (EBSCO). Reference lists from published articles and overviews were scrutinized for additional inclusion.
The search strategy was constructed and performed by an information specialist using appropriate controlled vocabulary and relevant text word terms. Searches were initially run on December 11, 2020 to inform two reports ^30^ , ^31^ sent to the Swedish National Board for Health and Welfare in June 2021 to support policy development. These reports address a wider set of questions than this systematic review, including eg, the effectiveness of screening programs. The part of the literature search concerning the question being answered in this systematic review was updated on January 16, 2023. Since a previous systematic review from the National Institute for Health and Care Excellence (NICE) ^12^ addressed this topic in 2014, a date limit was set to articles published in 2014 or later. The searches were limited to English, Swedish, Norwegian and Danish and the publication type case reports, comments, editorials, and letters were excluded. Duplicates were removed using EndNote's duplicate identification strategy and then manually. ^32^ The detailed search strategy can be found in Table S1.
All titles and abstracts were screened independently by two review authors using the Rayyan software. ^33^ All studies selected by at least one person were checked independently in full‐text by two authors. Discrepancies were resolved by discussion with the whole review team.
One review author extracted the data from primary studies, while a second author checked the integrity of the extracted information. Study authors were contacted when the data provided in the article was insufficient to calculate 2 by 2 tables (i.e., true positive, false positive, false negative, true negative).
The data extraction form included the following information from primary author and year; country of origin; participant characteristics; index test methods; reference test methods; test accuracy data; and information needed to assess risk of bias.
Risk of bias in primary DTA studies was assessed independently by two review authors, using QUADAS‐2. ^34^ Any disagreements were resolved by discussion, either with a third review author, or, when necessary, with the whole review team.
As effect measures, we used sensitivity and specificity. We extracted relevant data to get a full 2 by 2 table for each cutoff of interest for all three conditions being (1) MDD, (2) anxiety disorder, or (3) MDD, anxiety disorder or both. If data was reported for more than one time point during pregnancy, we selected the earliest. For the EPDS we selected cutoffs ≥10 to ≥14, based on a span of +/− 1 relative to the suggested cutoffs ≥11 and ≥13 in Levis et al. ^18^
To provide clinically informative estimates of sensitivity and specificity we conducted separate analyses of test data at different thresholds. Where sufficient data was available, we summarized diagnostic accuracy data using exact bivariate binomial distributions ^35^ by fitting a generalized linear mixed model using the glmer function in the package lme4 in the software R (described in detail in chapter “6.2 Meta‐analysis using glmer,” by Partlett and Takwoingi ^36^ ). The model was set to use study‐level random effects and one point per axis for evaluating the adaptive Gauss‐Hermite approximation to the log‐likelihood (Laplace approximation). The forest and SROC plots were generated in RevMan ^37^ using parameters obtained from the glmer function. The assessment of heterogeneity was based on a qualitative inspection of the forest and SROC plots and considered when rating the certainty of evidence (see Inconsistency below).
We present summary estimates of sensitivity and specificity, together with 95% confidence intervals (CI) and the Youden's J statistic, showing the overall performance when combining sensitivity and specificity (J = sensitivity + specificity – 1). In addition, we also present likelihood ratios derived from the estimates of sensitivity and specificity (LR+ = sensitivity/[1‐specificity]; LR‐ = [1‐sensitivity]/specificity) and illustrate what the estimates of sensitivity and specificity would mean in terms of negative and positive predictive values at a specific prevalence.
Publication bias was addressed in our Grading of Recommendations Assessment, Development and Evaluation (GRADE) assessment (see below). No test for funnel plot asymmetry was used.
The certainty of the evidence for each outcome was assessed using GRADE by evaluating the following five domains ^38^ , ^39^ : Risk of Ratings depended on the QUADAS‐2 assessments, focusing on patient selection, index test, reference standard, and flow and timing. The risk of bias of the largest studies included for the individual outcome were given more weight because of their larger impact on the summary estimate.Imprecision: For these ratings the width of the CI of the summary estimate was considered. Since the CI is affected by inconsistency between the estimates of the individual studies (below), we did not rate down for imprecision when that would have led to double penalization.Inconsistency: Assessments were based on the similarity between point estimates and the extent that CIs overlap.Indirectness: Ratings were primarily based on the QUADAS‐2 assessments of applicability, focusing on patient selection, index test and reference standard.Publication Judgments were made regarding for‐profit interest, or knowledge about studies that were conducted but not published.
The final certainty of evidence for each outcome can be rated as high, moderate, low or very low, where serious or very serious concerns regarding a domain would warrant rating down one or two levels respectively.
In cases where the body of evidence consists of only one study, inconsistency cannot be judged and therefore the certainty of evidence should not be downrated for this domain according to GRADE. ^40^ However, since diagnostic accuracy estimates in general are sensitive to context, in this review we chose to rate down for indirectness when the estimates came from only one study, due to concerns regarding transferring the results from a single diagnostic accuracy study to the general population. Also, when an estimate from a single study was used, concerns regarding the calculation of that estimate were considered in the precision domain.
The review process is summarized in the PRISMA flow diagram in Figure 1. The systematic literature search in databases resulted in 8276 records (titles and abstracts), of which 198 were retrieved in full text. In total, 23 articles met the inclusion criteria, seven of which were identified in a NICE review. ^12^ Two additional studies published before 2014 were identified in the reference lists of the included systematic reviews, both from Levis et al. ^18^ A list of articles excluded after reading them in full‐text, including reasons for exclusion, can be found in Table S2.

Of the 23 included articles, seven were systematic reviews with similar or broader inclusion criteria than ours. ^10^ , ^16^ , ^17^ , ^18^ , ^20^ , ^41^ , ^42^ These were not included in the analyses but reference lists were checked for potential additional studies. The remaining 16 papers were articles that reported on 14 primary DTA studies, including 4970 women in total. No DTA studies were identified that compared different screening alternatives.
A total of 16 DTA articles that reported findings from 14 different studies were included. The diagnostic accuracy of the EPDS was assessed in 12 articles ^7^ , ^43^ , ^44^ , ^45^ , ^46^ , ^47^ , ^48^ , ^49^ , ^50^ , ^51^ , ^52^ , ^53^ and one study respectively assessed the instruments GAD‐2, ^54^ MGMQ, ^55^ PHQ‐9, ^56^ PHQ‐2 ^57^ and Whooley Questions. ^53^ No articles that assessed the diagnostic accuracy of the EPDS‐3A or GAD‐7 were found. Most studies included women recruited during routine antenatal care appointments, and the number of participants per study ranged from 32 ^50^ to 845. ^49^ Data was collected in different trimesters, with some studies reporting data from several trimesters. ^47^ , ^48^ , ^49^ Four studies were conducted in the United States, ^47^ , ^50^ , ^56^ , ^57^ two in Australia, ^43^ , ^55^ and the remaining eight in Europe. Five different clinical interview protocols were used to establish diagnoses according to DSM‐IV, DSM‐5 or ICD‐10 different versions of the Structured Clinical Interview for DSM diagnoses (SCID), ^43^ , ^48^ , ^50^ , ^51^ , ^53^ , ^54^ , ^56^ the Mini International Neuropsychiatric Interview (MINI), ^7^ , ^45^ , ^52^ , ^55^ the Clinical Interview Schedule – revised (CIS‐R), ^46^ the Standardized Psychiatric Interview (SPI), ^44^ and the Composite International Diagnostic Interview (CIDI). ^47^ , ^49^ , ^57^ Further study characteristics are presented in Table S3.
The overall risk of bias regarding diagnostic accuracy was assessed to be low in three studies, ^43^ , ^46^ , ^51^ moderate in four studies (that were presented in 6 articles), ^7^ , ^48^ , ^49^ , ^52^ , ^53^ , ^54^ and high in seven studies, ^44^ , ^45^ , ^47^ , ^50^ , ^55^ , ^56^ , ^57^ see Table S3. Domain specific assessments are presented in Table 1. Several studies used a sample that was selected, at least in part, based on previous symptom‐screening or a case–control approach. ^7^ , ^47^ , ^52^ , ^53^ , ^54^ , ^57^ In some studies, patient selection was neither random, nor consecutive, ^44^ , ^45^ , ^50^ selection procedures were changed along the way, ^55^ or patient exclusions were assessed as inappropriate. ^56^ Others presented only data for optimal cutoff points, ^47^ , ^50^ used lay interviewers for the reference standards, ^47^ , ^49^ , ^57^ or reported unclear ^53^ , ^54^ or unlikely ^47^ , ^50^ , ^55^ blinding procedures. In some cases, the interval between the index test and reference standard exceeded 2 weeks or was unclear, ^7^ , ^52^ , ^53^ , ^56^ and in a couple of studies it was unclear if all participants were included in the analyses. ^44^ , ^47^
Problems with applicability were most commonly related to participants belonging to specific risk groups (eg high risk pregnancies, home visiting programs, or refugees), ^43^ , ^45^ , ^50^ , ^56^ , ^57^ self‐report tests being administered as interviews, ^47^ , ^50^ , ^56^ or how the target condition was defined in the reference standard used. ^44^ , ^46^ , ^47^ In some studies, the assessment of the risk of bias and applicability was complicated by unclear reporting of study procedures, especially regarding blinding and time between the index test and reference standard.
The results from the individual studies can be found in the table of study characteristics (Table S3).
Sufficient data was available for synthesizing the diagnostic accuracy of the EPDS for identification of MDD at cutoffs ≥10 to ≥14. One study ^43^ was excluded from the synthesis due to applicability concerns (Dari speaking refugees in Australia). Forest plots and SROC plots are presented in Figures S1–S5 (Figure a show the forest plots and Figure b the corresponding SROC plots). For the syntheses we also calculated the Youden's index (J), the positive and negative predictive values based on a prevalence of 5%, ^58^ and the positive and negative likelihood ratios, see Table 2.
Assessments of the certainty of evidence of all results included in this review can be found in Tables 3, 4, 5. Apart from the results for identifying cases of MDD using the EPDS presented above, all estimates of sensitivity and specificity were assessed as having very low certainty, mainly due to sparse data, concerns due to risk of bias, and applicability issues.
The most commonly studied screening instrument in this systematic review was the EPDS. While Park and Kim ^20^ chose to synthesize studies using different cutoff points on the EPDS, yielding a pooled sensitivity of 0.81 and a specificity 0.87 in pregnant women, we aimed to look at the sensitivity and specificity of specific cutoffs. When sensitivity and specificity were assessed repeatedly during pregnancy, we aligned with clinical guidelines emphasizing early detection ^12^ and thus selected the earliest timepoint.
In further contrast to previous reviews, ^16^ , ^17^ , ^20^ our focus of interest was only the test accuracy during pregnancy in Western countries. Interestingly, there was a tendency for our summary estimates to be somewhat lower for sensitivity and somewhat higher for specificity compared to the findings from NICE, ^12^ supposedly relating to the difference in our respective inclusion criteria. Despite 9 years having passed since the NICE report was published in 2014, we identified only five additional articles, stemming from three projects, that assessed the diagnostic accuracy of screening instruments for this population. ^7^ , ^52^ , ^53^ , ^54^ , ^55^
In a recent systematic review that included studies of both pregnant and postpartum women from countries worldwide, ^18^ the authors concluded that an EPDS cutoff value of ≥11 would maximize sensitivity and specificity throughout the perinatal period (Youden's J = 0.66–0.73 to be compared with our figure of J = 0.65 at ≥11). Our analyses, however, indicate that a cutoff point of ≥10 would be the most suitable cutoff point for maximizing the combination of sensitivity and specificity in pregnant women (J = 0.71). However, these calculations build on the assumption of sensitivity and specificity being equally important, which might not always be the case. For the purpose of screening in pregnant populations, it can be argued that high sensitivity is more important, given that regularly scheduled antenatal check‐ups will likely provide multiple unobtrusive opportunities to reassess whether a woman whose responses are in the lower range for a positive identification, indeed has mental health issues that warrant further assessment and care. However, at ≥10, the estimated sensitivity of 0.84 was rather low, and with the low positive predictive value of 0.25 (Table 2), most identified cases would be false positives. This suggests that the EPDS might not be perfectly suitable to detect cases of major depression during pregnancy. If sensitivity is prioritized, the Wooley questions might be an alternative. In a recent systematic review, ^16^ this dichotomously scored instrument was found highly sensitive (0.95, 95% CI: 0.81–0.99) but rather low in specificity (0.60, 95% CI: 0.44–0.74) in the perinatal population (4/5 included studies had a sample of pregnant women). It is also important to remember that identification of women scoring above any predefined cutoff point is not likely to be effective unless it is followed by clinical assessment by a mental health specialist and appropriate intervention.
Current NICE guidelines for antenatal identification of women with anxiety disorders are mainly based on data from nonpregnant populations. ^12^ Of the five DTA articles published after the NICE report, ^7^ , ^52^ , ^53^ , ^54^ , ^55^ only three address screening tools for identifying women with anxiety disorders. ^7^ , ^54^ , ^55^ Hence, there is still a great need for instruments aimed at identifying cases of antenatal anxiety to be validated and assessed.
A strength of this review is the use of GRADE for assessing the certainty of the body of evidence for each outcome. Our GRADE assessments show that there are limitations in the body of included evidence that should be taken into account when interpreting the results. Many estimates were based on results from a single study and the certainty of evidence was assessed as very low. Our assessment of the pooled DTA estimates for identification of cases of MDD using the EPDS varied from low to moderate, with several of the analyses revealing an inconsistency between the results reported by the individual studies. All but three studies included in this systematic review were judged to have a moderate or high risk of bias, with participant selection being the most problematic domain. We generally found fewer issues regarding applicability. Focusing on the context of Western countries, we chose to exclude studies from countries outside Europe, USA, Australia and New Zeeland. We also excluded studies with selected populations (eg psychiatric samples and teenagers). As a result, problems with applicability were most often due to participants belonging to specific risk groups (eg high‐risk pregnancies, home visiting programs or refugees). Assessment of the risk of bias and applicability was in some cases complicated by unclear reporting of study procedures, especially regarding blinding and the time that elapsed between when the index test and reference standard were administered.
There is a great need for studies assessing the diagnostic accuracy of instruments that aim to identify cases of anxiety or depression during pregnancy. In this review, most of the included studies evaluated how accurately the EPDS could identify cases of antenatal MDD. A cutoff value of ≥10 was found to be the most suitable for maximizing the combination of sensitivity and specificity. However, our analysis indicates that when used in an antenatal population, the EPDS has a relatively low sensitivity, regardless of cutoff point, while our estimates of specificity are generally higher. Based on our first and broad searches, we also acknowledge the need for primary studies evaluating possible benefits and harms of structured approaches for identification of depression and anxiety during pregnancy. Lastly, we encourage clear reporting on study details such as blinding and timing of the test and the reference standard in primary studies.
In the absence of evidence on effects of structured programs for case identification in these populations, clinical practitioners and stakeholders are left to make decisions of how to identify cases of clinical depression and anxiety during pregnancy based on their clinical judgment and the evidence from other populations. Although current international guidelines differ in some respects, they generally share the notion that early identification, followed by thorough psychosocial and clinical assessment and intervention are desirable. ^12^ , ^13^ , ^14^ , ^15^
At present, assessment using any tool may best serve as a complement embedded in a broader professional interview, which aims to identify those in need of further clinical assessment and intervention, whilst minimizing the potential negative effects which can arise from false positive diagnoses, such as individual stigma and strain on health care resources. It is thus important that healthcare professionals who are expected to effectively identify cases have appropriate training, access to mental health specialists, and enough time to engage in a broad and meaningful conversation with the women in their care. ^59^ , ^60^ Clear pathways need to be set in place to ensure that people who are identified are promptly provided with a thorough diagnostic assessment and effective interventions.
This systematic review was conducted at The Swedish Agency for Health Technology Assessment and Assessment of Social Services (SBU), under the lead of author Monica Hultcrantz. Authors Elisabet Rondung, Pamela Massoudi, Katri Nieminen, and Birgitta Wickberg were subject experts in the project group. All authors contributed to the study design. Martina Lundqvist structured the analytic framework for the review questions. Author Klas Moberg performed the literature search. Elisabet Rondung, Pamela Massoudi, Katri Nieminen, Birgitta Wickberg and Åke Grundberg performed the study selection, and Elisabet Rondung, Pamela Massoudi, Katri Nieminen, Birgitta Wickberg the risk of bias and applicability assessments. Nathalie Peira and Rebecca Silverstein performed the data extraction and tabulation, and Nathalie Peira performed the data analysis. Monica Hultcrantz and Nathalie Peira assessed the certainty of evidence. The first draft of the manuscript was written by Elisabet Rondung and Monica Hultcrantz. Authors Rebecca Silverstein, Nathalie Peira, Monica Hultcrantz and Elisabet Rondung provided figures and tables. Pamela Massoudi, Katri Nieminen, and Birgitta Wickberg made substantial contribution in the planning of the manuscript and commented on previous versions of the manuscript. Rebecca Silverstein reviewed and refined the language. Elisabet Rondung and Monica Hultcrantz finalized the manuscript. All authors read and approved the final manuscript.
This systematic review was conducted at The Swedish Agency for Health Technology Assessment and Assessment of Social Services (SBU) and financed by the Swedish government.
None.