Authors: Francesco Innocenti, Math JJM Candel, Frans ES Tan, Gerard JP van Breukelen
Categories: Articles, Cross-national comparisons, informative cluster size, maximin design, optimal design, sample size calculation, two-stage sampling
Source: Statistical Methods in Medical Research
To estimate the mean of a quantitative variable in a hierarchical population, it is logistically convenient to sample in two stages (two-stage sampling), i.e. selecting first clusters, and then individuals from the sampled clusters. Allowing cluster size to vary in the population and to be related to the mean of the outcome variable of interest (informative cluster size), the following competing sampling designs are sampling clusters with probability proportional to cluster size, and then the same number of individuals per cluster; drawing clusters with equal probability, and then the same percentage of individuals per cluster; and selecting clusters with equal probability, and then the same number of individuals per cluster. For each design, optimal sample sizes are derived under a budget constraint. The three optimal two-stage sampling designs are compared, in terms of efficiency, with each other and with simple random sampling of individuals. Sampling clusters with probability proportional to size is recommended. To overcome the dependency of the optimal design on unknown nuisance parameters, maximin designs are derived. The results are illustrated, assuming probability proportional to size sampling of clusters, with the planning of a hypothetical survey to compare adolescent alcohol consumption between France and Italy.
Keywords: Cross-national comparisons, informative cluster size, maximin design, optimal design, sample size calculation, two-stage sampling
For the purpose of estimating the mean or prevalence of an outcome variable (e.g. alcohol consumption or smoking) in a hierarchical population (e.g. students within schools, patients within general practices), or of comparing subpopulations with respect to such a mean or prevalence, it is often convenient, for economic or logistic reasons, to sample in two first, clusters (e.g. schools, general practices) are sampled and then individuals (e.g. students, patients) are drawn from the sampled clusters.^1–3^ Examples of these multi-stage sampling designs include school-based surveys for monitoring substance use among adolescents,^4–6^ and national surveys for estimating the average length of stay for discharges from hospitals,^ 7 ^ or nursing homes.^ 8 ^ The topic of this paper is the efficient design of two-stage sampling (TSS) schemes for estimating the mean of a quantitative outcome variable in a two-level population.
In practice, clusters usually vary in size (e.g. small versus large schools) and then, to estimate the population mean, a sample can be drawn with at least three alternative TSS sampling clusters with probability proportional to cluster size, and then sampling the same number of individuals from each selected cluster (TSS1); sampling clusters with equal probability, and then sampling the same percentage of individuals from each sampled cluster (TSS2); sampling clusters with equal probability, and then sampling the same number of individuals per cluster (TSS3). These three TSS schemes will be considered in this paper and compared with Simple Random Sampling (SRS) of individuals.
Additionally to cluster size variation, further complications arise with informative cluster sizes, that is, when cluster size is related to the outcome of interest.^ 9 ^,^ 10 ^ For instance, cluster size is informative when the amount of alcohol consumed by an adolescent is related to the number of students enrolled in the school, as small schools might provide a more supportive environment,^11–13^ or when the number of patients registered to a general practice affects its efficacy in preventing expensive hospitalisations,^ 14 ^ thus impacting on public expenditure on health per patient. Informative cluster sizes not only can have direct policy implications, such as introducing a limit to school or general practice size, they also have consequences for statistical data analysis and sample size planning. In informative cluster size literature (see the review by Seaman et al., ^ 9 ^ and references therein), the main focus has been on how to handle informative cluster size when the target of inference is the association between the outcome variable and some covariates (e.g. a risk factor). For instance, Seaman et al.^ 9 ^ have discussed several methods to make cluster-specific inferences with Generalized Linear Mixed Models and population-average inferences with Generalized Estimating Equations when cluster size is informative. Innocenti et al., ^ 15 ^ instead, have investigated a different the implications of informative cluster size for unbiased and efficient estimation of a population mean in surveys conducted with the three aforementioned TSS schemes. The present paper is also about mean estimation for these three TSS schemes when cluster size is informative, but focuses instead on sample size planning, and the consequences of informative cluster size for the required sample sizes and budget.
Innocenti et al.’s results^ 15 ^ are the starting point of this paper and therefore summarized here. First, there are two definitions of overall mean in a two-level population, namely the average of all individual outcomes and the average of all cluster-specific means. These two definitions coincide only if cluster sizes are either equal or non-informative. Second, when cluster size is informative, estimation of the mean of all individual outcomes (i.e. the definition used in this paper) is unbiased under TSS1 with the unweighted average of cluster means, and asymptotically unbiased under TSS2 and TSS3 with the average of cluster means weighted by cluster size. In contrast, when cluster size is non-informative, the unweighted average of cluster means is unbiased for all sampling schemes, but optimally efficient for TSS1 and TSS3 only. Third, under the constraint of a fixed total sample size, SRS is more efficient than any TSS scheme, TSS3 is the least efficient TSS scheme, and TSS1 is the most efficient for many cluster size distributions. Indeed, when cluster size is informative, the relative efficiency of these sampling schemes depends on some features of the cluster size distribution in the population, such as the coefficient of variation, the skewness, and the kurtosis. However, when cluster size is non-informative, TSS1 and TSS3 are equally efficient and outperform TSS2. Fourth, the two inferential paradigms in survey sampling, namely the model-based^ 3 ^ and the design-based approach,^ 1 ^,^ 2 ^ give similar results in terms of unbiased and efficient estimation of the average of all individual outcomes with the three aforementioned TSS schemes, at least if the model assumptions are met. Furthermore, sample size planning and sampling schemes comparisons, which are the topics of this paper, are much more feasible with the assumption of a model for the outcome variable of interest.^ 15 ^ For these two reasons, the model-based approach is adopted here.
This work extends the results of Innocenti et al.^ 15 ^ in the following ways. First, for each of the three aforementioned TSS schemes, the optimal design is derived. Here, the optimal design is defined as that design (i.e. number of clusters and number of individuals per cluster) that minimizes the sampling variance of the population mean estimator subject to a cost constraint. Second, the three optimal TSS schemes are compared with SRS and with each other under the constraint of a fixed budget. Third, to take care of uncertainty with respect to model parameters and distributional features of cluster size, as a practical alternative, maximin designs are derived. Fourth, sample size calculations for making comparisons between populations are derived and illustrated.
This paper is structured as follows. In section 2, the assumptions of this paper are presented, as well as the sampling schemes and the corresponding mean estimators. Furthermore, the findings of a simulation study to assess the accuracy of some results in Innocenti et al.^ 15 ^ that are relevant to the present paper are summarized. In section 3, the optimal design for each TSS scheme is derived, and these optimal TSS designs are compared with each other and with SRS for a fixed budget. Furthermore, the consequences of ignoring informative cluster size at the design phase of a study are investigated. Section 4 deals with the maximin approach, that is, a strategy to solve the dependency of the optimal design on unknown nuisance parameters. Section 5 provides a procedure for computing sample sizes for surveys aimed to make cross-population comparisons, and the procedure is illustrated in planning a survey for comparing the average alcohol consumption among adolescents in France and Italy. Section 6 offers some final remarks. The mathematical derivations of the results, the description of the simulation study discussed in section 2, and additional figures and tables can be found in the Supplementary Material 1 (S.M.1). The Supplementary Material 2 (S.M.2) provides the R^ 16 ^ code of the simulation study and other R codes to apply some of the mathematical results of this paper.
The results of Innocenti et al.^ 15 ^ and this paper are based on the following assumptions (the notation used in the main text is summarized in the Appendix).
**Assumption ** The population is composed of K clusters and each cluster j contains Nj individuals, that is, in the population clusters vary in size ( Nj ). The population size is Npop=∑j=1KNj .
**Assumption ** Sampling is either SRS of individuals in one stage, or else TSS. In TSS, we first sample k clusters, and then sample n or nj individuals per selected cluster j . In case of TSS, the population is very large relative to the sample size at each design level (i.e. kK→0 and n¯θN→0 , where n¯=∑j=1knjk is the average sample size per sampled cluster, and θN=NpopK is the population mean of cluster size). In case of SRS, Npop is very large relative to m , the number of individuals sampled (i.e. mNpop→0 ).
**Assumption ** The outcome variable Yij is quantitative (e.g. alcohol consumption) and measured at the individual level. Further, Yij shows variation at the cluster level as well as at the individual level. Therefore, sampling error occurs at each design level. This is taken into account by assuming the following two-level random intercept model for the outcome of the i -th individual from the j -th cluster^ 3 ^,^ 17 ^
where εij∼N0,σε2 , and cluster effect uj and individual effect εij are unrelated (i.e. uj⊥εij ). The distribution of uj will be defined in the next assumption.
**Assumption ** Cluster effect uj is linearly related to cluster size Nj , that is, uj=α0+α1Nj+νj=α1Nj−θN+νj , where α0=−α1θN for model identifiability, νj∼N0,σν2 , and νj is the component of cluster effect uj that does not depend on cluster size (i.e. νj⊥Nj ). Thus, the conditional distribution of uj given Nj is uj|Nj∼Nα1Nj−θN,σν2 .
Innocenti et al.^ 15 ^ show that β0 in model (1) is the average of all cluster-specific means in the population, and differs from the average of all individual outcomes in the population μ , unless cluster size is non-informative ( α1=0 ) or constant across clusters, as can be seen from the following expression
where θN , τN=σNθN , and σN2 are, respectively, the population mean, the coefficient of variation, and the variance of cluster size. The distinction between β0 and μ comes from considering the distribution of cluster effect uj over either the population of clusters (which yields β0 ) or the population of individuals (which yields μ ).^ 15 ^ This paper focuses on μ .
With the aim of estimating μ , the three aforementioned TSS schemes are studied in this paper. For each of these TSS schemes and SRS, Table 1 summarizes the sampling procedure (i.e. sample size and inclusion probability per design stage) and the required knowledge before sampling. Furthermore, Table 1 shows the population mean estimator μ^ and the sampling variance Vμ^ for each sampling scheme. Denote by ρuN=Euj−EujNj−ENjσuσN=EujNj−θNσuσN the correlation between uj and Nj , where Euj=0 and Vuj=σu2=σν2+α12σN2 , and by ψ=ρuN21−ρuN2 the degree of informativeness of cluster size. From Table 1, note that n¯=∑j=1knjk=n for TSS1 and TSS3, while n¯=∑j=1knjk=p∑j=1kNjk=pN¯ for TSS2, where N¯ is the average population size of the k sampled clusters (not to be confused with n¯ , that is, the average sample size of the sampled clusters). Furthermore, for TSS2 n=En¯=pEN¯=pθN . The sampling variances in Table 1 are functions of the total unexplained outcome variance σy2=σν2+σε2 , the intraclass correlation coefficient ρ=σν2σy2∈0,1 , the sample sizes ( k , n ), the parameter ψ , and some features of the cluster size distribution in the the coefficient of variation τN , the skewness ζN , and (for TSS2 and TSS3 only) the kurtosis ηN . When cluster size is non-informative ( ψ=0 ), Vμ^ depends only on σy2 , ρ , k , n , and (for TSS2 and TSS3 only) τN . The estimators μ^ associated with SRS and TSS1 are unbiased, and their sampling variances Vμ^ are exact expressions.^ 15 ^
20.74 ).
For any sampling scheme, the precision of the estimator μ^ , and thus also the width of a confidence interval for μ and the statistical power for testing a hypothesis on μ , depends on the number of clusters and on the sample size per cluster (Table 1). This raises the question of the best combination of sample sizes at each design stage (i.e. sampling many clusters versus sampling many individuals per cluster). Define the optimal design as that design (i.e. number of clusters and number of individuals per cluster), which minimizes Vμ^ subject to a cost constraint, given that time and budget are limited in practice. For TSS, the cost constraint is assumed to be C=kc2+c1n , where C is the budget for sampling and measuring (excluding costs for constructing the sampling frame and other costs not related to sample size). From now on C is called the research budget. Furthermore, c2 is the average cost for sampling a cluster, c1 is the average cost for sampling an individual from a sampled cluster, and c2+c1n is the cost per cluster including the costs for sampling n individuals from that cluster (recall that for TSS2 n=pθN ). For SRS, the cost constraint is C=c0+csrsm , where m is the number of individuals to sample, csrs is the average cost for sampling an individual directly from the population, and c0 represents the extra-cost due to constructing the sampling frame for a SRS compared with the sampling frame for a TSS.
For each TSS scheme, the optimal design (i.e. the optimal sample sizes k* and n* ) for estimating μ and the optimal variance Vμ^* (i.e. Vμ^ under the optimal design) are given in Table 2 (for proofs, see section 2.2 of S.M.1). For TSS2, one can obtain the optimal proportion of individuals to sample per cluster p* from the optimal n* , by dividing n* as given in Table 2 (TSS2 row) by θN . The optimal TSS2 and TSS3 designs depend on two approximations of Vμ^ : the first-order Taylor approximation mentioned in section 2 and evaluated in S.M.1 (section 1), which underlies the equations in Table 1, and an approximation based on large k (i.e. k such that τN2k≈0 , k−1k≈1 , and k−3k−1≈1 ) to simplify the expressions in Table 1. These two approximations give the following equations (for details, see section 2.1 of S.M.1)
and
where for TSS2, n=pθN . Recall from section 2 that, for TSS2 and TSS3, k must be large anyway, because the estimators μ^TSS2 and μ^TSS3 given in Table 1 are only asymptotically unbiased. As a special case, ψ=0 gives the optimal design and optimal variance for non-informative cluster size (for which case β0=μ ), which under TSS1 coincide with the equations available for cluster randomized trials (for instance, see Moerbeek et al. ^ 19 ^). There is no such equivalence under TSS2 due to sample size variation between clusters, and under TSS3 due to weighting cluster means by cluster size if informative cluster size is assumed in the design phase. Indeed, under non-informative cluster size, no weighting is needed under TSS3,^ 15 ^and then the optimal design equations for TSS1 apply to TSS3 as well.
Note from Table 2 that the optimal number of clusters k* and the optimal number of individuals per cluster n* are inversely related, and that n* is an increasing function of the cluster-to-individual cost ratio cr=c2c1>1 and a decreasing function of ρ and ψ . These relations between the optimal design and cr , ρ , and ψ hold, under TSS1, for ζN>τN−1τN , and always under TSS2 and TSS3 (for proof, see section 2.1 of S.M.1). The condition ζN>τN−1τN is met by all the distributions in Tables S.2 and S.7 (S.M.1). Hence, this condition is assumed to be satisfied when considering results for TSS1 in the sequel.
The optimal number of individuals per cluster n* for TSS1 and TSS3 is plotted in Figure 1, for two real-life cluster size the general practice list size distribution in England, and the public high school size distribution in Italy (both distributions are shown in Figure S.1, S.M.1). The behaviour of n* for other cluster size distributions is shown in Figures S.2 and S.4 (S.M.1) for TSS1 and TSS3, respectively, and in Figure S.3 (S.M.1) for TSS2. In most scenarios in Figure 1 and Figures S.2–S.4 (S.M.1), the difference between n* for ψ=0.35 (i.e. ρuN=±0.51 ) and n* for ψ=0 (i.e. ρuN=0 ) is small, which means that the ratio of Vμ^ under the design assuming ψ=0.35 to Vμ^ under the design assuming ψ=0 , when the true ψ=0.35 , is close to 1. So, the optimal designs in Table 2 are quite robust against misspecification of ψ , in the sense of being efficient relative to the optimal design for the true ψ and given a fixed research budget C . However, ignoring informativeness can lead to serious underestimation of the sampling variance of the mean estimator, and thereby also of the budget needed, as will be seen below. Further, the optimal design depends not only on ψ , but also on ρ and the cluster size distribution ( τN , ζN , ηN ). That dependence will be addressed in section 4.
Figure 1. Optimal number of individuals per cluster n* under TSS1 (left column) and TSS3 (right column), as a function of ρ, for different values of c
rand ψ (curves), and different cluster size distributions (rows). The cluster size distributions are shown in Figure S.1 (S.M.1). Note that ψ = 0.35 corresponds to ρuN=±0.51.
An example will now show that (a) given a study budget, the optimal design is robust against misspecification of cluster size informativeness, but (b) the budget needed is very sensitive to misspecification. Suppose we plan a survey to estimate μ in the population of all patients of all general practices in England. The parameters of the general practice patient list size distribution are τN=0.633 , ζN=2.12 , and ηN=14.549 (Table S.2, S.M.1). Furthermore, suppose that ρ=0.05 , cr=10 , and C/c1=1000 . The optimal TSS1 samples n*=10.74 individuals and k*=48.22 clusters assuming ψ=1/3 , and n*=13.78 and k*=42.04 assuming ψ=0 (see Table 2, TSS1 row). If the true ψ=1/3 , Vμ^=σy2×0.00354 for the design correctly assuming ψ=1/3 , and Vμ^=σy2×0.00360 for the design incorrectly assuming ψ=0 (see variance equation in Table 1, TSS1 row), giving a variance ratio 0.00354/0.00360 = 0.983 . Additional results for TSS1, TSS2, and TSS3 are given in Table S.8 (S.M.1), which shows that even in some more extreme cases (e.g. ψ=1 , i.e. ρuN=±0.707 ) the variance ratio still exceeds 0.8. The example given here and those in Table S.8 (S.M.1) show that the optimal designs in Table 2 are quite robust against misspecification of ψ , in the sense of being efficient relative to the optimal design for the true ψ and given a fixed research budget C .
However, ignoring informativeness can lead to serious underestimation of the budget needed. Suppose one wants to test the null hypothesis H0 that μ=μ0 against the alternative hypothesis H1 that μ≠μ0 . The budget that guarantees the desired power level 1−γ for the chosen type I error rate α , is then obtained by equating Vμ^* in Table 2 with μ−μ0z1−γ+z1−α22 , where zq is the q th percentile of the standard normal distribution. This gives C=gρ,ψz1−γ+z1−α22d02 , where gρ,ψ is the numerator of Vμ^* in Table 2 excluding σy2 , and d0=μ−μ0σy is the standardized difference between true mean and mean according to H0 . Since gρ,ψ is an increasing function of c1 , c2 , and ψ , the required budget C for the desired power level also increases with c1 , c2 , and ψ . Likewise, C increases with ρ , at least up to ρ=0.5 (for proofs, see section 2.2 in S.M.1). The required budget C to detect a standardized difference of medium size ( d0=0.5 ), with 90% power and two-tailed α=0.05 , is plotted in Figure 2 for TSS1 and TSS3, as function of ψ , for the general practice list size distribution in England and the public high school size distribution in Italy, and assuming c1=10 . As can be seen in Figure 2, the research budget C is not robust against misspecification of ψ . For example, the required budget C for the optimal TSS1, assuming the English general practice list size distribution, cr=30 , c1=10 , and ρ=0.10 (Figure 2, left column, first row), is underestimated by 29% if one incorrectly assumes ψ=0 when the true ψ=0.35 . The required budget C is also shown, for other cluster size distributions, in Figures S.5 and S.7 (S.M.1) for TSS1 and TSS3, respectively, and in Figure S.6 (S.M.1) for TSS2. These figures show that C increases with ρ , c2 , and ψ , and that the impact of the cluster size distribution on C becomes more relevant as ψ increases. Hence, ignoring informative cluster size at the design phase of the survey can lead to underestimating the required budget for the chosen effect size and desired power level. Finally, for the desired power level, the required budget is smallest with the optimal TSS1, and largest with the optimal TSS3.
Figure 2. Budget C needed for the optimal design to detect a standardized difference between hypothesized and true population mean of medium size (d
0= 0.5), with 90% power using a two-tailed test with α = 0.05, as a function of ψ, for different values of ρ and cr(curves) with c1= 10, different sampling schemes (columns), and different cluster size distributions (rows). The cluster size distributions are shown in Figure S.1 (S.M.1). Note that ψ∈[0,1.3] corresponds to ρuN∈[−0.75,+0.75].
We now compare the efficiency of the optimal designs in Table 2 with each other and with SRS, under the constraint of a fixed research budget. The relative efficiency ( RE ) of the optimal designs for two sampling schemes is defined as the ratio of their optimal variances Vμ^* in Table 2, more specifically, RE D1 vs D2=VD2μ^VD1μ^ . These RE s are shown in Table 3 (for proofs, see section 2.3, S.M.1), which also gives the sufficient (but not necessary) conditions under which each RE is smaller than one.
The RE of a TSS scheme compared with SRS (Table 3, first three rows) is composed of three ratios. The first ratio is a function of ρ , ψ , τN , ζN , ηN , and cr , and is always smaller than one for ψ=0 , and also for ψ≠0 at least under the conditions for ζN given in the rightmost column of Table 3 (for proofs, see section 2.3, S.M.1). The other two components of RETSS vs SRS are the ratio csrsc1 , for the costs per individual in SRS relative to TSS, and the budget ratio CC−c0 . Since sampling an individual directly from the population will be more expensive than sampling an individual after having sampled the cluster to which he/she belongs (i.e. csrs>c1 ), and constructing the sampling frame for a SRS has extra-costs compared with constructing the sampling frame for a TSS (i.e. c0>0 ), the ratios csrsc1 and CC−c0 will always be at least one and often larger than one. As a result, the RE can become larger than one, implying that SRS can be less efficient than TSS under the constraint of a fixed budget.
The RE s of the optimal TSS1 and TSS3 versus SRS are shown in Figure 3, for the general practice list size distribution in England and the public high school size distribution in Italy, and assuming csrsc1=CC−c0=1 (note that values greater than 1 give a higher RE of TSS versus SRS). Further, Figures S.8–S.10 (S.M.1) show the RE s of the optimal TSS1, TSS2, and TSS3 versus SRS for other cluster size distributions. For ψ=0 , the RE of any optimal TSS versus SRS is a decreasing function of (i) cr (Table 3), (ii) ρ (at least for ρ≤0.5 , see Figure 3, and Figures S.8–S.10 in S.M.1), and (iii), only for TSS2 and TSS3, τN (Table 3). For ψ≠0 , the patterns remain almost the same as before and the RE s also do not seem to vary much across cluster size distributions (Figure 3, and Figures S.8–S.10 in S.M.1).
Figure 3. Relative efficiency of the optimal TSS1 versus SRS (left column), and of the optimal TSS3 versus SRS (right column), for a given research budget C and assuming (c
srs/c1) = (C/(C-c0)) = 1 (values greater than 1 give a higher RE of TSS versus SRS), as a function of ρ, for different values of crand ψ (curves), and different cluster size distributions (rows). The cluster size distributions are shown in Figure S.1 (S.M.1). Note that ψ = 0.35 corresponds to ρuN= ±0.51.
The RE s of the three TSS schemes compared with each other (Table 3, last three rows) are functions of ρ , cr , ψ , τN , ζN , and ηN . The optimal TSS2 is more efficient than the optimal TSS3 since RETSS3 vs TSS2<1 (unless τN=0 , Table 3, or ρuN≈±1 ,^ 15 ^ since in both cases RETSS3 vs TSS2=1 ). The RE s of TSS2 and TSS3 versus TSS1 are smaller than one, and so the optimal TSS1 is the most efficient TSS scheme, at least for cluster size distributions satisfying the conditions in Table 3 (rightmost column), such as all distributions in Table S.7 (S.M.1). For other cluster size distributions, one must compute the RE for that particular distribution to see whether RE<1 . However, for ψ=0 , the RE s in the last three rows of Table 3 are all smaller than one for any cluster size distribution, making TSS1 the most efficient TSS scheme, followed by TSS2. Note that this only holds if informative cluster size ( ψ≠0 ) is assumed at the design stage, such that in TSS3 cluster means are weighted by cluster size to estimate μ (Table 1). If non-informative cluster size ( ψ=0 ) is assumed already in the design stage, then no weighting is needed for TSS3,^ 15 ^ and TSS3 then is as efficient as TSS1.
The RE of the optimal TSS2 and TSS3 versus the optimal TSS1 are shown in Figure 4, for the general practice list size distribution in England and the public high school size distribution in Italy, and in Figures S.11–S.12 (S.M.1) for other four cluster size distributions. For ψ=0 , these reduce to RETSS2 vs TSS1=crρ+1−ρ2crρτN2+1+1−ρ2 and RETSS3 vs TSS1=1τN2+1 , which are both decreasing functions of τN , but RETSS2 vs TSS1 also decreases as ρ and/or cr increases. For ψ≠0 , the patterns are the same as before with two major differences. First, both RE s decrease as ηN increases (Table 3). Second, for ψ=0.35 , both RE s differ at most 6% from their values at ψ=0 (Figure 4, and Figures S.11–S.12 in S.M.1), except for the English general practice (GP) list size distribution that, having an extreme kurtosis (i.e. ηN=14.55 ), shows a drop in RE (compared with the case ψ=0 ) larger than 20% . Note that TSS1 is the most efficient design in Figure 4 and Figures S.11–S.12 (S.M.1).
Figure 4. Relative efficiency of the optimal TSS2 versus the optimal TSS1 (left column), and of the optimal TSS3 versus the optimal TSS1 (right column), for a given research budget, as a function of ρ, for different values of c
rand ψ (curves), and different cluster size distributions (rows). The cluster size distributions are shown in Figure S.1 (S.M.1). Note that ψ = 0.35 corresponds to ρuN= ±0.51.
In section 3.2 it has been noticed that the optimal designs in Table 2 require a priori knowledge of some nuisance parameters (i.e. ρ , τN , ζN , ηN , and ψ ). This is known as the local optimality problem in optimal design literature.^ 20 ^,^ 21 ^ Basically, this means that the optimal design is optimal only for certain values of these nuisance parameters. In this paper, the local optimality problem is solved taking a maximin approach.^20–22^ This approach has been applied in several contexts, such as longitudinal studies,^23–25^ fMRI experiments,^ 26 ^ cluster randomized and multicentre trials,^27–29^ cost-effectiveness studies,^ 30 ^,^ 31 ^ life-event studies,^ 32 ^ test construction,^ 33 ^ and biological and pharmacological studies.^34–37^ The maximin approach is composed of the following
The resulting design is called the maximin design, which is the optimal design for the worst-case scenario, as defined by that set of parameter values chosen in step 3. The advantage of the maximin design is that it not only maximizes the efficiency and the power in the worst-case scenario, but it also guarantees at least that same efficiency and power level for all the other parameter values within the parameter space. Indeed, Vμ^ is smaller and the power for hypothesis testing on μ is larger, for all other parameter values than for the worst-case values chosen in step 3, given any fixed sample size (i.e. k and n ).
Following the four steps above, we now explain how to find the maximin design for each sampling scheme. The optimal design for TSS1 depends on ρ , τN , ζN , and ψ. However, to draw a TSS1 sample we need to know the cluster size distribution in the population anyway, which means that τN and ζN are also known before sampling. Thus, for TSS1, only ρ and ψ are unknown. The maximin design for TSS1 is obtained by plugging into the optimal sample sizes equations (Table 2, TSS1 row) the largest realistic values of ρ and ψ (for proofs, see section 3.1 in S.M.1). Unlike for TSS1, when sampling with TSS2 or TSS3 the researcher needs no prior knowledge of the whole cluster size distribution. Indeed, if such information is available, sampling with TSS1 is a better choice (Table 3). The maximin design for TSS2 and TSS3 is obtained by plugging into the optimal design equations (Table 2) the upper-bounds of the ranges for ρ , ζN, ηN , and ψ , and the worst-case value of τN (for proofs, see section 3.1 in S.M.1). The latter value can be obtained with an R function given in S.M.2 (section 2), which searches numerically for the value of τN that maximizes Vμ^ (i.e. equations (3) and (4)) within its range of plausible values, given the worst-case values for ρ , ζN, ηN , and ψ . For several upper-bounds for ρ , ζN , ηN , and ψ , a numerical evaluation was performed and this always gave τN=1 as worst-case value of τN within the range 0,1 (for details, see section 3.2 in S.M.1).
To be on the safe side in sample size planning, one can assume for ρ the parameter range 0, 0.10 in health and medical research,^ 38 ^,^ 39 ^ and 0, 0.25 in educational research.^ 18 ^,^ 40 ^ Lacking empirical evidence for ψ or ρuN , we propose ψ∈0, 0.35 , which corresponds to ρuN∈−0.51, 0.51 . The range τN∈0,1 can be justified by considering Table S.7 (S.M.1), and the extreme cases of an exponential cluster size distribution, for which τN=1 , and of a binary distribution with half of all clusters having size 2 and the other half having size 2θN−2 , for which τN≈1 . Finally, for ζN and ηN, the ranges ζN∈0.5, 2 and ηN∈3, 15 can be chosen based on Table S.7 (S.M.1). Since Vμ^ under TSS2 and TSS3 is an increasing function of ζN and ηN (at least if τN≤1 , which will usually hold), assuming positive skewness and positive excess kurtosis (i.e. ηN−3 > 0 ) is a safe choice.
As mentioned in section 3.1, the optimal design for TSS2 and TSS3 depends on two the first-order Taylor series approximation used to derive Vμ^ for TSS2 and TSS3 in Table 1, and the large k approximation to simplify the equations in Table 1 into equations (3) and (4). Since the maximin design is the optimal design for the worst-case scenario, the same approximations also underlie the maximin design. Based on the simulation study and the numerical evaluation discussed in S.M.1 (sections 1 and 3.3), it turned out that each approximation induces a bias of at most 5% in the V(μ^) used to derive the optimal/maximin design if the optimal/maximin kMD ≥20 , or, for ζN≥2 and ηN≥9 , kMD≥100 . Since V(μ^)∝1k , a simple solution is to increase the maximin kMD with 10% to ensure sufficient power at the expense of a 10% higher budget C . However, if the maximin kMD<20 or (for ζN≥2 and ηN≥9 ) kMD<100 both approximations are biased by more than 5% . A solution is to first increase C such that maximin kMD ≥20 or (for ζN≥2 and ηN≥9 ) kMD≥100 , and then further increase C by 10% .
The results of the previous sections allow to efficiently plan a survey not only for estimating a mean, but also for comparing different populations, if the samples are independent. An example of such a study is the ESPAD study,^ 6 ^ which compares substance use among 15–16-year-old students across 35 European countries. For a fixed separate budget per population, the optimal design per population is given in Table 2 and the maximin design in section 4. However, the design can be further optimized by constraining the total budget (i.e. the sum of the separate budgets) instead of each separate budget and finding the optimal (or maximin) budget split between populations (for details, see section 4 of S.M.1). For the case of comparing two populations, this optimization was formalized into a procedure to compute maximin sample sizes per population and the maximin budget split between populations, obtained by extending Van Breukelen and Candel^ 28 ^ to TSS1 with informative cluster sizes and different cluster size distributions per population. This procedure for comparing two populations is implemented in an R code given in section 4 in S.M.2. To use this program, the researcher needs to specify c1 and c2 per population, τN and ζN of the cluster size distribution of each population, the largest plausible values for ρ and ψ , a range for the ratio of the outcome standard deviations ( σy ) between the two populations, the smallest difference μF−μI that is worthwhile being detected, the maximum sum of outcome variances in both populations Vmax , the power level 1−γ , and the type I error rate α . The R code (S.M.2, section 4) returns the maximin sample sizes per population and the maximin budget split. The steps of this procedure are given in S.M.1 (section 4). This procedure is presented only for TSS1, because it is the most efficient sampling scheme for many cluster size distributions.
Let us demonstrate the procedure with the following example. Suppose that we want to plan a survey to estimate and compare the average alcohol consumption among high school students between France and Italy. Similar to the ESPAD study,^ 6 ^ alcohol consumption Yij is measured as the average volume of ethanol (in centilitres) consumed on the last drinking day. Based on adolescent health literature, at the design stage, school size (i.e. total number of students) can be assumed to be informative, that is, related to alcohol consumption. Indeed, it has been found that school size and school connectedness, broadly defined as the degree of belonging at school, are inversely related,^ 12 ^,^ 13 ^ as well as school connectedness and alcohol use.^ 11 ^ TSS1 is the most efficient two-stage sampling scheme for both high school size distributions (this can be verified by checking the conditions in the rightmost column of Table 3, with the numbers given in the second and third row of Table S.7 of S.M.1), and so it is chosen for both populations. Suppose that we want to test the null hypothesis H0 that μF=μI against the alternative hypothesis H1 that μF≠μI , where μF and μI are the population means of alcohol consumption in France and Italy, respectively. Since the French and the Italian samples are independent, we can apply the procedure above to determine how many schools and how many students per school one has to sample per country, and how to split the total budget between countries.
The results are shown in Table 4 for four different cost scenarios. Two largest plausible values are assumed for ρ and ψ , respectively, ρmax=0.1,0.2 and ψmax=0,0.35 . This combination of costs and model parameters (Table 4, first six columns) gives a total of 4 × 2×2 = 16 scenarios, each corresponding to a row in Table 4. The seventh column in Table 4 gives the maximin budget split CFCI (i.e. the ratio of the budget for France, CF , to that for Italy, CI ), and from the eighth to the eleventh column the maximin sample sizes per country are shown. Finally, the rightmost column of Table 4 shows the total budget required to detect a standardized difference of medium size ( d=μF−μIVmax2=0.5 ), with 90% power using a two-tailed test with α=0.05 . From Table 4, it can be seen that the maximin nMD per country is an increasing function of cr , a decreasing function of ρ and ψ , and is inversely related to the maximin kMD . Furthermore, the maximin budget split CFCI=1 only for ψ=0 and homogeneous costs ( c1,F=c1,I and c2,F=c2,I ). In all other scenarios CFCI<1 , meaning that more budget is allocated to the Italian sample than to the French sample. Given that ρmax and ψmax are the same for both countries, CFCI<1 because (i) sampling a student is more expensive in Italy than in France ( c1,F<c1,I ), or (ii) sampling a school is more expensive in Italy than in France ( c2,F<c2,I ), or (iii) only for ψ=0.35 , the school size distribution in Italy is such that τNζN−τN is larger than in France (see Tables S.7 and S.9 of S.M.1). Finally, the total budget C required for the desired power is larger for ψ=0.35 than for ψ=0 (Table 4, rightmost column), suggesting that ignoring informative cluster size at the design stage has the consequence of determining a research budget which is too low for the desired power level. Specifically, informative cluster size requires C to increase with 23–32% depending on the scenario (the larger ρ and/or c2,I , the larger this relative increase, see Table 4, rightmost column).
To estimate an overall mean, two-stage sampling is a logistically convenient way to collect data from a multilevel population. In practice, resources (time and money) for sampling are limited. Thus, this paper presents optimal sample sizes per design stage that either maximize the precision of the population mean estimate for the available research budget, or minimize the research budget for the required precision for estimation. Such optimal designs were derived for three TSS sampling clusters with probability proportional to cluster size, and then the same number of individuals per cluster (TSS1); sampling clusters with equal probability, and then the same percentage of individuals per cluster (TSS2); and sampling clusters with equal probability, and then the same number of individuals per cluster (TSS3).
The optimal sample size equations were derived allowing cluster size to be informative, that is, to be related to the outcome variable of interest. It turned out that the optimal designs given in Table 2 are quite robust against misspecification of the degree of informativeness of cluster size ψ . As shown in section 3.2 and in Table S.8 (S.M.1), the relative efficiency of the optimal TSS1 assuming ψ=0 (i.e. non-informative cluster size) versus the optimal TSS1 assuming ψ>0 (i.e. informative cluster size), when the true ψ>0 was close to one. Nevertheless, ignoring informative cluster size is risky for two reasons. First, assuming ψ=0 one would be tempted to combine the unweighted average of cluster means with TSS3, because this strategy (i.e. combination of sampling scheme and estimator) is unbiased and efficient for ψ=0 . However, this strategy is biased and inefficient if the true ψ>0 . Thus, assuming ψ>0 is always prudent because it leads to combining the unweighted average of cluster means with TSS1, that is, choosing a strategy which is unbiased and highly efficient both for informative and non-informative cluster size. Second, assuming ψ=0 can lead to underestimating the research budget for the desired power level, because the research budget is an increasing function of ψ (see Figure 2, and Table 4, rightmost column). This applies not only to TSS1, but also if, because of practical constraints, one has to choose TSS2 or TSS3 as a sampling scheme. For these two reasons, we recommend assuming ψ>0 at the design stage of the survey.
The optimal designs of the three TSS schemes were compared with each other and with SRS under the constraint of a fixed budget. In contrast to what was the case under the constraint of a fixed total sample size,^ 15 ^ SRS can be less efficient than TSS, because it is more expensive to construct a sampling frame of all individuals in the population than of those from the selected clusters only ( c0>0 ), and because it is more costly to sample and measure geographically dispersed individuals than those that are grouped in a natural cluster (e.g. school, general practice) ( csrs>c1 ). Under informative cluster size, the optimal TSS1 was shown to be the most efficient sampling scheme for many cluster size distributions, followed by TSS2, and then TSS3. We thus recommend TSS1, provided all cluster sizes are known before sampling.
The optimal design depends on several unknown parameters (i.e. the intraclass correlation ρ , the informativeness parameter ψ , and the cluster size distribution’s coefficient of variation τN , skewness ζN , and kurtosis ηN ). To address this issue the maximin approach was proposed. For the considered TSS schemes, this strategy consists of plugging the worst-case value for each unknown parameter into the optimal design equations in Table 2. For ρ , ψ , and ηN , the largest plausible value is the worst-case value. If all plausible values for τN≤1 , then the largest plausible value for ζN is also the worst-case value. The worst-case value for τN can be obtained with an R code, given in S.M.2 (section 2). However, a numerical evaluation showed that if the largest plausible value for τN is 1, this is the worst-case value for τN . The R code also returns the worst-case value for ζN in the rather unrealistic case that some plausible values for τN>1 . The maximin approach has the advantages of being relatively simple to implement, and being robust against misspecification of the unknown parameters by maximizing the minimum efficiency over the ranges of their plausible values. An alternative approach is to obtain estimates of the nuisance parameters from a pilot study and use these in the sample size calculation. However, ρ risks to be underestimated (and thus the main survey to be under-powered), unless the pilot study samples a large number of clusters and of individuals per cluster, which means a sizeable portion of the limited resources for the main survey has to be devoted to the pilot study.^ 41 ^ The underestimation is likely to be even more severe for skewness and kurtosis, given that their traditional estimators are biased downwards unless the sample size is large or (only for the skewness) cluster size is normally distributed.^ 42 ^ For all these reasons, we recommend the maximin approach. Relatedly, to improve the planning of future surveys, empirical studies should report values of these nuisance parameters like in Table S.7 (S.M.1).
The results of this paper also allow to efficiently plan surveys for comparing different populations, provided the samples are independent. For TSS1, a procedure to derive maximin sample sizes and maximin budget split between populations was obtained by extending Van Breukelen and Candel’s^ 28 ^ findings to informative cluster size. Analogous extensions for TSS2 and TSS3 could be explored. However, when either cluster size is non-informative ( ψ=0 ), or the cluster size distribution as well as the informativeness parameter α1 is the same in both populations (e.g. treated and control groups in a cluster randomized trial), we have that μF−μI=β0,F−β0,I (see equation (2)) and then the equations given in this paper reduce to simpler expressions as also derived by Van Breukelen and Candel^ 28 ^ (i.e. those for TSS1 with ψ=0 ).
Finally, in this paper the model-based approach to survey sampling was adopted. However, the results of this paper are valid also under the design-based approach, provided model (1) and assumption 4 hold and inference is then based on the sampling scheme.^ 15 ^ Future research could extend the results of this paper by considering dichotomous outcomes, three-level populations, and by deriving the optimal design for longitudinal studies to monitor trends.
Francesco Innocenti https://orcid.org/0000-0001-6113-8992
Math JJM Candel https://orcid.org/0000-0002-2229-1131
Gerard JP van Breukelen https://orcid.org/0000-0003-0949-0272