What the design effect and intraclass correlation coefficient (ICC) actually mean, and how to set them correctly for cluster and multi-stage sampling.
People who live in the same village tend to be more alike than two people picked at random from the whole population, similar access to services, similar local conditions, similar exposure to the same programmes. That similarity has a direct, calculable cost in required sample size. Design effect and ICC are how that cost is calculated, and how it is corrected for.
Clustering, in a sampling context, means selecting whole groups of people together rather than selecting individuals independently from the entire population. A village, a school, a health facility catchment area, or a workplace are all examples of clusters. Instead of drawing a simple random sample of individuals scattered across an entire region, a cluster design first selects a sample of clusters, then samples individuals within each selected cluster.
Clustering is common, and often necessary, for a practical reason: a complete individual-level list of the population rarely exists, particularly across a large or dispersed area. A list of villages, schools, or facilities is usually far easier to obtain and work from than a list of every person in the population. Clustering makes fieldwork logistically manageable, since a data collection team can work through one village at a time rather than travelling to scattered, unrelated households across an entire region.
The cost of this convenience is statistical, not logistical, and it is what the rest of this guide explains.
Simple random sampling assumes every observation provides fully independent information. Cluster sampling does not. Once something is known about one person in a village, a little is already known about their neighbour, simply because they share a context, similar services, similar conditions, similar exposure to the same local factors. Each additional person from the same cluster adds less new information than a fully independent person would.
The design effect, DEFF, quantifies exactly how much less. It is a multiplier applied to the sample size that would be needed under simple random sampling, to arrive at the sample size actually needed under a cluster design, for the same level of precision.
This is not the formula being cautious. It reflects a genuine loss of statistical information caused by within-cluster similarity, and ignoring it produces a study that looks adequately sized on paper but is not, in practice, precise enough to support the conclusions drawn from it.
Two things drive DEFF up, sampling more people per cluster, m, and clusters being more internally homogeneous, ρ. DEFF equals 1 exactly when ρ equals 0, meaning clusters are not actually more alike internally than the population at large. In that case, clustering costs nothing, and cluster sampling behaves just like simple random sampling.
The intraclass correlation measures how much of the total variation in an outcome is between clusters versus within them. A ρ near 0 means clusters are essentially interchangeable, most variation is between individuals regardless of which village they belong to. A ρ near 1 means clusters are highly distinct, knowing the village reveals almost everything about the individual.
| Typical ρ | Indicator type |
|---|---|
| 0.01 to 0.03 | Individual attitudes, knowledge, and awareness questions |
| 0.03 to 0.07 | Behavioural indicators, practice, uptake, usage |
| 0.05 to 0.15 | Health and household-level indicators, immunisation, water access, nutrition |
These ranges are starting points, not universal truths. Where possible, use ρ from a prior round of the same survey or a closely comparable study, rather than a generic benchmark.
Ignoring the design effect, whether by overlooking it entirely or leaving it at its default of 1 when the sample is genuinely clustered, understates the true sample size needed. The study still recruits the number calculated under the simple random sample formula, but that number is not actually sufficient to achieve the intended precision under a cluster design. The practical consequences follow directly from that gap.
From a purely statistical efficiency standpoint, the ideal DEFF is 1, meaning no clustering effect at all, since that extracts the maximum amount of information from every person sampled. In practice, a DEFF of exactly 1 is rare in a genuine cluster design, because some degree of within-cluster similarity almost always exists.
There is no single ideal value beyond that principle. What counts as a reasonable DEFF depends on the indicator and the design.
| DEFF range | What it suggests |
|---|---|
| 1.0 to 1.5 | Low clustering effect. Common for attitude or knowledge indicators, or designs with small, tightly controlled cluster sizes. |
| 1.5 to 2.5 | Typical for many community and household surveys. A reasonable planning default in the absence of better information. |
| 2.5 and above | Substantial clustering effect. Common for some health and household indicators, or designs with large cluster sizes. Worth reconsidering the sampling design if the resulting sample size becomes impractical. |
A DEFF that seems unexpectedly high is a signal worth investigating, not simply accepting, since it may indicate that clusters are unusually large, that ρ is genuinely high for that indicator, or that a design change could meaningfully reduce the required sample.
Because DEFF grows with m, people sampled per cluster, but ρ is largely fixed by the indicator itself, the most effective lever available for reducing DEFF is usually the number of people sampled within each cluster, not the indicator.
A few further principles help keep a cluster design efficient.
Two approaches to setting DEFF are available, and they are not mutually exclusive across a research programme, different studies may reasonably use different approaches depending on what evidence exists.
When neither a study-specific DEFF nor a confident ρ estimate is available, 1.5 to 2.0 is a reasonable, defensible planning default for most community-level surveys. State this explicitly as an assumption in the study protocol, and revisit it after a pilot round if the precision of the final estimate matters a great deal to the decision the study is meant to inform.
Design effect is not limited to a single tool in AnalyZ Solutions. It is available, using the same mechanism, across every sample size tool where clustering is a realistic possibility, Cross-sectional, RCT, Case-Control, Change in Proportion, Change in Mean, Longitudinal, Change in Rate, and the single incidence rate estimate. The intraclass correlation and design effect are entered, or calculated, in exactly the same way regardless of which calculation you are running.
Design effect defaults to 1 in every tool across AnalyZ Solutions, meaning no clustering effect assumed, and is always editable, whether or not the chosen sampling approach is cluster-based.
A household water-access survey samples 15 households per village. Based on a prior round, ρ is estimated at 0.06 for this indicator.
Applied to a base simple random sample of 385, this DEFF pushes the requirement to roughly 709 households, a substantial jump that would be easy to miss if DEFF were left at its default of 1.
Calculate sample size considering DEFF and ICC. Free, browser based.
Try it out