Sampling/ Randomized Controlled Trial
Sample Size Estimation, Measuring Change

How to Calculate Sample Size for a Randomized Controlled Trial

A randomized controlled trial compares outcomes between a treatment arm and a control arm assigned at random. This guide explains the underlying statistics first, then shows how to apply them using AnalyZ Solutions.

7 min read Sample Size Estimation Intermediate

A randomized controlled trial assigns participants to a treatment arm and a control arm at random, then compares outcomes between them. Random assignment is what allows any observed difference to be attributed to the intervention itself, but that comparison is only trustworthy if each arm is large enough to reliably detect a real effect. This guide first explains what an RCT is and when it applies, then the general formula behind the calculation, then walks through how to apply it in AnalyZ Solutions.

01What an RCT is, and when it applies

An RCT is a study design in which participants, or sometimes whole groups such as clinics or villages, are assigned to a treatment or a control condition using a random process, such as a coin flip, a random number generator, or a lottery. Because assignment is random and not chosen by the participants, the researcher, or circumstance, the two arms should be similar to each other on average in every way except whether they received the intervention, both in ways that were measured and in ways that were not.

An RCT is a strong fit when three conditions hold together.

Why randomisation is often called the gold standard

The core problem in evaluating any intervention without randomisation is confounding: people or places that receive a programme often differ from those that do not, for reasons connected to the outcome itself. A programme that targets motivated communities, better resourced clinics, or households already on an upward trajectory will look effective even if it changed nothing, simply because of who ended up in the treatment group. Random assignment breaks this link. On average, across a large enough sample, the treatment and control arms start out balanced on every characteristic, measured or not, so any difference that emerges after the intervention can be attributed to the intervention itself rather than to pre-existing differences between the groups. This is what allows an RCT to support a causal claim, not just a correlation, and is why it is often described as the strongest available design for estimating impact.

Common pitfalls

02The general sample size formula

An RCT can compare either a proportion outcome (a rate, such as percentage employed) or a continuous outcome (a numeric score, such as average income). Each uses its own formula.

For a proportion outcome, comparing the rate in the control arm against the rate in the treatment arm:

n per arm = (Z1−α/2√(2p̄q̄) + Z1−β√(p₁q₁+p₂q₂))2 / (p₁−p₂)2 p₁ and p₂ are the expected outcome rates in the control and treatment arms. p̄ is their average. q = 1 − p in each case.

For a continuous outcome, comparing the mean in each arm:

n per arm = 2(Z1−α/2 + Z1−β)2 × σ2 / Δ2 σ is the expected standard deviation of the outcome, assumed similar in both arms. Δ is the expected difference in means between arms.

Each symbol plays a specific role, and each is a genuine judgment call rather than a fixed input.

How to set confidence level and power in practice

95% confidence and 80% power are conventional defaults, appropriate for most trials with no specific reason to depart from them. Two situations commonly justify moving away from these defaults.

Relaxing either value below the conventional default, rather than raising it, is a much less common and generally harder to justify choice, since it directly increases the chance of either a false positive or a missed effect, whichever value is relaxed.

An underpowered trial is not just a wasted opportunity. A no significant effect result from an underpowered trial is genuinely uninformative. It cannot distinguish the intervention does not work from the trial was too small to tell. A confidence level set too low, meanwhile, increases the risk that an apparent effect is actually just noise, potentially leading to scaling an intervention that does not really work. Register the power calculation in the trial protocol before data collection begins. Once data collection is complete, there is no way to fix an underpowered trial after the fact.

Setting the minimum detectable effect

The gap between p₁ and p₂, or the value of Δ, is the single input most likely to be set unrealistically in practice. It should reflect the smallest effect that would still be considered a meaningful, worthwhile result for the intervention being tested, not the effect the programme team hopes to see in a best case scenario. A trial sized around an optimistic effect that does not materialise will very likely be underpowered to detect the real, smaller effect the intervention actually produces. Where possible, base this figure on effects found in similar interventions in comparable settings, published evaluation literature, or a pilot study, rather than an aspirational target.

One-tailed or two-tailed

The formulas above use Z1−α/2, the two-tailed critical value, appropriate when testing for a difference in either direction. A one-tailed test, appropriate only when a change in one specific direction is the sole interest, uses Z1−α instead, a smaller value that produces a smaller required sample for the same nominal confidence level. Two-tailed is the safer, more conventional default, and is expected by most funders and ethics reviewers unless there is a specific, defensible reason a change in the unexpected direction genuinely does not need to be detected.

Accounting for cluster randomisation

If whole groups, such as clinics, schools, or villages, are randomised to arms rather than individuals, a design effect must be applied. It defaults to 1, meaning no clustering effect, and is always editable. Skipping this adjustment for a cluster-randomised trial will understate the required sample, often substantially, because outcomes within the same cluster tend to be correlated. See the dedicated guide on Design Effect and ICC for how to set this correctly.

How to Calculate This in AnalyZ Solutions

The RCT tool in the Sampling module applies both formulas above automatically, along with attrition and design effect adjustments.

  1. Choose your outcome type. Select Proportion outcome or Continuous (mean) outcome. This determines which formula and which inputs follow.
  2. Set confidence and power. Enter both as numbers. Defaults are 95% confidence and 80% power.
  3. Enter the effect you want to detect. For a proportion outcome, enter the expected rate in the control group and the expected rate in the treatment group. For a continuous outcome, enter the expected standard deviation and the expected difference between arms.
  4. Set expected attrition. This inflates the per-arm recruitment target so the analysed sample still meets the statistical minimum after losses to follow-up.
  5. Choose one-tailed or two-tailed. Two-tailed is selected by default and is appropriate for almost all trials.
  6. Set a design effect, if cluster-randomised. If whole groups, such as clinics or schools, are randomised rather than individuals, enter a design effect directly or calculate one from an intraclass correlation and average cluster size. It defaults to 1, meaning no clustering effect, and is always editable.

Watch calculating sample size for a RCT study in the AnalyZ Solutions interface

03Worked example

Suppose you are evaluating a job readiness training programme. Currently 40% of a comparable population finds employment within six months. You expect the programme to lift this to 55%. You want 80% power at 95% confidence, two-tailed, and expect 15% attrition by endline.

Inputs
Outcome typeProportion
Confidence level95%
Power80%
Control group rate40%
Treatment group rate55%
Test typeTwo-tailed
Expected attrition15%
~171
to recruit per arm, 342 total

04Common mistakes to avoid

Frequently Asked Questions
When is an RCT not the right choice, even if it is the strongest design in principle?
When withholding or delaying the intervention from a control group would be unethical or impractical, when random assignment is not logistically or politically feasible, or when the intervention is already universally available. In these situations, a non-randomised comparison such as those covered in Change in Proportion or Change in Mean is often the realistic alternative.
Should I power my trial for every outcome I plan to measure?
No. Power the trial for its pre-specified primary outcome. Secondary outcomes can still be reported, but powering for every outcome measured inflates the required sample unnecessarily.
What if I want unequal arm sizes, such as two treatment participants for every one control?
The tool assumes equal arm sizes, which is standard for most impact evaluations. Unequal allocation typically requires a larger total sample for the same power, and is usually only worthwhile when the treatment is substantially more costly to deliver than to measure.
How is this different from the Change in Proportion tool?
The statistics are closely related, but RCT specifically assumes random assignment to arms and adds attrition handling. Use Change in Proportion when there is no randomisation, such as a natural comparison or a single group tracked over time.
When should I use a one-tailed test instead of two-tailed?
Only when a change in one specific direction, such as an improvement, is the sole outcome of interest, and a change in the opposite direction would not need to be detected or reported. This is uncommon in most programme evaluation contexts, which is why two-tailed remains the default.
Does randomisation guarantee the treatment and control arms are identical at baseline?
No, not exactly, and not in every individual trial. Randomisation guarantees that, on average across many hypothetical repetitions of the same trial, the arms will be balanced. In any single trial, some imbalance by chance is possible, particularly with a small sample, which is one more reason an adequately powered sample size matters even with proper randomisation.

Ready to power your trial?

Open the RCT tool in AnalyZ Solutions. Free, browser based, your data never leaves your device.

Try it out
Related Guides