Sampling/
Change in Mean
Sample Size Estimation, Measuring Change
How to Calculate Sample Size to Detect a Change in a Numeric Outcome
A change in mean compares a numeric average across two groups, or the same group at two points in time. This guide explains the underlying statistics first, then shows how to apply them using AnalyZ Solutions.
6 min read
Sample Size Estimation
Intermediate
Not every before and after comparison involves a percentage. When the outcome is a numeric average, a test score, an income level, an attendance count, this tool applies the equivalent logic for continuous data. This guide first explains the general formula behind the calculation, then walks through how to apply it in AnalyZ Solutions.
01The general sample size formula
As with Change in Proportion, the first decision is whether you are comparing two independent groups or the same group measured twice.
Two independent groups
n per group = 2(Z1−α/2 + Z1−β)2 × σ2 / Δ2
Same group, before and after
n = (Z1−α/2 + Z1−β)2 × 2σ2(1−ρ) / Δ2
Choosing your confidence level and power
Confidence level and power are judgment calls, not fixed rules, and they deserve as much thought as the effect size itself.
Confidence level reflects how willing you are to risk a false alarm, concluding a change happened when it actually did not. 95% is the standard default, meaning a 5% chance of that false alarm. A higher level, such as 99%, is worth considering before a costly or hard to reverse decision. A lower level, such as 90%, is sometimes acceptable for lower stakes, exploratory work.
Power reflects how willing you are to risk missing a real change entirely. 80% power is the conventional minimum, meaning a 20% chance of failing to detect a genuine change of the size specified. 90% power is worth considering for high stakes decisions where missing a real effect would be especially costly.
Undermining either value carries a real cost. A confidence level set too low increases the risk of reporting a change that is actually just noise. Power set too low means a genuinely effective programme has a meaningful chance of producing a study that fails to detect its own impact. Once data collection is complete, there is no way to fix this after the fact.
The standard deviation assumption drives this calculation more than any other input. Unlike a proportion, there is no built in worst case default for σ. Pull this figure from a pilot study, a prior round of the same instrument, or comparable published data wherever possible, not a guess.
How to Calculate This in AnalyZ Solutions
The Change in Mean tool applies both formulas automatically based on the design you select.
- Choose your design. Select Two independent groups or Same group, before and after.
- Set confidence and power. Enter both as numbers.
- Enter the expected standard deviation and the change to detect. Both should be in the same units as the outcome you are measuring.
- For the paired design, enter an assumed correlation. 0.5 is a reasonable default if unsure.
- Choose one-tailed or two-tailed. Two-tailed is selected by default.
- Set a design effect, if sampling by cluster. Defaults to 1 and is always editable.
VIDEO WALKTHROUGH PLACEHOLDER, 90 SECONDS
A short screen recording showing these steps in the AnalyZ Solutions interface can be embedded here.
02Worked example
A literacy programme wants to detect a change in average test score among the same cohort of learners, tested before and after a six month course. The expected standard deviation of scores is 12 points, the programme targets a 6 point improvement, and the assumed correlation between someone's before and after score is 0.5.
Inputs
DesignPaired
Confidence level95%
Power80%
Expected standard deviation12
Change to detect6
Correlation (ρ)0.5
~32
pairs, learners tested twice
The same effect size and power calculated as two independent groups of learners, rather than the same learners tracked twice, would require about 63 people per group, roughly double, again illustrating the efficiency of paired designs whenever tracking the same individuals is feasible.
Frequently Asked Questions
How is this different from Change in Proportion?
The design logic, independent groups versus paired before and after, is identical. The difference is the outcome type. Use Change in Proportion when the outcome is a percentage or rate, and Change in Mean when it is a numeric average.
What if this is a randomised controlled trial specifically?
Where do I get a realistic standard deviation if I have no pilot data?
Check published studies using the same or a similar outcome measure in a comparable population. If truly nothing is available, a conservative approach is to use a wider standard deviation than expected, which produces a larger, safer sample size rather than a smaller, riskier one.
Ready to size your comparison?
Open the Change in Mean tool in AnalyZ Solutions. Free, browser based.
Try it out
Related Guides