Sampling/
Change in Rate
Sample Size Estimation, Measuring Change
How to Calculate Sample Size to Detect a Change in an Incidence Rate
How to calculate sample size for comparing two incidence rates measured over person-time, rather than two proportions.
7 min read
Sample Size Estimation
Intermediate
Some outcomes are not naturally a percentage. A disease incidence, an accident frequency, or a service utilisation count, measured while people are followed for different lengths of time, is a true incidence rate, not a proportion. This tool compares two such rates. This guide first explains what an incidence rate is and when it applies, then the general formula behind the calculation, then walks through how to apply it in AnalyZ Solutions.
01What an incidence rate is, and when it applies
A proportion answers what share of a fixed group experienced an outcome. An incidence rate answers how frequently an outcome occurs per unit of time at risk, events divided by person-time, typically expressed as events per person-year. The distinction matters whenever follow-up time is not the same for everyone. If one participant is observed for two months and another for two years before the study ends, a simple percentage of who experienced the outcome treats those two very differently exposed people as equivalent. An incidence rate does not.
Use a rate-based comparison rather than a proportion-based one when two conditions hold together.
- Follow-up time genuinely varies across participants. People enter the study at different times, are lost to follow-up at different points, or the outcome can occur at any point during a variable observation window.
- The comparison is between two rates, not one. If you are estimating a single rate, use the companion tool for a single incidence rate estimate instead. This tool is specifically for detecting a difference between two rates, such as an intervention group and a comparison group.
If everyone in your study is followed for exactly the same fixed period, and you only want to compare the percentage who experienced the outcome in each group, that is a proportion comparison, not a rate comparison. Use
Change in Proportion or
RCT instead, both of which are simpler to interpret and report when the underlying quantity really is a proportion.
02The general sample size formula
Comparing two incidence rates uses a different formula from comparing two proportions, because the statistical behaviour of a rate is different from the statistical behaviour of a proportion. A proportion is bounded between 0 and 1, and its variance depends on p(1−p). A rate has no upper bound, and for a Poisson process, its variance is simply equal to the rate itself.
T = (Z1−α/2√(2λ̄) + Z1−β√(λ₁+λ₂))2 / (λ₁−λ₂)2
Each symbol plays a specific role.
- Z1−α/2 (confidence level). How sure you want to be that a detected difference is real, rather than random noise, at 95% confidence this is 1.96.
- Z1−β (power). How likely you want to be to actually detect the difference, if it is really there, at 80% power this is approximately 0.84.
- λ₁ and λ₂ (the two rates). Your expected rate in each group, in events per person-year. Set these from prior surveillance data, published literature, or a pilot period, in the same units you intend to report.
The result of this formula is person-time, not a headcount. To convert it into a number of participants to enrol, divide by how much follow-up time each participant is expected to contribute on average.
headcount = T / average follow-up per person
Entering a rate ratio instead of two rates
If you already know one rate and want to specify how large a relative difference you want to detect, enter the incidence rate ratio instead of the second rate directly. An incidence rate ratio of 1.5 means the second group's rate is expected to be 50% higher than the first. This is mathematically equivalent to entering both rates directly, since the second rate is simply the first rate multiplied by the ratio.
Choosing your confidence level and power
95% confidence and 80% power are conventional defaults. Raise power to 90% when missing a real difference between rates would be especially costly, for example when a result determines whether a control intervention is scaled. Tighten confidence below 0.05 before a costly or hard to reverse recommendation based on the result. As with any comparison, undermining either value carries a real cost: too low a confidence level increases the risk of reporting a difference that is actually noise, and too low a power means a real difference has a meaningful chance of going undetected, an outcome indistinguishable afterward from there being no real difference at all.
One-tailed or two-tailed
The formula above uses Z1−α/2, the two-tailed critical value, appropriate when testing for a difference in either direction. A one-tailed test, appropriate only when the direction of any true difference is known in advance and the opposite direction is of no interest, uses Z1−α instead, producing a smaller required sample for the same nominal confidence level. Two-tailed is the safer, more conventional default.
Accounting for cluster sampling
If participants are sampled by cluster, for example recruited from a sample of clinics rather than individually across an entire population, a design effect should be applied. It defaults to 1, meaning no clustering effect, and is always editable. See the dedicated guide on Design Effect and ICC for how to set this correctly.
How to Calculate This in AnalyZ Solutions
- Choose how to enter your rates. Select "Enter both rates directly" if you have separate estimates for each group, or "Enter rate and incidence rate ratio" if you want to specify the relative difference you are trying to detect.
- Set confidence and power. Enter both as numbers. Default values are 95% and 80%, respectively.
- Enter your incidence rate, or rates, in events per person-year. If your data is naturally expressed in a different unit, such as per 1,000 person-years, convert to events per person-year first by dividing accordingly.
- Enter the average follow-up per person, in years. This converts the person-time result into a headcount. It should reflect your actual planned study duration and expected retention, not an idealised full year of observation for everyone.
- Set your expected non-response rate. This inflates the recruitment target so the completed sample still meets the statistical minimum.
- Choose one-tailed or two-tailed. Two-tailed is selected by default.
- Set a design effect, if sampling by cluster. Defaults to 1 and is always editable.
- Calculate. The result shows participants per group, the person-time required, the minimum headcount before non-response, and the implied incidence rate ratio your inputs describe.
Watch calculating sample size to detect chaneg in rate in the AnalyZ Solutions interface
03Worked example
A study is comparing disease incidence between a group receiving a preventive intervention and a comparison group. The comparison group's expected incidence is 0.10 events per person-year, and the study should detect a rate 50% higher, an incidence rate ratio of 1.5, at 95% confidence and 80% power, two-tailed. Participants are expected to contribute one year of follow-up on average, with 10% non-response.
Inputs
Confidence level95%
Power80%
Rate, group 10.10 per person-year
Incidence rate ratio to detect1.5
Average follow-up1 year
Non-response rate10%
Test typeTwo-tailed
~873
participants to recruit per group
Before the non-response adjustment, this requires about 785 person-years of follow-up per group, or 785 participants at one year of average follow-up each. A comparatively modest expected difference between rates of 0.10 and 0.15 per person-year requires a substantial sample, which is typical for rate comparisons involving relatively uncommon events.
04Common mistakes to avoid
- Using this tool when follow-up time is actually fixed and equal for everyone. If every participant is observed for the same period, a proportion-based comparison is simpler and just as valid. Reserve this tool for genuinely variable follow-up.
- Mixing units between the two rates, or between the rates and the follow-up time. Both rates must be in the same time unit, and that unit must match the follow-up time field. Convert everything to events per person-year before entering it.
- Assuming an idealised full year of follow-up for every participant. The average follow-up figure should reflect realistic retention and study duration, not the best case where nobody is lost before the study ends.
- Forgetting the design effect for a clustered sample. If participants are recruited from a sample of sites rather than individually, skipping the design effect will understate the true required sample.
Frequently Asked Questions
How is an incidence rate different from a proportion?
A proportion is the share of a fixed group that experienced an outcome, bounded between 0 and 1. An incidence rate is events divided by person-time of observation, and has no upper bound. Use a rate when follow-up time varies across participants, and a proportion when everyone is observed for the same fixed period.
What if I only want to estimate one rate, not compare two?
Use the companion Incidence Rate (Single Estimate) tool instead, found under Cross-sectional. It estimates a single rate to a target precision rather than testing for a difference between two rates.
My rates are reported per 1,000 or per 100,000 person-years. How do I convert?
Divide by the denominator to convert to events per person-year. A rate of 12 per 1,000 person-years becomes 0.012 events per person-year. Apply the same conversion to both rates so they remain comparable.
Why does the result show person-time as well as a headcount?
The underlying statistical calculation is genuinely in person-time, since that is what determines statistical precision for a rate. The headcount is a practical conversion, based on your stated average follow-up, to help with planning recruitment. If actual follow-up varies substantially from your assumption, the true required headcount will differ from this estimate.
Does this tool assume person-time accrues at a constant rate?
Yes. It assumes participants contribute a broadly similar average follow-up and that the rate does not change sharply over the observation period. If accrual is very uneven, for example most participants are lost early or enrol very late, treat the headcount conversion as an approximation rather than an exact figure.
Ready to calculate the sample size involving rate comparison?
Open the sample size tool in AnalyZ Solutions. Free, browser based, your data never leaves your device.
Try it out
Related Guides