Sampling/
Power Analysis Overview
Power Analysis, Overview
Is Your Study Big Enough? A Guide to Statistical Power
Power Analysis evaluates the probability a statistical test will detect a real effect. This guide explains the underlying statistics first, then shows how to apply them using AnalyZ Solutions.
8 min read
Power Analysis
Beginner
Sample Size Estimation answers how many people do I need to estimate this value precisely. Power Analysis answers a related but different question: if I run this specific statistical test, how likely am I to detect a real effect, or how many people do I need to have a good chance of detecting it. This guide first explains what power means and how to choose it, then walks through how to apply it in AnalyZ Solutions.
01What statistical power and significance level mean
Power, formally 1 minus beta, is the probability that a study correctly detects an effect that is really there. A power of 80% means that if the effect genuinely exists at the size specified, there is an 80% chance the test comes back statistically significant, and a 20% chance of missing it entirely, purely due to sample size.
Significance level, alpha, is the counterpart risk. It is the probability of concluding an effect exists when it actually does not, a false positive. A significance level of 0.05, corresponding to 95% confidence, is the standard default across most fields.
How to choose these values
80% power and a 0.05 significance level are conventional defaults, not fixed rules. Consider raising power to 90% when missing a real effect would be especially costly, for example when the result determines whether a programme continues. Consider tightening significance level below 0.05 when a false positive would be especially costly, for example before recommending a policy change to a wide population based on the result. Relaxing either value below its convention is sometimes acceptable for lower stakes, exploratory work, but should be a deliberate choice rather than a default.
A no significant effect result is not the same as no effect. If a study was underpowered, a null result is nearly uninformative. It is equally consistent with there is no effect and there is a real effect, but the study could not detect it. A significance level set too loosely carries the opposite risk, an increased chance of concluding an effect is real when it is actually noise. This is why both power and significance level justification are expected before data collection, not calculated afterward to explain a disappointing or a surprising result. Once data collection is complete, there is no way to fix an underpowered study after the fact.
02Two ways to use each tool
Every test in Power Analysis works in either direction.
- Achieved power. A sample size is already known or fixed, and the tool reports the power that sample size provides. Useful when working backward from a fixed budget or an existing dataset.
- Required sample size. A target power is specified, usually 80% or 90%, and the tool reports the minimum sample needed to reach it.
After calculating, a power curve chart shows how power changes across a range of sample sizes for small, medium, and large effects, a fuller picture than any single number can give.
03Choosing the right test
| Test | Use when | Effect size |
| Two-sample t-test | Comparing means between two independent groups | Cohen's d |
| One-sample t-test | Testing whether a group mean differs from a known reference value | Cohen's d |
| Paired t-test | Comparing two measurements on the same individuals, pre and post | Cohen's dz |
| Chi-square test | Testing association between two categorical variables | Cohen's w |
| McNemar's test | Detecting a change in a binary outcome measured twice on the same subjects | Discordant pair proportions |
| One-way ANOVA | Comparing means across three or more independent groups | Cohen's f |
| Linear regression | Detecting a given R² with one or more predictors | Cohen's f² |
| Logistic regression | Binary outcome model with multiple predictors | Odds ratio and EPV |
| Incidence rate ratio | Comparing two incidence rates, events per person-time, rather than two proportions or means | Incidence rate ratio (IRR) |
Incidence rate ratio is different in kind from the other eight tests. Every other test in this table compares proportions or means from a fixed sample. Incidence rate ratio compares two Poisson rates, events divided by person-time of observation, and is the right choice specifically when follow-up time genuinely varies across participants. If everyone in your study is observed for the same fixed period, your outcome is more naturally a proportion, and the two-sample or chi-square tests above apply instead.
Paired versus independent, once more. If the same people are measured twice, use the paired t-test for a continuous outcome or McNemar's test for a binary outcome, not the two-sample or chi-square versions. Paired designs are almost always more efficient when the option is available.
04Understanding effect size
Every test requires specifying how large an effect should be detectable, expressed in standardised units so it is comparable across different outcome scales. Cohen's benchmarks are the field standard.
| Effect size | Small | Medium | Large |
| Cohen's d, t-tests | 0.2 | 0.5 | 0.8 |
| Cohen's w, chi-square | 0.1 | 0.3 | 0.5 |
| Cohen's f, ANOVA | 0.10 | 0.25 | 0.40 |
| Cohen's f², regression | 0.02 | 0.15 | 0.35 |
When there is no prior basis for the effect size, no pilot data, no comparable published study, a medium effect is the standard planning default. Incidence rate ratio does not have an equivalent standardised Cohen benchmark, since a rate ratio is already expressed in directly interpretable terms, for example an IRR of 1.5 meaning a rate 50% higher in one group. Base this figure on prior surveillance data or published literature for a comparable outcome, rather than a generic small, medium, or large preset.
05Logistic regression's extra rule, EPV
Logistic regression has a second constraint beyond power, Events Per Variable, the number of outcome events divided by the number of predictors in the model.
EPV = events / predictors
A conventional minimum of 10 events per predictor helps avoid an unstable, overfit model, one whose estimated relationships would likely look quite different if the study were repeated with a new sample. A model with 5 predictors, for example, needs at least 50 outcome events, not 50 total participants, to meet this minimum. Since the total sample required to accumulate a given number of events depends on how common the outcome is, a rarer outcome requires a much larger total sample to reach the same EPV.
The logistic regression tool calculates the required sample from both constraints, the power-based minimum and the EPV-based minimum, and reports whichever is larger as the binding requirement. It also reports the actual EPV implied by the result, so it is visible which constraint determined the final sample size.
If the EPV rule is binding, meaning it produces a larger required sample than the power calculation alone, reducing the number of predictors in the model, if that can be justified on substantive grounds, lowers the EPV-based requirement directly, since EPV depends on the number of predictors as well as the number of events.
How to Use This in AnalyZ Solutions
- Choose the test matching your planned analysis. Select from the nine tests listed above, grouped by what you are comparing: means between groups, a categorical association, a change measured twice on the same subjects, a regression model, or two incidence rates. If you are unsure which test fits your design, the "Choosing the right test" table above maps each one to a specific use case.
- Choose achieved power or required sample size. Select achieved power if your sample size is already fixed, by budget, an existing dataset, or another constraint, and you want to know what power that gives you. Select required sample size if you are still planning and want to work forward from a target power to the minimum n needed.
- Set your significance level and effect size. 0.05 is the standard significance level default. For effect size, use a value calculated from pilot data or comparable published research if one is available, or one of Cohen's small, medium, or large benchmarks otherwise, as discussed above. This single input drives the result more than any other, so treat it as a genuine assumption worth documenting, not an arbitrary setting.
- For logistic regression, also enter the number of predictors and your target EPV. The tool calculates both the power-based minimum sample size and the EPV-based minimum, and reports whichever is larger as the binding requirement, along with the actual EPV implied by that result.
- For incidence rate ratio, enter your rates in events per person-year, and person-time rather than a headcount. Enter both rates directly, or one rate plus the incidence rate ratio to detect. In achieved power mode, enter the person-time already available per group; in required sample size mode, the tool solves for the person-time needed instead. This test assumes equal person-time contributed to each group.
- Calculate. In achieved power mode, the result shows the power your specified sample size provides. In required sample size mode, it shows the minimum n needed to reach your target power, at your chosen significance level and effect size.
- Review the power curve. After calculating, view the chart showing how power changes across a range of sample sizes for small, medium, and large effects, with your own inputs marked. This makes visible how sensitive your result is to the effect size assumption, not just to the sample size itself.
Watch power analysis in the AnalyZ Solutions interface
06Worked example
An evaluation plans to compare average scores on a knowledge test between two independent groups, using a two-sample t-test. Based on similar studies, a medium effect, Cohen's d of 0.5, is a realistic planning assumption. The team wants 80% power at a significance level of 0.05.
Inputs
TestTwo-sample t-test
ModeRequired sample size
Significance level0.05
Power target80%
Effect size, Cohen's d0.5
Running the same inputs in achieved power mode, entering 64 per group directly, confirms the result reaches approximately 80% power at this effect size. The power curve chart shows that the same sample would achieve noticeably lower power against a small effect, Cohen's d of 0.2, and higher power against a large effect, Cohen's d of 0.8, illustrating why the effect size assumption drives the result as much as the sample size itself.
07Common mistakes to avoid
- Choosing an optimistic effect size rather than a realistic one. Powering a study around a large effect when the true effect is more likely small or medium produces a study with far less power than intended, in practice.
- Using an independent groups test when the design is paired. If the same individuals are measured twice, use the paired t-test or McNemar's test, not the two-sample or chi-square versions, which understate the available power for a paired design.
- Ignoring EPV when planning a logistic regression model. A sample sized only from the power calculation can still leave the model unstable if the number of events relative to predictors is too low. Check both constraints, not just power.
- Calculating power after the study is already complete to explain a null result. This, sometimes called post-hoc power, is not informative and is generally discouraged. Power should be calculated and justified before data collection, as part of planning the study.
- Using incidence rate ratio when follow-up time is actually fixed and equal for everyone. If every participant is observed for the same period, a proportion-based test such as chi-square or the two-sample t-test is simpler and equally valid. Reserve incidence rate ratio for genuinely variable follow-up time.
Frequently Asked Questions
How is this different from Sample Size Estimation?
Sample Size Estimation is generally built around estimating a value to a stated precision, or, for the Measuring Change designs, a specific real-world comparison such as an RCT or a before and after study. Power Analysis is built around named statistical tests run directly in analysis software, useful whenever a plan does not map cleanly onto one of the study-design tools.
What power should I target, 80% or 90%?
80% is the conventional minimum across most fields. Consider 90% for high stakes decisions where missing a real effect would be especially costly, such as a go or no-go funding decision.
Can I use a custom effect size instead of Cohen's benchmarks?
Yes. Every tool accepts a custom effect size value, not only the small, medium, and large presets. A pilot-derived estimate is always preferable to a generic benchmark when one is available.
What is the difference between achieved power and required sample size mode?
Achieved power mode takes a sample size, whether fixed by budget, an existing dataset, or another constraint, and reports the power that sample size provides. Required sample size mode takes a target power and reports the minimum sample needed to reach it. The same underlying formula is used in both directions.
Why does the EPV requirement sometimes exceed the power-based requirement for logistic regression?
EPV depends on the number of predictors in the model and how common the outcome is, independent of the effect size being tested. A model with many predictors, or an outcome that occurs rarely, can require a larger sample to maintain model stability than the power calculation alone would suggest, particularly when the odds ratio being detected is large and therefore easy to detect statistically.
Is post-hoc power analysis, calculated after a study finds a non-significant result, useful?
Generally not. Power calculated from the effect size actually observed in a completed study is mathematically determined by the p-value already obtained, and does not provide new information about whether the original null result reflects a true absence of effect or an underpowered study. Power should be planned before data collection, not calculated afterward.
How is incidence rate ratio different from the Change in Rate tool under Sample Size Estimation?
Both use the same underlying idea, comparing two incidence rates measured over person-time. Incidence rate ratio here is a lean, direct statistical test, usable in either achieved power or required sample size mode, and works in person-time only. Change in Rate under Sample Size Estimation is a fuller study-design tool, it converts person-time into an actual headcount using your stated follow-up, and adds non-response and clustering adjustments on top. Use this tool for a quick power check, and Change in Rate when planning full study recruitment.
Ready to check your study's power?
Open Power Analysis in AnalyZ Solutions. Free, browser based, your data never leaves your device.
Try it out
Related Guides