Sample Size & Statistical Power Calculator

Free sample size and power calculator. Work out the n you need from a Cohen’s d effect size, or the power a planned experiment actually has, with grant-ready justification text to paste into a proposal. Runs in your browser.

How it works

Sample size is derived from the standardised effect size (Cohen’s d), the desired significance level, and the target power. The tool reports the n per group required, and inverts the calculation to report the power achieved by a sample size you already have — which is the useful direction when reviewing a completed experiment.

Frequently asked questions

How do I calculate sample size for a t-test?

Choose the smallest difference between group means that would change your conclusion, divide it by the pooled standard deviation from pilot or published data to get Cohen’s d, then set your significance level (usually 0.05) and target power (usually 80%). The calculator returns n per group. For a two-sample t-test at 80% power and alpha 0.05 the requirement is roughly 16/d² per group, so a d of 0.5 needs about 64 per group.

Can I calculate power after the experiment is finished?

You can, but observed or post-hoc power — power recomputed from the effect you happened to measure — tells you nothing the p-value has not already told you, because the two are mathematically tied: a non-significant result always yields low observed power. It is not evidence that a study was underpowered. What is legitimate after the fact is asking what effect size your sample size could have detected, which this tool reports, and which reviewers do accept.

What effect size should I assume if I have no pilot data?

Use the smallest effect that would be scientifically meaningful rather than the one you hope to see — the question is what size of difference would change what anyone does next. If you must fall back on convention, Cohen’s benchmarks are d = 0.2 small, 0.5 medium, 0.8 large, but they were never intended as substitutes for domain judgement and reviewers increasingly say so. Assuming a large effect because it yields a comfortable n is the most common way a power calculation becomes fiction.

Does sample size mean biological or technical replicates?

Biological replicates. n is the number of independent experimental units — separate animals, separate cultures, separate preparations. Technical replicates measure the same unit more than once; they reduce measurement noise and should be averaged before analysis, not counted towards n. Treating them as independent is pseudoreplication, and it inflates significance rather than power.

How many replicates do I need for my experiment?

It depends on the size of the effect you need to detect and the variability of your measurement. Sample size is driven by the standardised effect size — the difference between group means divided by the pooled standard deviation. As a rough guide at 80% power and a 0.05 significance level, detecting a large effect (d = 0.8) takes about 26 per group, a medium effect (d = 0.5) about 64, and a small effect (d = 0.2) about 394.

What is statistical power and why does 80% matter?

Power is the probability that a study will detect an effect that genuinely exists. At 80% power, one in five real effects will be missed. The 80% convention is a pragmatic balance between the cost of larger experiments and the risk of false negatives, not a statistical law — safety-critical work often targets 90% or higher.

Method reference

Cohen J. (1988) Statistical Power Analysis for the Behavioral Sciences, 2nd ed.

SciKeepLoading SciKeep…