Statistical Power Calculations: How to Size Your Sample Before Applying for Grants
Calculate sample size per group with Cohen effect sizes and alpha thresholds, and generate grant-ready power analysis boilerplate.
Most sample sizes in biomedical research are chosen by precedent — "the last paper used six mice, so we will use six mice." Power analysis replaces that with a calculation, and the calculation usually says you need more animals than you planned.
**The short answer.** Statistical power is the probability that your experiment detects a real effect of a given size. Convention is 80%. To compute the sample size you need three inputs: the effect size you care about, the variability of your measurement, and your significance threshold. Get any of them from a pilot study and you will underestimate the sample size, often badly.
**What power actually means.** Power is 1 − β, where β is the chance of a Type II error: a real effect that your experiment fails to detect. At 80% power, one experiment in five that studies a genuinely real effect of the size you specified will return a non-significant result. That is the deal you are accepting, and it is worth saying out loud, because 80% is a convention, not a law. For an expensive or irreversible experiment, 90% is often the better trade.
The four quantities — sample size, effect size, alpha, and power — are locked together. Fix any three and the fourth follows. That is the whole of power analysis.
**The formula, for two independent groups.**
n per group = 2 × ((z(α/2) + z(β)) / d)²
where d is Cohen's d, the standardised effect size: the difference between the group means divided by the pooled standard deviation.
At α = 0.05 (two-tailed) and 80% power, z(α/2) = 1.96 and z(β) = 0.84, so the numerator is a constant:
n per group ≈ 2 × (2.80 / d)² ≈ 15.7 / d²
That single expression is worth memorising, because it makes the cost of small effects immediate.
**A worked example.**
You are measuring tumour volume in treated versus control mice. Pilot data: control mean 450 mm³, treated mean 320 mm³, pooled SD 140 mm³.
d = |450 − 320| / 140 = 0.93
n per group ≈ 15.7 / 0.93² = 15.7 / 0.865 ≈ 18.2 → 19 mice per group, 38 total.
Now suppose the true effect is smaller than your pilot suggested — say the real difference is 90 mm³, giving d = 0.64:
n ≈ 15.7 / 0.41 ≈ 38 per group, 76 total.
A 30% reduction in the effect size doubled the animals required. Power scales with the square of d, which is why underestimating the effect size is the single most expensive mistake in experimental design.
**Why pilot studies mislead you.** This is the part most guides skip.
A small pilot gives you a noisy estimate of both the effect size and the standard deviation. Noise is symmetric, but its consequences are not: you only proceed to the full experiment when the pilot looks encouraging, which means you preferentially proceed on pilots that overestimated the effect. That is a selection effect, and it biases your power calculation in the optimistic direction every time.
There is also a hard statistical floor: an SD estimated from n=5 has a 95% confidence interval running roughly from 0.6× to 2.0× the true value. Plug the low end into the formula and you will underpower by a factor of three.
Two defences. First, use the smallest effect size that would be **scientifically** meaningful, not the effect your pilot happened to show — if a 20% reduction in tumour volume is the smallest result that would change what anyone does next, power for 20%. Second, if you must use pilot variance, use the upper confidence bound of the SD rather than the point estimate.
**Post-hoc power is not a thing.** If your experiment returned p = 0.08 and a reviewer asks for "observed power", the answer is that observed power is a deterministic function of the p-value and carries no information beyond it. High p-value, low observed power, always. It cannot tell you whether you were underpowered for a real effect or correctly detected the absence of one.
What is legitimate after the fact is a confidence interval on the effect size. "The difference was 40 mm³, 95% CI −15 to 95" tells a reader exactly what the experiment could and could not rule out. That is the honest version of the question post-hoc power is trying to ask.
**Multiple comparisons multiply the cost.** If you are testing four doses against a control, you are running four tests, and your familywise error rate is no longer 0.05. Under Bonferroni, each test runs at α = 0.0125, so z(α/2) rises from 1.96 to 2.50, the numerator constant rises from 2.80 to 3.34, and the sample size per group rises by about 42%.
This is a design decision, not an analysis decision. Deciding to correct after you have collected the data means you are underpowered for the analysis you are actually running.
**Paired designs are cheaper, sometimes dramatically.** If each subject can serve as its own control — before/after, contralateral limb, same donor across conditions — the relevant SD becomes the SD of the *differences*, not the pooled SD of the raw values. When the measurement has high between-subject variability and the treatment effect is consistent, the SD of differences can be a fraction of the pooled SD, and the required n falls with the square of that ratio.
For a measurement with high inter-animal variability, a paired design can cut the animal count by more than half. That is an ethics argument as much as a statistical one.
**Writing this into a grant.** NIH and most funders now expect an explicit sample-size justification, and reviewers read it. A sufficient one names five things:
The primary outcome and how it is measured. The smallest effect size that would be meaningful, and where that number comes from. The variance estimate, and its source. Alpha, power, and whether the test is one- or two-tailed. The resulting n, plus any inflation for expected attrition.
One paragraph, five facts. "Assuming a 25% reduction in infarct volume (the smallest effect considered clinically meaningful, per Smith 2023) and a pooled SD of 18% observed in our published cohort, detecting this difference with 80% power at α = 0.05 two-tailed requires 17 animals per group. Allowing for 10% attrition, we will enrol 19 per group."
**What SciKeep does with this.** The power calculator takes effect size, variance and alpha and returns n, and it will also run the calculation in reverse — given the n you can afford, what is the smallest effect you could detect? That reverse question is frequently the more useful one, because it tells you before you start whether the experiment is worth running at all.
It also generates the grant paragraph in the format above, with the assumptions stated explicitly, because an unstated assumption is the thing reviewers flag.
**References.**
Cohen J (1988). Statistical Power Analysis for the Behavioral Sciences, 2nd ed. Routledge. doi:10.4324/9780203771587
Button KS et al. (2013). Power failure: why small sample size undermines the reliability of neuroscience. Nature Reviews Neuroscience 14:365–376. doi:10.1038/nrn3475
Hoenig JM, Heisey DM (2001). The abuse of power: the pervasive fallacy of power calculations for data analysis. The American Statistician 55(1):19–24. doi:10.1198/000313001300339897
Faul F et al. (2007). G*Power 3: a flexible statistical power analysis program. Behavior Research Methods 39:175–191. doi:10.3758/BF03193146
Percie du Sert N et al. (2020). The ARRIVE guidelines 2.0. PLOS Biology 18(7):e3000410. doi:10.1371/journal.pbio.3000410