Biological vs Technical Replicates in qPCR: What Counts as n?
What counts as n in qPCR? Average technical replicates within each biological replicate, then run statistics on biological replicates only — with Livak 2^-ddCt and Pfaffl worked examples.
If you ran three wells of the same cDNA and called that n=3, this article is about the mistake that follows — and it is the most common statistical error in published qPCR.
**The short answer.** Technical replicates measure your pipetting. Biological replicates measure biology. Only biological replicates count toward n, and only they can support a p-value about a treatment effect. Technical replicates are averaged within each biological replicate before any statistics are run. If you have one biological replicate per condition, you have no n, and no honest p-value exists no matter how many wells you loaded.
**Why this is not pedantry.** The two kinds of replicate estimate different variances. Technical variance is the spread you get from loading the same material more than once: pipetting error, well-to-well thermal variation, optical noise. On a decent instrument with a decent operator, that is small — typically under 0.2 Cq. Biological variance is the spread between independently treated units: different flasks, different animals, different differentiation batches. That is usually far larger, and it is the only variance your hypothesis is about.
When you feed three technical wells into a t-test as n=3, you are asking the test to compare group means against technical variance. Technical variance is small, so the denominator of your test statistic is small, so t is large, so p is small. You will get a significant result almost regardless of whether the treatment did anything. Hurlbert named this pseudoreplication in 1984, and forty years later it remains the reason a large fraction of qPCR figures do not replicate.
**A worked example, with numbers.**
Suppose you treat one flask of cells with a drug and one flask with vehicle, and you run each in triplicate on the plate. Target gene, normalised to a reference gene, Cq values:
Vehicle flask, ΔCq per well: 4.02, 4.11, 3.97 Treated flask, ΔCq per well: 2.88, 2.95, 2.91
Analysed as n=3 per group, an unpaired t-test on these six numbers gives t ≈ 21 and p ≈ 3×10⁻⁵. Fold change 2^-ΔΔCt = 2^-(2.913 − 4.033) = 2.17. It looks like a clean, highly significant 2.2-fold induction.
It is not a result at all. The p-value is answering the question "are these two flasks different?" — to which the answer is yes, and would have been yes if both flasks had received vehicle, because two flasks are always slightly different. The experiment contains one independent observation per condition. n=1.
Now run it properly. Three independent flasks per condition, each with three technical wells. Average the technical wells within each flask first:
Vehicle ΔCq per flask: 4.03, 3.72, 4.41 Treated ΔCq per flask: 2.91, 3.44, 2.60
Now n=3 per group. t ≈ 3.4, p ≈ 0.027. Fold change 2.11 — almost the same point estimate. But the p-value moved by three orders of magnitude, and the confidence interval is now wide enough to be honest about what three flasks can tell you.
Same fold change. Completely different claim. The first version is not a stronger version of the second; it is a different, false statement.
**Where the technical replicates actually go.** Averaging them within each biological replicate is not throwing information away — it is putting the information where it belongs. The technical spread tells you whether your pipetting and your plate setup are sound. If technical Cq values within a well group differ by more than about 0.3–0.5 Cq, something is wrong: a bubble, a bad dilution, a poorly mixed master mix. That is a QC signal, and worth acting on. It is not evidence about your drug.
A useful discipline: report both. "n = 3 biological replicates, each measured in technical triplicate (mean technical CV 1.2%)" tells a reviewer that you know the difference and that your bench work is clean.
**Why statistics run on ΔCq, not on fold change.** This is the second error, and it travels with the first.
Cq values are logarithmic — each cycle is roughly a doubling. ΔCq is therefore a log-ratio, and log-ratios are approximately normally distributed, which is what a t-test assumes. Fold change is 2^-ΔΔCq, an exponentiated quantity, and it is log-normal: skewed, with a hard floor at zero and no ceiling.
Run a t-test on fold-change values and you are testing a skewed variable with a normality assumption. The consequences are systematic rather than random: means are pulled upward by the long right tail, error bars become asymmetric in a way that symmetric SEM bars misrepresent, and downregulation is compressed into the interval between 0 and 1 while upregulation stretches to infinity. A 2-fold increase and a 2-fold decrease are the same magnitude of biological effect, but as fold changes they are 2.0 and 0.5 — visibly unequal.
Do the statistics on ΔCq. Convert to fold change only for display, and back-transform your confidence interval rather than computing one on the fold-change scale.
**The reference gene problem.** ΔCq is only meaningful if the reference gene is genuinely stable across your conditions. There is no universally stable reference gene, and the popular ones are popular for historical reasons rather than measured ones.
GAPDH is glycolytic and shifts under hypoxia, altered glucose, and many metabolic drugs. ACTB changes with confluence, cytoskeletal drugs and differentiation state. 18S is abundant enough that it often sits 10–15 cycles away from your target, which puts it in a different part of the amplification range.
Validate in your own conditions. Run two or three candidates, check that their Cq values do not shift systematically between groups, and normalise to the geometric mean of the stable ones. Vandesompele's geNorm work (2002) is the standard method; NormFinder and BestKeeper do the same job. If your reference gene moves with treatment, every fold change you compute is wrong by exactly that amount, and no amount of replication will reveal it.
**Efficiency correction, and when it matters.** The 2^-ΔΔCq formula assumes both amplicons double every cycle — 100% efficiency. Real primer pairs run between roughly 90% and 110%, and the error compounds with ΔCq.
If your target amplifies at 95% and your reference at 105%, an apparent 2-fold change over a ΔΔCq of 1 is off by several percent. Over a ΔΔCq of 5, the error is large enough to change the conclusion. Pfaffl's 2001 formulation corrects for this by raising each measured efficiency to its own ΔCq rather than assuming 2. Use it when your efficiencies are known and differ, and run a standard curve to know them. The MIQE guidelines (Bustin et al. 2009) ask for efficiency to be reported; it is one of the more commonly skipped requirements.
**What SciKeep does with this.** The qPCR tool asks for your design before it asks for numbers: how many biological replicates, how many technical wells within each. It averages technical wells within each biological replicate automatically, runs the statistics on ΔCq, and reports n as the number of biological replicates.
If you enter one biological replicate per condition, it will not produce a p-value. It shows the fold change, states that technical replicates measure pipetting rather than biology, and asks for biological replicates. That refusal is deliberate. Every other calculator will happily return p = 3×10⁻⁵ for the first dataset above, and a reviewer who checks will reject the paper.
**Checklist before you submit.**
Report n as biological replicates and say so explicitly. State how technical replicates were handled. Name your reference genes and say how you validated them in your conditions. Report amplification efficiency for each assay. Say whether statistics were run on ΔCq or on fold change. Show individual biological replicate points on the figure rather than bars alone — a SuperPlot makes the replicate structure visible at a glance.
Every one of those is a MIQE requirement, and every one is routinely omitted.
**References.**
Bustin SA et al. (2009). The MIQE guidelines: minimum information for publication of quantitative real-time PCR experiments. Clinical Chemistry 55(4):611–622. doi:10.1373/clinchem.2008.112797
Livak KJ, Schmittgen TD (2001). Analysis of relative gene expression data using real-time quantitative PCR and the 2^-ΔΔCT method. Methods 25(4):402–408. doi:10.1006/meth.2001.1262
Pfaffl MW (2001). A new mathematical model for relative quantification in real-time RT-PCR. Nucleic Acids Research 29(9):e45. doi:10.1093/nar/29.9.e45
Vandesompele J et al. (2002). Accurate normalization of real-time quantitative RT-PCR data by geometric averaging of multiple internal control genes. Genome Biology 3(7):research0034. doi:10.1186/gb-2002-3-7-research0034
Hurlbert SH (1984). Pseudoreplication and the design of ecological field experiments. Ecological Monographs 54(2):187–211. doi:10.2307/1942661