SuperPlots: How to Solve the N=1 Pseudoreplication Crisis in Cell Biology Microscopy
Treating 500 cells from one dish as N=500 inflates false positives. How SuperPlots combine single-cell spread with independent experiment means.
A bar chart of 3,000 individual cells, from three dishes, with error bars and p < 0.0001, is one of the most common figures in cell biology. It is also, in most cases, statistically meaningless. SuperPlots are the fix, and they take about ten minutes to adopt.
**The short answer.** If you treated three dishes and measured a thousand cells in each, your n is 3, not 3,000. Cells within a dish are not independent — they share a passage, a medium change, an incubator position and a treatment event. Pooling them inflates your apparent sample size a thousandfold and makes p-values essentially arbitrary. A SuperPlot shows both levels at once: every cell, colour-coded by which replicate it came from, with the three replicate means overlaid as large symbols. Statistics run on the three means.
**Why pooling breaks the test.** Every standard statistical test assumes independent observations. Cells in the same dish violate that assumption in a specific, directional way: they are more similar to each other than to cells in another dish, because they share everything except their individual biology.
The consequence is not subtle. Standard error scales as 1/√n. Going from n=3 to n=3,000 shrinks your error bars by a factor of about 32, and shrinks the denominator of your test statistic by the same amount. Almost any difference becomes significant. Lord et al. (2020) demonstrated this directly: pooled analyses of the same data routinely returned p < 0.0001 where the correct replicate-level analysis returned p > 0.05.
The tell is a figure where the p-value is spectacular and the effect is small. If a 6% difference in mean fluorescence has p = 10⁻¹², the p-value is describing your cell count, not your biology.
**A worked example.**
Three independent experiments, ~1,000 cells each, measuring nuclear translocation of a transcription factor.
Per-replicate means, control: 0.62, 0.71, 0.58 Per-replicate means, treated: 0.78, 0.89, 0.69
Pooled across all cells (n ≈ 3,000 per group), a t-test gives p < 10⁻¹⁵. Presented that way, the result looks unassailable.
Run it correctly on the three replicate means: t ≈ 2.6, p ≈ 0.06. The effect is real-looking and consistent in direction across all three replicates, but three replicates cannot establish it at α = 0.05. The honest conclusion is "consistent trend, n=3, not significant" — and the obvious next step is a fourth and fifth replicate.
Note what did *not* change: the effect size. Both analyses say treatment raises the ratio by about 0.15. The disagreement is entirely about how confident you are entitled to be.
**What a SuperPlot looks like.** Three layers on one axis:
The individual cells, plotted as small semi-transparent points, jittered so density is visible — this shows the distribution, the spread, and any bimodality that a bar would hide.
Colour by biological replicate. Replicate 1 in one colour, replicate 2 in another, replicate 3 in a third. This is the layer that does the work: it makes replicate-to-replicate variability visible immediately. If the three colours form three separated clouds, your between-replicate variance dominates and no cell-level p-value means anything.
The replicate means, as three large symbols in the matching colours, with the mean and error bar computed across those three symbols only.
A reader can then see the full distribution, judge the consistency across replicates, and read a p-value that comes from the correct level. Nothing is hidden and nothing is inflated.
**When cell-level data does belong in the analysis.** Sometimes the cell-level distribution is the biology — you are studying heterogeneity, or a subpopulation that appears only under treatment, and the mean is not the interesting quantity.
In that case the answer is not pooling but a hierarchical model: a linear mixed-effects model with replicate as a random effect and treatment as a fixed effect. This uses all the cells while correctly accounting for the nesting, and it estimates both the within-replicate and between-replicate variance components explicitly.
The mixed model is strictly better than averaging when you have unbalanced replicates — 200 cells in one dish and 2,000 in another — because it weights appropriately instead of treating both dish means as equally precise. It needs more replicates than three to estimate the random effect well; five or six is a reasonable floor.
**What counts as a biological replicate.** This is where most of the remaining disagreement lives.
Independent means the units did not share the thing you are testing. Separate dishes seeded from the same flask on the same day, treated from the same drug dilution, imaged in the same session are only partly independent — they share the passage, the drug dilution and the imaging session. Genuinely independent replicates are performed on different days, from different passages, with freshly prepared reagents.
Three dishes split from one flask and treated in parallel are a technical triplicate wearing a biological label. This is uncomfortable, because it means many published n=3 results are really n=1, but it is the standard the field is moving toward.
**Making the change.** You do not need to redo experiments. The data is already there; it was collapsed too early. What you need is to keep the replicate identity attached to every measurement — one extra column in your export — and stop averaging across replicates before analysis.
For figures, state in the legend exactly what n refers to: "n = 3 independent experiments; each point is one cell, coloured by experiment; large symbols are experiment means; p from unpaired t-test on experiment means." A reviewer reading that knows precisely what was done.
**What SciKeep does with this.** The imaging tools tag every segmented cell with the image and replicate it came from, so the replicate structure survives into the analysis. The SuperPlot view renders all three layers from that structure, and the reported p-value comes from the replicate means.
When there are enough replicates the tool also offers the mixed-model analysis and shows the two p-values side by side. When there are fewer than three biological replicates it declines to produce a p-value at all and explains why — because with n=1 or n=2, no test is answering the question you are asking.
**References.**
Lord SJ, Velle KB, Mullins RD, Fritz-Laylin LK (2020). SuperPlots: communicating reproducibility and variability in cell biology. Journal of Cell Biology 219(6):e202001064. doi:10.1083/jcb.202001064
Lazic SE, Clarke-Williams CJ, Munafò MR (2018). What exactly is 'N' in cell culture and animal experiments? PLOS Biology 16(4):e2005282. doi:10.1371/journal.pbio.2005282
Aarts E, Verhage M, Veenvliet JV, Dolan CV, van der Sluis S (2014). A solution to dependency: using multilevel analysis to accommodate nested data. Nature Neuroscience 17:491–496. doi:10.1038/nn.3648
Hurlbert SH (1984). Pseudoreplication and the design of ecological field experiments. Ecological Monographs 54(2):187–211. doi:10.2307/1942661