Tool validation status — what has actually been checked

Every SciKeep tool carries a verification grade, and the grade reflects what has been tested against something external — not how finished the tool looks. We publish the whole table, including the tools that score badly, because a product claiming uniform excellence is not one you should trust with a number that ends up in a paper.

The rule we hold ourselves to: a score goes up only after the test that justifies it is written, and a test in the project fails if a tool is graded "verified" without one. Jump to the verified-against table or the corrections log.

Verified against reference values (25 tools)

The tool's own code has automated checks against published reference values or hand-worked cases; every check is listed in the table below. Plain-arithmetic calculators without a test are included, and marked as having none.

Cell Counter & Image Segmentation — 9/10

MEASURED AGAIN ON 21 SEP 2026, on the current code, on five datasets; and THE DEFAULT WAS CHANGED THE SAME DAY. The low threshold of the two-level mask was raised from 0.50 to 0.85. It was chosen on BBBC039 and BBBC038 alone, by a protocol fixed before the run, then confirmed on four sets that took no part in choosing it (0.50 to 0.85): BBBC020 78.3% to 79.9%, TNBC 29.5% to 35.1%, NuInsSeg 21.2% to 22.8%, and a second NuInsSeg sample of different patches, fetched afterwards, 22.0% to 25.5%. Every set improved, by 3.1 points on average. The limit: the plain single threshold scores about as well overall on those held-out sets (81.8%, 36.6%, 22.7%, 25.9%), so this corrects a value that was too low; it does not show that a two-level mask beats a single threshold. This is the segmentation stage only (mask, seeds, watershed, minimum-area gate), not the whole interface, and every score was cross-checked against an independent implementation. F1 at IoU 0.5 at the new setting, with the plain single threshold in brackets: BBBC039 (U2OS nuclei, 200 fields, 19,389 nuclei) 87.0% (85.7%); BBBC020 (mouse macrophages, 700 nuclei) 79.9% (81.8%); BBBC038 (30-plus experiments, 670 fields, 29,461 nuclei) 71.2% (70.2%); TNBC (H&E breast cancer, 50 patches, 4,028 nuclei) 35.1% (36.6%); NuInsSeg (H&E, 62 patches across 31 organs, 3,248 nuclei) 22.8% (22.7%), and 25.5% (25.9%) on the second sample of 62 patches. Detail: BBBC039 misses 3.8%, splits 0.2%, merges 0.6%, counts 0.6% high, and predicted areas are about 16% larger than the annotated interior, which excludes the boundary ring; the change did not alter it. BBBC020 has precision 74.8%, recall 85.9% and a count 15% high. BBBC038 has precision 79.3%, recall 64.6%, 7.8% merged and a count 18% low, and is the fluorescence set to quote for will it work on MY images. TNBC has precision 25.8%, recall 54.8%, 5.3% split, 7.3% merged and a count 2.1 times too high (8,556 objects against 4,028). NuInsSeg has precision 16.8%, recall 35.7%, 9.6% split, 7.1% merged and a count 2.1 times too high, and varies widely by organ. So on H&E tissue counts are a screen, not a measurement. The benchmark applies H&E stain handling to every patch, but the tool auto-detected histology in only 9 of the first 62 NuInsSeg patches and 13 of the second 62 (45 of 50 TNBC patches), so on your own H&E images the tool may not switch to those settings by itself. WHY THE 0.50 SETTING WAS WRONG: it was chosen on 17 Aug to correct nuclei that came out about 38% too small, which was largely a symptom of a watershed defect fixed on 19 Aug. Without that defect the plain threshold gives a median predicted-to-true area ratio of 0.99 on BBBC039, and at 0.50 the two-level mask scored below the plain threshold on four of the five sets. The 3-pixel growth cap, described further down as best on all sets, is not: any cap from 1 to 5 px scores within 1.1 points on BBBC039 and BBBC038. It was not changed. WHY THE FIGURES BELOW ARE LOWER: on the samples behind them, the code as of 17 Aug scored BBBC039 63.8% (at 0.50 and today: 89.9%), BBBC038 60.0% (69.5%), BBBC020 62.9% (78.3%), TNBC 24.6% on its first 25 patches (36.0%) and NuInsSeg 12.0% (21.2%). The difference is one line in the watershed, fixed on 19 Aug: the flood stalled wherever the distance map rose and dropped about two-thirds of each rotated or elongated nucleus. Reverting that one line reproduces every old figure exactly, on all five datasets. THE COMPARISON WITH DEEP-LEARNING TOOLS WAS RE-RUN ON 21 SEP 2026, on both NuInsSeg samples of 62 patches (3,248 and 3,108 nuclei), scored with the same IoU 0.5 matching: the Cellpose nuclei model (version 3.1.1.3) scores F1 64.9% on the first sample and 68.6% on the second; StarDist 2D_versatile_he (0.9.2) scores 45.2% and 52.1%. The first sample reproduces the August figures exactly, as it should, since neither tool uses the watershed. Against the classical engine at the new setting (22.8% and 25.5%) the gap to Cellpose is about 42 to 43 points and to StarDist 22 to 27, not the 53 quoted below. The three fail differently: the classical engine finds 2.1 times too many objects; Cellpose finds 28% and 19% too few (precision 78% and 77%, recall 56% and 62%); StarDist finds 62% and 54% too few (precision 83%, recall 31% and 38%). So no method here gives a reliable H&E count, and for H&E work Cellpose is the better choice. OUR OWN TRAINED MODEL (67.2% against Cellpose 63.2% on seven held-out organs, quoted below) WAS NOT RE-RUN; it needs a training run on the full 1.6 GB NuInsSeg dataset. EARLIER MEASUREMENTS, on the code as of 17 Aug 2026. Benchmarked against TWO independent hand-annotated datasets, and improved on both. A two-level (hysteresis) mask replaced the single Otsu threshold after the benchmark showed every nucleus came out about 38% too small (median predicted/true area 0.615): Otsu separates bright cores from background and discards the dimmer rim that is still nucleus. Growth is capped at 3 pixels, a value chosen because it was the best on BOTH datasets — uncapped growth improved one and destroyed the other, which is exactly the overfitting the second dataset exists to catch. BBBC039 (1,733 U2OS nuclei): F1 at IoU 0.5 rose from 55.3% to 63.8%, missed nuclei fell from 12.2% to 6.4%, and the object count is now within 0.1% of the human count (was -2.7%). BBBC020 (700 mouse macrophage nuclei, a deliberately different morphology): F1 rose from 47.8% to 62.9%, missed fell from 6.0% to 4.0%. THE GENERALISATION TEST, 17 Aug figures, superseded by the BBBC038 re-measurement at the start of this note: BBBC038, the Kaggle 2018 Data Science Bowl collection, spans 30-plus experiments — human, mouse and fly nuclei, fluorescence, brightfield and histology stains, varied magnification and illumination. Across 1,696 nuclei sampled evenly through the collection, F1 is 60.0% (up from 58.3% before the two-level mask). Holding within four points of the single-preparation sets is the evidence the method generalises rather than being fitted to one morphology, and the 3-pixel growth cap proved independently best on this set too, having been chosen on the other two. Over-segmentation stays low: splits 0.0-0.8%. Merges are the honest weak point on diverse data at 3.3%, against 0.0-0.3% on the uniform sets. What this means in practice: counts are H&E HISTOLOGY, measured at last and the weakest result here: TNBC (Naylor et al. 2019; hand-annotated breast-cancer patches) gives F1 24.6% over 1,797 nuclei at the settings the app actually ships, and 20.5% over the full 50 patches. An earlier draft of this note said 33.9%. That figure was real but was produced with hand-tuned benchmark flags (split sensitivity 4.0, minimum area 80px) that no user gets by default, so it overstated what the tool does — the shipped-settings number replaces it, and the benchmark is meant to run at the shipped defaults (the benchmark scripts are in scripts/benchmarks and are run by hand, not by npm test) so a tuned run cannot be quoted again. Tissue is confluent — there are no gaps between cells — which is the hardest case for a watershed. Colour deconvolution (Ruifrok & Johnston 2001) is what lifted it off the floor, separating haematoxylin from eosin, because inverting the red channel makes the pink stroma bright too and the tool was finding tissue rather than nuclei; without it the same patches score 20.5%. The modality is auto-detected as histology in 45 of 50 patches, so those settings apply themselves. USE H&E COUNTS AS A SCREEN, NOT A MEASUREMENT: at F1 0.25 only about a quarter of objects match a human annotation, and a dedicated histology model (HoVer-Net, StarDist, CellPose) will do far better. It is reported because measuring it honestly is more useful than not measuring it. BROADER H&E, AND IT IS WORSE. TNBC is 50 patches of ONE tissue from one lab, which cannot separate "confluent tissue is hard" from "this preparation is hard". NuInsSeg (Mahbod et al., Sci Data 2024; CC-BY-4.0) is the same question across 31 human and mouse organs. On 62 patches sampled evenly across all 31, F1 is 12.0% (precision 9.5%, recall 16.3%) — HALF the TNBC figure, and the per-organ spread runs from 52.9% (cerebellum) to 0.0% on several tissues. So the honest H&E summary is 12-25% depending on tissue, not 25%. MEASURED AGAINST THE ALTERNATIVES, on the identical 62 patches with identical IoU 0.5 scoring: Cellpose "nuclei" reaches F1 64.9% and StarDist 2D_versatile_he 45.2%, against our 12.0%. That is a 53-point gap, and no amount of threshold tuning closes it — a watershed needs gaps between objects and confluent tissue has none. Anyone doing serious H&E nuclear morphometry should use Cellpose, StarDist or HoVer-Net, and this note says so because the measurement says so. Those tools were installed and RUN as published binaries for this comparison; no code was copied from either. THERE IS NOW A DEEP-LEARNING PATH, IT IS OPT-IN, AND THE MODEL IS OURS. Shipping Cellpose's weights was the obvious move and is not permitted: Cellpose states that ALL its models are trained on CC-BY-NC data, and NonCommercial excludes a paid product — the BSD-3 licence on their source does not change what is behind the weights, and StarDist's H&E model traces back to a non-commercial annotation set the same way. So a model was trained here from scratch on NuInsSeg (CC-BY-4.0, commercial use permitted with attribution), 665 hand-annotated H&E patches across 31 organs. IT ALSO TURNED OUT BETTER. Scored on SEVEN WHOLE ORGANS held out of training, with Cellpose run on the identical patches: ours F1 67.2% (precision 66.3%, recall 68.3%), Cellpose 63.2% (precision 78.5%, recall 52.9%), classical watershed 12.0%. Cellpose is the more precise; ours finds a third more of the nuclei that are actually there, which is what a count is for. That is not general superiority — it is that ours was trained on tissue sections specifically. The split is by ORGAN rather than by patch, because patches from one organ are near-duplicates and a random split would report memorisation; every figure here is generalisation to tissue the model has never seen. It is 3.9 MB against Cellpose's 13.3 MB. Per-organ F1 ranges from 76.5% to 45.0%, so check the overlay on your own tissue rather than assuming the mean. It is OPT-IN because it costs a one-time ~17 MB download and every other tool here runs with nothing downloaded. YOUR IMAGES ARE STILL NOT UPLOADED: inference is local, there is no server, and the privacy guarantee is unchanged. Only the SEGMENTATION step is replaced; every measurement, the learned detection-quality gate, duplicate suppression and border clearing run identically, so the two paths give comparable numbers and one set of export columns. The mask-recovery half — turning the network's vector field into objects — is our own code, written from the published algorithm and checked during development against the Python reference (not an automated test in the current suite): it reproduced 29 of 29 objects at median IoU 1.000. If the model fails to load, the classical engine runs instead rather than the analysis failing. Please cite NuInsSeg if you publish results from this path. THE SHAPE FLAG IS NOT A DAMAGE CALL, and this was measured. The tool marks objects "irregular" from circularity, and users read the colour as cell damage. Applying the old fixed cut of 0.58 to BBBC039 flagged 40.2% of OUR detections but only 0.9% of the HUMAN outlines of the very same healthy nuclei — so roughly 39 of every 40 flags were our own boundary roughness, not biology. Circularity divides by perimeter SQUARED, and a pixelated outline is punished before biology is involved. The threshold is now derived from each IMAGE'S OWN circularity distribution (median minus 3 robust SDs), because every outline in a field shares the same roughness: false positives on healthy nuclei fall from 40.2% to 0.0%, while genuine shape outliers are still caught (checked during development; no automated test). Even so, this measures SHAPE. Nothing in this tool observes membrane integrity, caspase activity or dye exclusion — confirm with a viability stain before reporting damage. Counts are good and boundaries are much better, but the 0.60-0.64 figures above predate the watershed fix (see the start of this note), and a dedicated deep-learning segmenter such as Cellpose or StarDist may still do better, so check the overlay on your own images before quoting per-cell morphometry. MULTIVERSE ANALYSIS re-runs the count across every defensible combination of settings and reports what fraction agree, so you learn whether your number survives the settings you happened to pick. BATCH ANALYSIS applies one fixed set of settings across many images grouped by condition and reports n as the number of IMAGES rather than pooled cells, so the handoff to the statistics tools is not pseudoreplicated. AUTO-DETECT reads each uploaded image and configures the tool from what is actually there: modality (fluorescence, brightfield, DAB immunohistochemistry or H&E), whether objects are darker than background and must be inverted, which channel carries the signal, typical object size, and how crowded the field is — then sets the size limits and split sensitivity to match and SHOWS you what it changed and why. Checked during development, not by an automated test, against synthetic fields of known modality and density: object radius is recovered within 25% at 5, 9 and 16 px, and foreground coverage within about 2 points from sparse (3.6%) to confluent (64.5%). Very large images are profiled from a systematic sample so the tab stays responsive; the segmentation itself still processes every pixel. Soft or out-of-focus fields are flagged, because boundaries — and therefore every area measurement — become unreliable before the count does.

Cell Culture & Doubling Time Tracker — 9/10

Fixed 23 Sep 2026: the tool's own title promised a doubling-time computation that the code never ran — the formula (Td = t·ln2/ln(Nt/N0)) existed elsewhere in the codebase with zero callers, while the tracker only ever showed a static, looked-up Cellosaurus reference value. It now computes each dish's OWN observed doubling time from its confluency readings since the last split, asserted in npm test including the exact identity case (an observed doubling returns exactly the elapsed time) and the refusal cases (no net growth, or no real elapsed time, return no number rather than a fabricated one). The published Cellosaurus value is still shown alongside it, unchanged, for comparison. LIMIT: confluency is used as a cell-count proxy, which holds pre-confluence in an evenly-spread monolayer but not for suspension culture, uneven seeding, or a dish already near confluent when the split was logged — treat a wildly-off observed rate as a measurement problem, not evidence the culture changed overnight.

PCR Primer Designer — 9/10

Tm is verified against SantaLucia 1998 nearest-neighbour values, and primers are now checked for multiple binding sites in the template you paste — including a repeated 3-prime 12 bases, which primes extension even when the full-length primer is unique. This is NOT a genome search: a primer with one site here can still bind elsewhere in the organism. Run shortlisted pairs through Primer-BLAST before ordering.

Multiple Sequence Alignment (MSA) Viewer — 9/10

Pairwise alignment now uses real AFFINE gap penalties (Gotoh 1982), asserted in npm test against hand-computed scores: identical 8-mers score exactly 16, a single substitution 13, and a 20-base deletion is aligned as ONE gap opening rather than twenty. It previously used a LINEAR penalty while the method notes described affine — which charges a long deletion by the base and so scatters small gaps where biology expects one indel. The gap-extension control was also inert: the slider moved but the function accepted only a gap-open cost. Both are now live. Progressive alignment now follows a real UPGMA guide tree (Sokal & Michener 1958) — node heights and size-weighted merges are asserted against a hand-computed tree — rather than the input order it used while displaying a "Clustal-W Lite & UPGMA" badge. Because progressive alignment is greedy, merging the most similar sequences first is what makes the early, frozen decisions the reliable ones. Scope: this is a Clustal-style progressive aligner, not an iterative refiner — for deep or highly divergent alignments prefer MAFFT or Clustal Omega.

Plasmid Map Viewer & Restriction Sites — 9/10

The full 20-enzyme panel is asserted in npm test against REBASE recognition sequences and overhangs, and every palindromic site equals its own reverse complement (a typo there is invisible in the drawing but shifts every predicted fragment size). The map does not check Dam/Dcm methylation: a site shown as cutting may be blocked in a dam+/dcm+ host. Methylation screening lives in the Sequence Toolkit.

Repeated Measures ANOVA & Mixed Model Calculator — 9/10

p-values are exact at every degrees of freedom, asserted in npm test against published t and chi-squared critical values. ANOVA is now followed by a real Tukey-Kramer HSD: the studentized range is integrated numerically and checked against published q tables, and a two-group Tukey reproduces the pooled t-test exactly. The previous version approximated q as t times root two and then applied a Bonferroni bound, which overstated every adjusted p and hid real differences. Bonferroni, Holm and Benjamini-Hochberg corrections, Welch unequal-variance t-tests, Cohen's d with the Hedges correction, confidence intervals and omega-squared are all available. Samples are screened for normality (D'Agostino-Pearson K-squared, n >= 8) as advice, not a veto. Kruskal-Wallis with tie correction runs alongside every ANOVA, with Dunn's post-hoc (Holm-adjusted) when it is significant — chi-squared p-values are asserted against textbook critical values, and a two-group Dunn reproduces Kruskal-Wallis exactly. REPEATED-MEASURES ANOVA covers the design most bench experiments actually have — the same passage or animal measured under every condition. Treating those as independent pushes the shared subject-level swing into the error term and hides real effects, so the subject is removed as its own term; Mauchly sphericity is checked and the Greenhouse-Geisser correction applied, with the uncorrected value shown beside it. It is pinned by the exact identity F = t-squared against a paired t-test, which fixes the entire sum-of-squares partition rather than matching one table value. A LINEAR MIXED-EFFECTS MODEL (random intercept, REML) handles what repeated measures must refuse: missing cells and unbalanced designs. A subject that lost one measurement keeps contributing the rest instead of being discarded entirely. Its omnibus test refits by maximum likelihood, because REML likelihoods from models with different fixed effects are not comparable — a standard and silent error. On balanced complete data its condition effect equals the plain difference in condition means (asserted in npm test). It has not been compared with R lme4 output. Still absent: crossed or multiply-nested random effects, and Satterthwaite degrees of freedom (the conservative between-subject count is used instead). For those, use R and lme4.

Sample Size & Statistical Power Calculator — 9/10

A MONTE CARLO SIMULATOR now runs beside the formula, because the closed form assumes normal data, equal variance and one independent observation per subject — and bench biology routinely breaks all three. It is validated against the formula where the formula is valid (n = 26 per group at d = 0.8 returns 79.9% against G*Power's 80% — measured during development, not an automated test), which is what licenses trusting it where the formula does not apply. Every run is seeded and the seed is reported, so a number you put in a grant can be reproduced exactly. It also quantifies the winner's curse: at n = 6 and d = 0.5 the power is 11% and the effect you would publish is inflated about 2.8x. MOST IMPORTANTLY it simulates NESTED designs. Measuring 100 cells from each of 3 dishes is n = 3, not n = 300 — and analysed as 300 the design looks well powered while, on data with NO real effect, that pooled test still declares significance around 80% of the time against the 5% you asked for. Pooling does not buy power; it manufactures findings. Sample sizes match G*Power exactly for a two-tailed independent t-test at alpha 0.05 and 80% power (d = 0.2/0.5/0.8/1.0 gives 394/64/26/17 per group), asserted in npm test. The textbook normal approximation is biased LOW by one subject per group across that whole range, always toward an underpowered study; Guenther's correction is applied.

Standard method (17 tools)

The published standard method, correct on review, with no external reference test yet. Verify against your own positive control before quoting absolute values.

Western Blot Densitometry Analyzer — 7/10

Band detection and quantification are tested (npm test) against synthetic lanes with known band amounts: bands are found by position rather than assumed, background is subtracted per lane, and normalisation arithmetic is checked by hand (double the loading, half the fold). Measured limits, on Gaussian bands of known integrated density: a band less than about 10% above background is NOT DETECTED rather than under-measured, and a detected band recovers about 9-10% low — a fixed cost of the integration window, the same for a faint band as for one 22x stronger. Saturated bands are flagged rather than quantified. Re-expose faint or saturated blots instead of trusting the number. (The note previously claimed faint bands were under-measured by roughly 20%; no code implemented that and no test measured it.)

Flow Cytometry Panel Designer — 8/10

Every ordered pair in a panel is assessed — asserted in npm test. The previous table covered 19 of 132 pairs and reported the other 86% as clean simply because they were absent; BV605 and BV711 could never raise a warning at all. Built-in reference spillover values (no source is cited for them) are labelled as such and kept separate from spectral estimates. YOUR INSTRUMENT IS AN INPUT, because spillover is not a property of a dye pair: it is the fraction of the donor's emission landing in the acceptor's DETECTOR, and a detector is a bandpass filter on a specific machine. Two peaks 90 nm apart spill heavily through a wide filter and barely at all through a narrow one. Choose your optical configuration and the tool integrates the donor emission over the acceptor's actual filter band, so the numbers change with the instrument — which is the honest behaviour, since the physical quantity does too. Asserted in npm test: the textbook FITC-into-PE spillover is detected, it is correctly ASYMMETRIC (emission tails run red, so the reverse is far smaller), and results genuinely differ between instrument configurations. What remains an approximation, and is labelled as such: the emission profile is parametric — asymmetric with a long red tail — not a measured spectrum, which would mean shipping licensed per-dye spectral curves. The filter bands themselves are the manufacturers' published layouts. Treat the output as a ranking of what to check, then run single-stain controls.

CRISPR gRNA Designer — 8/10

Off-target sites are now genuinely SEARCHED and SCORED. Both strands of the sequence you submit are scanned for NGG-PAM sites within 4 mismatches, and each is scored with the published Hsu-Zhang 2013 position weights — asserted in npm test to return exactly 100 for a perfect match, exactly (1 - w) x 100 for a single mismatch at each weighted position, and to fall monotonically as mismatches accumulate. This is NOT genome-wide: it covers the sequence you paste, so a guide that looks clean here can still cut elsewhere in your organism. Guide discovery and ranking are asserted in npm test: malformed spacers are rejected rather than given a passing grade, poly-T is penalised hardest because it truncates Pol III transcription, extreme GC ranks below balanced guides (repetitive seeds are penalised too, though that is not separately tested), and scoring is deterministic and bounded. The score orders candidates — it is NOT a predicted cutting efficiency, and NO genome-wide off-target search is performed. Run shortlisted guides through Cas-OFFinder, CRISPOR or Benchling before ordering oligos.

CRISPR Indel Analysis & Frameshift Caller — 8/10

Indel calling is asserted in npm test: deletions of 1-9 bp and insertions are sized exactly, a single clean deletion is not split into scattered gaps, substitutions are not counted as indels, the result does not change with how much flanking sequence you paste OR with the two reads being trimmed differently from each other (end gaps are free — until 24 Sep 2026 they were not, and a read trimmed 12 bp was reported as a 12 bp deletion), and a negative control returns "no edit" rather than a floor value. It deliberately reports NO editing efficiency — that is a population property and needs a Sanger chromatogram, not consensus text. CHECKED 24 SEP 2026 whether the original ICE paper (Synthego) publishes a redistributable validation set to test against directly: it does — github.com/synthego-open/ice ships real .ab1 traces with an asserted expected result (their own test suite: results['ice'] == 77 for their "good_example" control/edited pair). It is not usable here even so: this tool takes plain text sequences, and ICE's own number comes from decomposing overlapping chromatogram peaks in a mixed edited population — exactly the computation this tool refuses to attempt, and exactly why a edited sample's own consensus base-calls are unreliable right at the cut site. Re-implementing that decomposition to chase one more digit of confidence is the kind of thing this tool's own refusal exists to avoid; the honest answer is that this tool and ICE do not measure the same thing, not that one needs to catch up to the other.

siRNA Knockdown Designer — 8/10

The microRNA seed screen was previously DEAD (RNA/DNA mismatch, so nothing ever matched), then fixed but still untested — and the fix was incomplete: asserted in npm test on 24 Sep 2026, writing the reference test turned up that 4 of the 5 stored seeds were themselves wrong (verified against miRBase's own published mature sequences). The check ran, compared two real strings, and got the biology wrong for miR-21-5p, let-7a-5p, miR-34a-5p and miR-126-3p — only miR-17-5p's stored seed was ever correct. All five now hold the real positions-2-8 seed of their cited miRBase accession. GC-window filtering (30% floor, a ceiling that scales with the floor to avoid inverting on the tool's 0-70% slider) is also asserted in npm test. Scope is 5 well-characterised human seeds, not the full miRBase catalogue — a clean result narrows risk, it does not clear the siRNA. The overall candidate score is a ranking heuristic, not validated against measured knockdown efficiency — design 3-4 independent siRNAs per transcript.

Methods Section Writer — 8/10

Generates DRAFT text from the parameters you entered, and — more usefully — tells you what the draft is MISSING. Every draft is checked against the ARRIVE 2.0, MIQE and SAMPL items that apply to it, and each gap names the guideline, its item number and what to write. This replaced a fixed "Compliance Validation Summary" that asserted RRIDs and catalogue numbers were included whether or not they were, congratulating the user for compliance it had never checked. The core limit is unchanged: it reports what you told it, so a wrong input becomes a wrong sentence stated with full confidence, and this text is destined for a manuscript where an unchecked number is the most expensive kind of error. It does NOT verify that the described method is what you performed, and it invents no citations. Read every sentence against your own records before submission.

Protocol & SOP Writer — 7/10

A fixed library of twelve reference protocols; it takes no input from you. The volumes, timings and concentrations are typical starting points written into the library, not validated for your cells, reagents or safety requirements. Have someone who has run the assay read it, and check each value against the manufacturer, before anyone follows it at the bench.

Pre-Submission Manuscript Audit — 8/10

Conformance against PUBLISHED, citable guidelines rather than our own list: the ARRIVE 2.0 Essential 10 (Percie du Sert et al., PLOS Biology 2020) for animal research, the essential detectable subset of MIQE (Bustin et al., Clinical Chemistry 2009) for qPCR, and SAMPL for statistical reporting. The applicable guideline is inferred from your text, so ARRIVE items are not thrown at an in-vitro paper. ARRIVE items carry their official numbers (MIQE and SAMPL do not number items, so those carry a topic), each links to the source, and every PASS quotes the sentence that triggered it so you can disagree with it. Asserted in npm test: detection routes correctly by manuscript type, a bare text fails, and every PASS quotes text that is really in the manuscript. IT REMAINS A CHECKLIST, NOT A REVIEW: it detects whether a topic is ADDRESSED. It cannot judge whether your statistics are appropriate, whether your sample size is justified, or whether your blinding worked. A clean result means nothing obvious is missing, not that the manuscript is sound.

Verified against — what each check compares with, and what it does not cover

Each row names a test in the project. A check counts only if the test runs the tool's own code, and the expected value comes from a published source or from arithmetic done by hand before the code was run.

Cell Counter & Image Segmentation — 9/10

  • The burned-in scale bar is detected and measured, so areas use the image's real calibration rather than an assumed one (Constructed test data: Constructed images with a bar of known length)
  • 12-bit sensor data stored in a 16-bit TIFF container is rescaled correctly, and images that are already correct are left alone (Constructed test data: Constructed TIFFs)
  • An image too large to decode is refused from its header, before any pixel data is read (Required behaviour: Design requirement)
  • Wound-healing, spheroid, puncta, neurite and scatter-gating analyses compute from real pixels rather than fixed demo numbers, and never emit an invented p-value (Constructed test data: Constructed images with known geometry)
  • Segmentation on BBBC039 (real hand-annotated U2OS nuclei, 20 fields / 1,733 nuclei) does not regress below F1 85% at IoU 0.5 — measured 24 Sep 2026 at 89.9% on this exact subset. Skipped automatically wherever the ~80 MB dataset isn't downloaded (it is not vendored), but runs as part of the normal test suite wherever it is (Published reference: Caicedo et al., Broad Bioimage Benchmark Collection BBBC039, hand-annotated ground truth)

Not covered: Five benchmarks were re-measured on the current code on 21 Sep 2026, after the two-level mask low threshold was raised from 0.50 to 0.85 (F1 at IoU 0.5, full sets: BBBC039 87.0%, BBBC020 79.9%, BBBC038 71.2%, TNBC 35.1%, NuInsSeg 22.8%, and 25.5% on a second sample; segmentation stage only, scored against hand annotation). The value was chosen on two sets and confirmed on four others, but a plain single threshold scores about as well on the held-out sets. On H&E tissue the tool over-counts by about 2 times and may not auto-detect the stain, so H&E counts are a screen only. Cellpose and StarDist were re-run on 21 Sep 2026 on two NuInsSeg samples (Cellpose F1 64.9% and 68.6%, StarDist 45.2% and 52.1%, against the classical engine at 22.8% and 25.5%). Our own trained H&E model was not re-run. Only BBBC039's 20-field subset runs automatically (see above) — the full 200-field BBBC039 run and the other four datasets (BBBC020, BBBC038, TNBC, NuInsSeg) are still run by hand only. Check the overlay on your own images.

Western Blot Densitometry Analyzer — 7/10

  • A band under about 10% above background is not detected at all, and a detected band recovers within 15% of its true integrated density — the same shortfall for a faint band as for one 22x stronger (Constructed test data: Gaussian bands of known integrated density (amplitude x sigma x sqrt(2pi)))
  • Per-lane p-values against the control are Holm-adjusted across the lanes tested, so the significance stars account for the number of comparisons (Required behaviour: Holm-Bonferroni; the same correction the group chart uses)
  • Ladder calibration, replicate statistics, normalisation without stand-in divisors, Holm-adjusted comparisons and SD/SEM/95% CI error bars behave as specified (Worked by hand: Hand-computed values on constructed lanes)

Not covered: Band densities are measured on synthetic lanes, not benchmarked against ImageJ on real film or ChemiDoc images. Saturated bands are flagged, not quantified.

qPCR ΔΔCt Calculator — Biological vs Technical Replicates — 9/10

  • 2^-ΔΔCt returns exactly 2.0 for ΔΔCt = -1, 1.0 for the baseline against itself, and treats a 2-fold fall as the mirror of a 2-fold rise (Worked by hand: Livak & Schmittgen 2001, worked by hand)
  • One biological replicate produces no p-value, with the reason stated; the fold change is still reported (Required behaviour: Pseudoreplication rule: n is biological replicates, not wells)
  • n is counted from biological replicates (n = 3, not 9 wells) and the p-value follows (Worked by hand: Hand count on a 3 x 3 design)
  • Pfaffl mode raises target and reference each to their own efficiency (1.9^2 / 2.1^1 = 1.719), and reduces exactly to Livak at 100% efficiency (Worked by hand: Pfaffl 2001 Nucleic Acids Res 29:e45, equation 1, worked by hand)

Not covered: No published qPCR dataset is used as a reference; the expected values are hand-computed from the Livak and Pfaffl equations. Reference-gene stability (geNorm M) uses unweighted Ct even in Pfaffl mode.

qPCR Efficiency Calculator (Standard Curve) — 9/10

  • E = 10^(-1/slope) - 1 gives exactly 100% at the textbook slope of -3.3219, and the MIQE window maps to the slopes the page quotes (Published reference: MIQE guidelines (Bustin et al. 2009); the definition of amplification efficiency)

Not covered: The fit itself is checked on constructed scatter, not on a published standard curve.

Nucleic Acid Concentration Calculator (A260) — 9/10

  • Concentration uses the 1 cm conversion factors (dsDNA 50, ssDNA 33, RNA 40 ng/µL per A260) and the dilution factor; purity ratios are left null when unmeasured (Published reference: Beer–Lambert; the standard A260 conversion factors)

Not covered: Instrument-specific path-length settings are not modelled.

4PL IC50 & Dose-Response Calculator — 9/10

  • Z' = 1 - 3(σp + σn)/|μp - μn| for controls 10/12/11 and 100/98/102 gives 0.899; overlapping controls are labelled unacceptable (Worked by hand: Zhang, Chung & Oldenburg 1999, worked by hand)
  • A 4PL fit to noiseless data with IC50 = 1 and Hill = 1 recovers both (Constructed test data: Constructed curve (ground truth known by construction))

Not covered: Fit confidence intervals, ROUT outlier removal and the 5PL model are not checked against reference output (GraphPad, R drc). The 4PL check uses noiseless data.

ELISA & BCA Standard Curve Calculator — 9/10

  • Every standard reads back off the fitted curve to its own concentration within 5%, including the TOP standard, and a reading beyond an asymptote is refused rather than returned as a number (Constructed test data: Round trip through a known 4PL (a 0.05, d 3.0, c 100, b 1.0))
  • OLS linear regression (BCA/Bradford) matches a hand-computed slope, intercept and R² on an asymmetric small dataset (Worked by hand: Worked by hand: x = 1,2,3,5 / y = 2,3,5,6 gives slope 36/35)
  • A 4PL fit to noiseless data from a known (a, d, c, b) curve recovers the bottom asymptote exactly (the conc=0 blank point reduces the equation to y=a) and the top asymptote, slope and inflection concentration within the grid search's own resolution (Constructed test data: Constructed curve (ground truth known by construction), same technique as the plate tool's own 4PL check)

Not covered: The 4PL grid search is coarse (b steps of 0.2, c steps of x1.5) and assumes the top standard concentration is roughly the same order of magnitude as the true inflection point — a top standard that does not reach near the plateau will under-recover the top asymptote, which is disclosed rather than hidden. Not checked against GraphPad or R drc output.

Drug Synergy Calculator — Bliss Independence Score — 9/10

  • Bliss expected effect: two 50% effects expect 75%; 30% and 40% expect 58%; observed minus expected is the score (Worked by hand: Bliss independence, worked by hand)

Not covered: Only Bliss is computed. There is no Loewe additivity and no Chou-Talalay combination index.

Cell Culture & Doubling Time Tracker — 9/10

  • Every curated doubling time sits inside the range its own Cellosaurus record publishes, and each accession belongs to the line it is named for (Published reference: Saved Cellosaurus records (api.cellosaurus.org, 24 Sep 2026))
  • Confluence bands (Monitor/Check Tomorrow/Split Soon/Critical) switch at exactly 50, 80 and 90% with no off-by-one (Worked by hand: The tool's own documented thresholds, worked by hand at each boundary)
  • An exact doubling (confluency 15% to 30%) over a known elapsed time returns that same elapsed time as the doubling time — the Td = t identity when Nt/N0 = 2 (Mathematical identity: Td = t·ln2/ln(Nt/N0), worked by hand)
  • No net growth since the last split, or no real elapsed time, returns no observed doubling time rather than a fabricated one (Required behaviour: Required behaviour: never invent a rate the data cannot support)

Not covered: Confluency stands in for cell count, which holds pre-confluence in an evenly-spread monolayer but not for suspension culture, uneven seeding, or a dish already near confluent at the last logged split. The published Cellosaurus doubling time is a static reference, not independently re-measured here.

Flow Cytometry Panel Designer — 8/10

  • A panel of n dyes returns all n(n-1) ordered pairs, none omitted; FITC into PE is flagged, and the effect is asymmetric; results change with the instrument configuration (Required behaviour: Design requirement (no pair assumed safe); textbook FITC/PE behaviour)

Not covered: Emission is a parametric model, not a measured spectrum. The built-in reference spillover values carry no source citation. Treat the output as a ranking of what to check, then run single-stain controls.

PCR Primer Designer — 9/10

  • The nearest-neighbour parameter table matches the paper, and Tm agrees with the published equation (including the salt correction on entropy) for four sequences (Published reference: SantaLucia 1998 (PNAS 95:1460), Table 1 and eq. 7, typed independently)
  • Mg²⁺ is converted to monovalent-equivalent concentration as [Na⁺] + 4·√[Mg²⁺] (Worked by hand: von Ahsen et al. 2001)

Not covered: Tm is not compared with another primer tool's output. The single-template binding-site check has no automated test, and no genome-wide specificity search is performed — run shortlisted pairs through Primer-BLAST.

Sanger Trace Chromatogram Viewer — 8/10

  • Basecalls are read from the ABIF specification's own field offsets (numelements +12, datasize +16, dataoffset +20 into each directory entry), cross-checked against the header's own root entry (Published reference: The ABIF format specification; a fixture built to it, not to the parser)
  • A datahandle of 0 is never mistaken for the data offset — the defect that made every genuine .ab1 fail to parse (Published reference: The ABIF specification)
  • A plain FASTA file is accepted as a sequence with its header stripped (Constructed test data: A constructed FASTA file with a known sequence)
  • Binary that is not a trace file is refused rather than mined for letters that happen to look like bases (Required behaviour: Required behaviour: an unreadable file must not produce a plausible-looking wrong answer)

Not covered: Mean Phred / %Q30 and reference-variant calling are not checked against a hand-computed or published case yet.

CRISPR gRNA Designer — 8/10

  • Off-target scoring uses the published Hsu-Zhang position weights: a perfect match scores 100, a single mismatch at any position scores (1 - w) × 100, and two mismatches follow the published formula (Published reference: Hsu et al., Nature Biotechnology 2013)
  • Malformed spacers are rejected, poly-T is flagged and ranked down, extreme GC ranks below balanced GC, scores are deterministic and within 0-100 (Required behaviour: Design requirements)
  • Cloning oligos follow the published U6 vector recipe: top CACC[G]N20, bottom AAAC[rc N20][C], with the extra G added only when the guide does not already start with G (Published reference: Sanjana et al. 2014 (lentiCRISPRv2) / Zhang lab pX330 protocol)

Not covered: The rank score is not a predicted cutting efficiency. Off-target search covers only the sequence you paste — not the genome. Cloning oligos are given for SpCas9 20-mers only; SaCas9 and Cas12a guides need their own vectors.

CRISPR Indel Analysis & Frameshift Caller — 8/10

  • Deletions of 1-9 bp and insertions are sized exactly; substitutions are not counted as indels; the call does not change with flanking sequence; an unedited sequence is "no edit" (Constructed test data: Constructed edits of known size)
  • End gaps are free, so a read trimmed differently from the reference reports "no edit", and a real 3 bp in-frame deletion keeps its size and frame whether or not the reads are trimmed (Constructed test data: Constructed reads trimmed at one or both ends)

Not covered: Reports no editing efficiency: that needs a Sanger chromatogram, not consensus text. Frame is taken from the pasted sequence, so a frameshift call assumes position 1 is a codon start and the amplicon is coding.

Multiple Sequence Alignment (MSA) Viewer — 9/10

  • Identical 8-mers score exactly 16 and one substitution 13; a 20-base deletion is ONE gap opening, not twenty (Worked by hand: Hand-computed scores (match +2, mismatch -1, affine gaps); Gotoh 1982)
  • The UPGMA guide tree has the right heights and weights merges by cluster size, on two hand-computed trees (Worked by hand: Sokal & Michener 1958, worked by hand)

Not covered: Progressive alignment is greedy and is not compared with Clustal Omega or MAFFT output; for deep or divergent alignments use those.

Plasmid Map Viewer & Restriction Sites — 9/10

  • All 20 panel enzymes have the REBASE recognition sequence and overhang (Published reference: REBASE)
  • Sites are found at known 1-based positions, including a site that spans the origin of a circular plasmid, and frequent cutters are separated from single cutters (Constructed test data: Constructed sequences with sites at known offsets)

Not covered: The map does not model Dam/Dcm methylation blocking, so a methylation-sensitive site is shown as cutting regardless of host strain.

Gibson & HiFi Assembly Planner — 9/10

  • Golden Gate parts are screened for the chosen enzyme's own recognition site on BOTH strands, and a part carrying one invalidates the assembly (Published reference: REBASE recognition sequences for BsaI, BsmBI, BbsI and PaqCI)
  • The Gibson-assembly overlap Tm returns exactly the same value (rounded to 1 decimal) as the shared, already-tested calculateTm engine under the documented conditions (50 mM Na+, 500 nM primer) for two different real sequences, and refuses (0) rather than reporting a Tm for a sequence with fewer than 2 real bases (Mathematical identity: Identity with calculateTm, already verified against SantaLucia 1998 (see primer))

Not covered: Fragment ordering is taken as given; the planner does not design overhangs or domesticate parts for you. Gibson overlap length is your choice, not solved for a target Tm.

3D Macromolecular PDB Viewer — 8/10

  • PDB coordinates are read from real records, accession numbers are validated, and pLDDT is shown only where it exists (Constructed test data: Constructed PDB text with known atoms)
  • The idealised B-DNA and α-helix models reproduce their published parameters (3.38 Å rise and 36° twist per base pair; 1.5 Å and 100° per residue) and are labelled as models (Published reference: Standard B-DNA and α-helix geometry)
  • Only the first conformer of an alternate-location residue is kept, so a residue has one backbone atom; insertion codes (52, 52A, 52B) count as separate residues (Published reference: PDB format columns 17 and 27)
  • An AlphaFold model uploaded as a file is recognised from its TITLE, so pLDDT colouring still applies, and a file without HELIX/SHEET records reports that it has none (Published reference: The header of a real AlphaFold PDB file)

Not covered: Rendering is not tested beyond metadata. Secondary structure is read from the file's HELIX/SHEET records, not computed from coordinates, and is not checked against DSSP. Surface calculation is not checked against another program.

DNA Reverse Complement, Translate & ORF Finder — 8/10

  • The codon table used for translation is the real, complete standard genetic code (NCBI table 1): the correct start and all three stop codons, all 64 codons defined, and the six-fold degenerate Leu/Ser/Arg families resolved correctly (Published reference: NCBI genetic code table 1)
  • Reverse complement is correct on a palindromic restriction site, a hand-worked non-palindromic case, and preserves letter case (Worked by hand: Hand-worked complementation)
  • Frame-1 translation of a constructed ORF matches its known peptide exactly, correctly stops at the first stop codon when asked to, and never translates a trailing partial codon (Constructed test data: A constructed ORF with a known translation)

Not covered: Only the reverse-complement and frame-1-translate tabs are tested. GC content, six-frame ORF-finding, restriction digestion (which reuses the same REBASE-sourced enzyme data as plasmid-viewer, but has its own untested fragment-length arithmetic), Dam/Dcm methylation screening and codon optimisation against usage tables have no automated check.

siRNA Knockdown Designer — 8/10

  • All five stored microRNA seeds are the real positions-2-8 seed of their named miRBase mature sequence (independently re-derived from the cited accession in the test, not copied from the source file) — catches the bug found 24 Sep 2026 where 4 of 5 stored seeds were wrong (Published reference: miRBase mature sequences (MIMAT0000076, MIMAT0000062, MIMAT0000070, MIMAT0000255, MIMAT0000445))
  • A guide carrying each real miRNA seed is correctly flagged as a match (in both DNA and RNA input form), a guide with none of the five seeds is not flagged, and a guide matching the OLD wrong let-7a window is no longer flagged (Published reference: The five verified seeds above)
  • GC% is computed correctly, the GC ceiling scales with the floor so the window can never invert on the tool's 0-70% slider, and candidates outside the floor/ceiling are excluded (Worked by hand: Hand-computed GC percentages and window arithmetic)

Not covered: The overall candidate score (asymmetry + GC + seed-GC terms, minus a miRBase-match penalty) is a ranking heuristic with no knockdown-efficiency ground truth to validate it against — it orders candidates, it does not predict silencing efficacy.

Protein MW, pI & Extinction Calculator — 10/10

  • Hen egg-white lysozyme is 129 residues, 14,313 Da and ε280 = 37,970 M⁻¹cm⁻¹ (37,470 with no disulfides), as ProtParam and UniProt report (Published reference: UniProt P00698; ExPASy ProtParam)
  • Human insulin A + B chains with three disulfides is 5,807.6 Da (Published reference: UniProt P01308 / DrugBank)

Not covered: Isoelectric point is NOT checked against ExPASy: it uses the Lehninger pKa set, not Bjellqvist, so it can differ by a few tenths.

Oligo & Primer Resuspension Calculator — 9/10

  • Tm comes from the same tested engine as the primer designer (Published reference: SantaLucia 1998 (see primer))
  • ε260 sums one published mononucleotide set (dA 15200, dC 7050, dG 12010, dT 8400) (Published reference: LGC Biosearch / Eurogentec extinction coefficients)
  • Molecular weight follows the anhydrous 5'-OH residue-sum formula (Worked by hand: Formula worked by hand)
  • Resuspension and dilution volumes are inline arithmetic (Inline arithmetic, no automated test: The formulas)

Not covered: ε260 is a nucleotide sum with no nearest-neighbour hypochromicity, so it runs roughly 5-10% high; resuspension and dilution arithmetic has no automated test.

Dilution & Molarity Calculator (C₁V₁ = C₂V₂) — 10/10

  • C₁V₁ = C₂V₂ gives 50 µL of a 10 mM stock for 10 mL at 50 µM, and 5 µL of 100 mM for 20 mL at 25 µM (Worked by hand: The formula, worked by hand)
  • A target above the stock, a zero stock or a zero volume is refused rather than shown as a negative or infinite volume (Required behaviour: Required behaviour)

Not covered: Units are fixed (stock mM, target µM, volume mL); there is no unit conversion or serial-dilution planner.

Buffer & Solution Recipe Generator — 8/10

  • PBS's four component masses (NaCl, KCl, Na2HPO4, KH2PO4) are each within 2.5% of molarity x molecular weight for their stated millimolar concentrations, and scale linearly with both volume and concentration multiple (Worked by hand: mass_g/L = molarity_mol/L x MW_g/mol, standard IUPAC molar masses)
  • TBS's two component masses (Tris base, NaCl) are each within 2.5% of molarity x molecular weight for their stated millimolar concentrations (Worked by hand: mass_g/L = molarity_mol/L x MW_g/mol, standard IUPAC molar masses)

Not covered: Only PBS and TBS are backed by a real molarity-x-MW derivation and a test. RIPA, HEPES, Tris, agarose, SDS running/transfer, TAE, TBE, NaCl and EDTA recipes remain hard-coded per-litre masses with no automated check.

PCR Master Mix Calculator — 10/10

  • Per-reaction volumes × reactions (+ overage) written inline in the component (Inline arithmetic, no automated test: The formula)

Not covered: No automated test on this tool. It is multiplication you can check by hand.

Lentiviral MOI & Titre Calculator — 10/10

  • At MOI 1, 63.2% of cells receive at least one particle and 26.4% two or more (Poisson); 1e5 cells at MOI 5 from a 1e8 TU/mL stock need 5 µL, and 5 µL delivers MOI 5 back (Worked by hand: Poisson transduction model, worked by hand)

Not covered: Assumes the titre you enter is functional (TU/mL), not physical particles.

Transfection Calculator — DNA & Reagent Volumes — 9/10

  • Lipofectamine 3000 amounts reproduce the manufacturer protocol table (96-, 24- and 6-well), with P3000 at 2 µL per µg DNA in every vessel (Published reference: Thermo Fisher Lipofectamine 3000 protocol and scaling table)
  • The growth areas used for scaling reproduce the manufacturer multiplication factors from 96-well to T-175 within 2% (Published reference: Thermo Fisher Lipofectamine 3000 scaling table)
  • Scaling an optimised condition multiplies every quantity by the area ratio and keeps the reagent-to-nucleic-acid ratio (Worked by hand: Area ratio, worked by hand)

Not covered: Only Lipofectamine 3000 has built-in amounts; for other reagents the starting condition is yours. Transfection efficiency is not predicted.

Cell Line & Mycoplasma Testing Registry — 9/10

  • A line is "Clear" below 30 days since its last mycoplasma test, "Test Due" from exactly 30 days, "OVERDUE" from exactly 60 days, and "Never tested" with no test date at all — every boundary checked on both sides (Worked by hand: The tool's own documented 30/60-day cadence, worked by hand at each boundary)
  • Elapsed days since a test date are computed correctly, including the same-day (0 days) case (Worked by hand: Hand-computed elapsed time)

Not covered: This is inventory/status bookkeeping, not a measurement — there is no external reference for "30 days" beyond the tool's own stated cadence, which is common practice, not a regulatory requirement.

Repeated Measures ANOVA & Mixed Model Calculator — 9/10

  • Student t and chi-squared p-values match published critical values (Published reference: Standard t and chi-squared tables)
  • A balanced one-way ANOVA matches a hand-computed example, and F = t² for two groups (Mathematical identity: Hand calculation; the F = t² identity)
  • Tukey critical values match published studentized-range tables at five (k, df) points (Published reference: Published q tables, α = 0.05)
  • Two-group Tukey reproduces the pooled t-test; repeated-measures ANOVA on two conditions gives F = t² of the paired t-test; two-group Dunn reproduces Kruskal-Wallis (Mathematical identity: Mathematical identities)
  • Holm, Bonferroni and Benjamini-Hochberg adjustments obey their defining orderings (Worked by hand: Definitions of the procedures)
  • Mixed-effects model on balanced complete data returns the raw difference in condition means (Mathematical identity: GLS = OLS means for balanced data)

Not covered: Greenhouse-Geisser and Mauchly sphericity, and the mixed model on unbalanced data or against another package (R/lme4), are not checked against reference output. Satterthwaite degrees of freedom are not implemented.

Sample Size & Statistical Power Calculator — 9/10

  • Two-sample t-test sample sizes match G*Power: d = 0.2 / 0.5 / 0.8 / 1.0 need 394 / 64 / 26 / 17 per group at α = 0.05, power 0.80 (Published reference: G*Power reference values)
  • Two-proportion sample size p1 = 0.5 vs p2 = 0.3 is 93 per group (Published reference: Fleiss two-proportion z-test textbook case)
  • Required events for a log-rank test equal Schoenfeld: 4(z_α + z_β)² / ln(HR)² (Published reference: Schoenfeld closed form)

Not covered: The Monte Carlo simulator (nested designs, winner's curse) is not covered by an automated test; its agreement with the formula (79.9% vs G*Power's 80% at n = 26, d = 0.8) was measured during development.

Kaplan-Meier Survival Curve Generator — 9/10

  • Product-limit survival for times 1, 2, 3 (censored), 4, 5 is 0.8, 0.6, 0.3, 0; median survival is 4 (Worked by hand: Kaplan & Meier 1958, worked by hand)
  • The log-rank p-value matches the exact chi-squared tail: events at 1,3 vs 2,4 give p = 0.433; 1,2 vs 3,4 give p = 0.0895 (Worked by hand: Mantel-Haenszel log-rank, worked by hand)

Not covered: Confidence bands and hazard ratios are not checked. Compared with R survival, only the hand-worked cases are covered. Before Sep 2026 the log-rank p-value was wrong — see the corrections log.

RNA-seq Volcano Plot Visualizer — 9/10

  • A gene is called "up" or "down" only with STRICT inequality on both the padj and log2FC cutoffs — sitting exactly at either threshold is "not significant", not a coin flip on floating-point rounding (Worked by hand: Design requirement, worked by hand at both boundaries and just past them)
  • -log10(padj) matches hand-computed values (padj=0.05 -> 1.301, padj=0.001 -> 3) (Worked by hand: Hand calculation)

Not covered: CSV column detection (matching header names like "padj"/"FDR"/"log2FC") and which genes get labelled on the plot are not tested.

Electronic Lab Notebook (ELN) — 9/10

  • The audit chain is recomputed and the interface reports the real result; altering a record's content, title or order breaks verification at that record (Required behaviour: Required behaviour of a hash chain)
  • The chained payload includes a SHA-256 of the entry text, so two records differing only in content have different hashes (Required behaviour: Required behaviour)
  • The hash function backing the audit ledger's tamper-evident chain really is SHA-256: it reproduces both of FIPS 180-4's own published test vectors (the digest of "abc" and of the empty string) exactly, and always returns a 64-character hex digest (Published reference: FIPS 180-4 Appendix B.1, the standard's own worked examples)
  • Changing any single field of a chained audit record changes its hash, which breaks the prevHash link the next record in the chain depends on — the mechanism that actually reveals tampering (Required behaviour: Required behaviour of a hash chain)

Not covered: Signing needs only a typed name and a checkbox: there is no second authentication component and no identity verification, so this does NOT meet 21 CFR 11.200. The ledger lives in your own browser storage, so it is tamper-evident, not tamper-proof — someone with the console can rewrite the whole chain. Entry CRUD and search/tagging are not tested.

-80C Freezer Inventory & Cryo Box Locator — 8/10

  • Items saved with values the old Add form invented (100 ng/µL, 50 µL, "Current Researcher") are cleared; values a user typed are kept (Required behaviour: Required behaviour)

Not covered: A record-keeping tool: positions and details are what you enter, and nothing checks them against the physical freezer.

Antibody Validation Tracker — 8/10

  • The RRID lookup parses the real SciCrunch resolver response, and each of the five offline records equals what the registry returns for that RRID (Published reference: Saved scicrunch.org/resolver responses (24 Sep 2026))
  • An RRID the registry cannot resolve returns nothing, rather than another antibody's details under the user's RRID (Required behaviour: Required behaviour)

Not covered: Inventory, validation status and dilution records are user-entered and not checked. The live resolver is only exercised through a saved response.

Publication Figure Panel Composer (300 DPI) — 9/10

  • Every journal preset width is the one that journal publishes: Nature 89/120/183, Cell 85/114/174, Science 57/121/184, PNAS 87/114/178 mm (Published reference: Each journal's own figure-preparation guidelines)
  • Scale-bar length is computed from the source image's own width and the pixel size you enter, and is refused when either is unknown or the bar would not fit (Worked by hand: bar px = (microns / micron-per-pixel) x (panel px / source px), worked by hand)
  • Canvas pixel dimensions follow the standard mm-to-pixel conversion (px = mm/25.4 x DPI) for real published journal column widths (Nature 89mm 1-col at 300 DPI, Cell 2-col at 600 DPI), and round-trip back to the physical width within 1mm rounding (Worked by hand: The standard millimetre-to-pixel/inch conversion, against Nature and Cell's own published submission widths)
  • The panel grid layout accounts for every margin, gap, header and caption band exactly once with no double-counted or missing space, for both a single panel and a multi-row/column grid (Worked by hand: Hand-derived layout identity: total height = 2 margins + header + N panels + N captions + (N-1) gaps)

Not covered: The actual SVG/canvas rendering (scale bars, channel badges, annotation overlays) is not tested, only the dimensional arithmetic feeding it.

Methods Section Writer — 8/10

  • The prompt forbids inventing values (centrifuge speeds, incubation times, dilutions) the user did not supply, and generated text carries a draft notice in the interface (Required behaviour: Design requirement, asserted on the prompt and the interface text)
  • The built-in templates for all twelve experiment types state only the fields you entered: no compliance claim, catalogue number or software version comes from the template (Required behaviour: Design requirement, asserted on every template with marker inputs)

Not covered: A model's output is not deterministic and is not tested. The draft reports what you told it; a wrong input becomes a wrong sentence. Read every sentence against your records.

Protocol & SOP Writer — 7/10

  • The prompt separates standard practice from guesswork, and generated output carries a draft notice in the interface (Required behaviour: Design requirement, asserted on the prompt and the interface text)

Not covered: The twelve library protocols are fixed text; their volumes, timings and concentrations are not checked against manufacturer protocols by any test.

Pre-Submission Manuscript Audit — 8/10

  • Animal text routes to ARRIVE, qPCR text to MIQE, SAMPL always applies; a bare methods sentence fails; every PASS quotes text that is really in the manuscript (Required behaviour: ARRIVE 2.0, MIQE, SAMPL)
  • The MIQE efficiency item needs an efficiency, a standard curve or a slope with its R²; gene names containing "R2" (HER2, NR2B) do not satisfy it (Required behaviour: MIQE (Bustin et al. 2009), efficiency item)
  • The suggested reviewer response does not claim a power analysis, antibody authentication or guideline compliance that the manuscript does not report (Required behaviour: Required behaviour)

Not covered: It is a checklist, not a review: it detects whether a topic is addressed, not whether the statistics are appropriate.

Corrections log

Errors we found in SciKeep, newest first, including ones that could have changed a result. Nothing is removed from this list.

2026-09-25 — Site: Raw LaTeX delimiters showed as literal text in three method notes

Usability, no effect on results

What was wrong: Four sentences across the MSA, master-mix and drug-synergy method notes were written with LaTeX math delimiters ($...$, $$...$$, \text{}, \times) that the article renderer does not interpret, so a reader saw the literal markup — "$N$", "$$E_{\text{expected}} = ...$$" — instead of the intended notation.

Who was affected: Anyone reading those three posts; no calculation or number was affected, only the prose.

What changed: Rewritten as the plain-text math notation the rest of the site uses (e.g. "Penalty = g_o + (k − 1) × g_e").

2026-09-24 — Site, Electronic Lab Notebook, Grant Justification Kit: Compliance and validation language overclaimed in five places, on top of the "immutable" fix earlier the same day

A statement on the site was wrong

What was wrong: A sweep for the same pattern as the "immutable" audit trail found five more: (1) the section introducing all 42 tools said every one was "validated against wet-lab benchmark data" — false for most of them; freezer, antibody and reagent trackers are record-keeping with nothing to benchmark, several calculators have no automated test at all, and the tool-by-tool validation page itself lists which are checked against a formula, hand arithmetic, or real image data instead. (2) The same overclaim appeared as the shorter tagline "Published methods, benchmarked and cited" in the hero and the pricing modal. (3) The Western blot tile's badge read "ImageJ Standard", naming a specific competing product as a match, while the validation page says lane densities are "not benchmarked against ImageJ on real film or ChemiDoc images." (4) Signing a notebook entry — a typed name and a checkbox, with no identity verification — stamped the record with the declaration "Certified accurate, complete, and reproducible experimental record under 21 CFR Part 11 standards", and the modal was titled "21 CFR Part 11 Electronic Signature", though the validation page states this signature "does NOT meet 21 CFR 11.200." (5) The Grant Justification Kit generated boilerplate meant to be pasted into a real NIH, NSF or Horizon Europe grant application asserting SciKeep's local execution was "ensuring compliance with institutional IRB, HIPAA, and IP governance protocols", would "guarantee... FAIR... compliance", was "fully aligning with NSF... mandates", and — in the modal's own subtitle — was flatly "Compliant with NIH NOT-OD-21-013 DMS Policy requirements". A researcher who pasted that text into a federal submission was making a compliance representation on SciKeep's behalf that the product cannot back up.

Who was affected: Anyone who read the 42-tool intro or either tagline as a validation claim, anyone who took a signed notebook entry's declaration as a Part 11 certification, and anyone who copied the grant boilerplate into an actual NIH, NSF or Horizon Europe application.

What changed: The tool-intro copy and both taglines now say tools are graded openly on the validation page rather than uniformly "benchmarked". The blot badge names its method ("Rectangular Lane Profile") instead of a competing product. The signature modal is titled "Electronic Signature" with a line stating it does not meet 21 CFR 11.200, and the stored declaration is the signer's own attestation rather than a certification claim. The grant kit's three templates now say a claim is "supported by" or "supports" the relevant requirement rather than "ensures compliance" or "guarantees" it, its subtitles no longer assert compliance status, and the modal carries a standing notice that it is a draft to edit, not a compliance certification.

2026-09-24 — Site, Electronic Lab Notebook: The audit ledger was still called "immutable" in nine places

A statement on the site was wrong

What was wrong: The 31 August correction removed "legally immutable" from the notebook itself, and the tool, the features page and the validation ledger have since said plainly that a hash chain held in your own browser is tamper-EVIDENT and not tamper-proof — someone with the developer console can rewrite the whole chain, and clearing site data removes it. The word survived everywhere else: the landing page (three places), the pricing modal, the comparison table, the FAQ, the Part 11 method note and its search description, and the notebook's own empty audit-trail message. A reader who saw only those surfaces would take the ledger to be something the product has never claimed to be in the places that describe how it works.

Who was affected: Anyone who judged SciKeep's Part 11 posture from the marketing pages, the pricing modal, the comparison table or the FAQ rather than from the tool and the validation page.

What changed: Every remaining "immutable" is now "tamper-evident", and the FAQ answer states outright that the chain gives detection rather than prevention, that the console can rewrite it and that clearing site data removes it. The comparison table row and the landing-page feature name were renamed to match. The word is left untouched in this log, where it records what was said.

2026-09-24 — Site: The isoelectric point article cited a paper about extinction coefficients as its source

A statement on the site was wrong

What was wrong: The method note "Why Your Protein's pI Matters" carried DOI 10.1002/pro.5560041120 — Pace et al. 1995, "How to measure and predict the molar absorption coefficient of a protein". That paper is the correct source for the article's ε280 formula and nothing else in it; it says nothing about isoelectric points. The article also stated the pKa set and the algorithm without saying that different pI calculators use different pKa tables, that ExPASy uses the Bjellqvist values rather than the textbook ones SciKeep uses, or that a folded protein's measured pI can sit half a pH unit or more from any sequence-derived number. Its markdown also rendered raw: bullet lists collapsed into run-on paragraphs and italics printed with their asterisks showing, across nine method notes.

Who was affected: Anyone who followed the citation expecting the source of the pI calculation, or who read the article as saying a calculated pI should agree with a published or measured one.

What changed: The citation is now Bjellqvist et al. 1993, Electrophoresis 14(10):1023–31, doi:10.1002/elps.11501401163 — the primary reference for predicting focusing positions from sequence. Pace 1995 is still cited, in the section it actually supports. The article was rewritten to state which pKa table SciKeep uses, why tools disagree, and why the folded protein disagrees with all of them; the validation page already recorded that the pI is not benchmarked against ExPASy. The article renderer now handles lists, line breaks and italics.

2026-09-24 — Sanger Trace Viewer, Reagent Tracker, Buffer Calculator: A schematic peak diagram was described as fluorescence channels, and three smaller mislabels

A statement on the site was wrong

What was wrong: The peak diagram is drawn from the called bases, one peak per base, but the page described it as "four fluorescence channels" and the legend named BigDye terminator dyes — a chemistry this view never reads. It cannot show a secondary peak the basecaller did not call, which is the main thing someone looks at a trace for. The reagent tracker opened on six invented orders with made-up people and grant codes, with no example notice, driving a red low-stock alert and a spend-by-grant panel that looked like the user's own. Its CSV export joined fields with a bare comma, so a reagent name containing a comma shifted every later column. The buffer list also offered "NaCl Stock (5M)" while the default produced 1 M, and showed a concentration control on the 0.5 M EDTA recipe that changed nothing.

Who was affected: Anyone who read the peak diagram as a real trace, took the reagent low-stock or spend figures as their own, or exported a reagent CSV containing a comma in a name.

What changed: The diagram is labelled a schematic and the dye names are gone; the reagent tracker shows an example notice and both CSV exports use the shared escaping exporter; the NaCl label no longer claims 5 M and the inert EDTA control is hidden.

What to do: Re-open any reagent CSV exported before today if a reagent name contained a comma.

2026-09-24 — Sequence Toolkit, Cloning Planner: Half of all sites for the Golden Gate enzymes were invisible, and Dcm warnings fired on unrelated motifs

A result could have been wrong

What was wrong: The restriction digest scanned only the forward strand. For palindromic enzymes that is the same answer either way, but for BsaI, BsmBI and BbsI — the enzymes people check before Golden Gate — roughly half of real sites sit on the reverse strand and were simply not found. The Dcm methylation check also flagged a site whenever a CCWGG motif appeared anywhere within ten bases, without requiring it to overlap the recognition sequence, so sites that would cut perfectly well were reported as blocked. The Gibson planner additionally demanded an overlap melting temperature between 48 and 55 °C; the manufacturer's guidance is a floor of about 48 °C with no upper limit, so ordinary overlaps were marked "not optimal". Toolkit descriptions also claimed six-frame translation, a sliding GC window, a mouse codon table and NCBI gene retrieval, none of which this tool has.

Who was affected: Any digest or Golden Gate check for a non-palindromic enzyme, any site reported as Dcm-blocked, and any Gibson overlap rejected for being too warm.

What changed: Both strands are scanned, Dcm requires a real overlap, the Gibson check enforces only the floor, and the descriptions list what the tool does.

What to do: Re-check any construct you cleared for internal BsaI, BsmBI or BbsI sites.

2026-09-24 — Electronic Lab Notebook: The audit trail recorded only signatures, and "hash chain verified" was never checked

A result could have been wrong

What was wrong: The ledger was written only when an entry was signed. Creating, editing, deleting and exporting recorded nothing, so a notebook of unsigned entries had an empty ledger and an entry could be rewritten with no trace — while the page said "Every create, edit, signature and deletion is recorded in an append-only ledger". The chained value also left the entry text out, so even a logged action did not commit to what the entry said. The green "✓ Cryptographically sealed — hash chain verified" line appeared whenever a signature existed; no verification code existed anywhere in the product. Editing a signed entry kept the signature, so a PI's name stayed on text they never saw. The "This month" count compared the month but not the year. Site copy also called the ledger "legally immutable", offered a "GLP compliance certificate (PDF)" that is a JSON file, and said SciKeep "complies with FDA requirements".

Who was affected: Anyone relying on the ledger as a record of what happened to their entries, or on the signature seal as evidence that an entry was unchanged.

What changed: Create, edit and delete are recorded, the chain covers a hash of the entry text, and the chain is recomputed in the browser with the real result shown. Editing a signed entry now clears the signature and says why. The compliance and certificate claims were corrected, and the tool now states plainly that signing needs only a typed name, which does not meet 21 CFR 11.200, and that a ledger in your own browser is tamper-evident rather than tamper-proof.

What to do: Treat any entry edited before today as unverified, and re-sign entries whose text changed after signature.

2026-09-24 — Western Blot Analyzer: Per-lane significance stars were uncorrected, and two stated limits were not what the tool does

A result could have been wrong

What was wrong: The per-lane table ran a separate Welch test of every treated lane against the control and printed stars from the raw p-value, with no correction for the number of comparisons — five treated lanes carries about a 23% chance of at least one false star — while the public page said a Holm correction was applied by default. That was true only of the group chart. The page also said a band under about 10% above background is "flagged as under-measured", and the confidence note put the shortfall at "roughly 20%": no code implemented either figure. Measured on bands of known density, a band below about 10% above background is not detected at all, and a detected band recovers about 9-10% low whatever its intensity. The offline record for vinculin also gave 116.6 kDa, which is an isoform mass, against an accession whose canonical protein is 123.8 kDa — while the tool's own dropdown said 124.

Who was affected: Significance stars in the per-lane table, and anyone who relied on the stated faint-band behaviour or the offline vinculin weight.

What changed: Lane p-values are Holm-adjusted across the lanes tested and the stars follow the adjusted value. The faint-band and pairwise-comparison descriptions now say what the code does, backed by tests that measure it. Vinculin reads 123.8 kDa.

What to do: Re-check any significance star taken from the per-lane table where more than one treated lane was compared.

2026-09-24 — RNA-Seq Volcano Plot: A "Real Data" badge sat above invented example tables, and raw p-values were plotted as adjusted

Showed fixed or invented output as analysis

What was wrong: The tool carried a green "✓ Real Data" badge, and its three example datasets were described as published DESeq2 output from named cell lines. They are illustrative tables written by hand; no source, accession or citation exists for any of them. Separately, if an uploaded table had no adjusted p-value column the raw p-value column was plotted and thresholded silently, while the axis, the slider and the exported column all still read "p-adj" — the tool's own FAQ warns that this overstates confidence. Rows with a blank or zero p-adj were dropped with no count, and a blank p-adj is ordinary DESeq2 output for independently filtered genes, often a third of a real table. Site copy also described the tool as applying "DESeq2 thresholds" and "DESeq2 & EdgeR statistical standards"; it runs no model or test of any kind.

Who was affected: Anyone who read the example data as measured results, or who uploaded a table without an adjusted p-value column, or who read the up/down/not-significant totals as covering their whole table.

What changed: The badge is gone and the examples are labelled illustrative. A table without an adjusted p-value column now says so plainly, and dropped rows are counted and explained. The copy says what the tool does: it plots and thresholds a table you already have.

What to do: If you exported significant genes from a table with only a raw p-value column, re-run with adjusted values.

2026-09-24 — Figure Panel Composer: Scale bars were drawn from an assumed calibration, and Science figures exported 9 mm too narrow

A result could have been wrong

What was wrong: Every scale bar was computed as if the source image were 512 pixels wide and every pixel 0.16 µm, and there was no control anywhere in the interface to enter your own pixel size. A 2048-pixel confocal frame therefore got a bar four times too long. The drawn bar was also capped at 40% of the panel while its micron label was printed unchanged, so once capped the bar and its label no longer agreed. Scale bars were on by default. Separately, the Science preset used 55, 120 and 175 mm where Science publishes 5.7, 12.1 and 18.4 cm, so a full-width figure exported 9 mm narrow and the journal would scale it up, dropping it below the 300 DPI the tool tells you to hit. PNAS 1.5-column was 115 mm rather than 114.

Who was affected: Any exported figure with a scale bar, and any figure sized to the Science or PNAS presets.

What changed: Scale bars need a pixel size you enter and the image's real width, both now used in the calculation; a bar that cannot be drawn to scale is not drawn, and the panel editor says why. All four journals' widths are now asserted against their published guidelines — the previous test read the width from the preset it was checking, so it would have passed on any number.

What to do: Re-measure the scale bar on any figure exported from this tool, and re-export anything sized for Science.

2026-09-24 — Cell Culture Tracker: Five published doubling times were not in the Cellosaurus record they were badged with

A result could have been wrong

What was wrong: Each line showed a doubling time beside a "Cellosaurus ✓" badge, but five of the ten values do not appear in the cited record: HeLa was 24 h where Cellosaurus reports 30-48 h, HEK293 20 h and HEK293T 18 h where it reports about 24-30 h, CHO-K1 16 h where it reports about 24 h, and NIH-3T3 18 h where it reports 20-21 h. "Jurkat" also carried the accession of the Jurkat E6.1 clone rather than the parent line. Separately, "Mark as Split" wrote a post-split confluency of 15% no matter what the split ratio was, and that figure is the starting point for the observed doubling time shown on the card — so a measurement presented as the dish's own was computed from a number the user never entered.

Who was affected: Anyone who planned a split or an experiment around the published doubling time for HeLa, HEK293, HEK293T, CHO-K1 or NIH-3T3, and every observed doubling time after a split.

What changed: All ten values were transcribed from their own Cellosaurus records and are checked against saved copies of those records by a test. The split action now asks for the confluency you actually see.

What to do: Re-check any timing planned around these figures — HeLa in particular was stated at half its published doubling time.

2026-09-24 — Buffer Calculator: The RIPA lysis buffer recipe was ten times too dilute

A result could have been wrong

What was wrong: RIPA's five amounts were written as grams per 100 mL but multiplied by the volume in litres. At the tool's default 100 mL batch it printed 0.08 g Tris-HCl, 0.09 g NaCl, 0.10 mL NP-40, 0.05 g deoxycholate and 0.01 g SDS — a tenth of every concentration, while the labels still read 50 mM, 150 mM, 1%, 0.5% and 0.1%. The other recipes in the same list use per-litre amounts and were correct.

Who was affected: Anyone who made RIPA from this recipe: the buffer had a tenth of the detergent and will not solubilise membrane or nuclear proteins.

What changed: RIPA now uses the same tested molarity x molecular weight arithmetic as PBS and TBS, with the detergents as real percentages of the batch volume.

What to do: Remake any RIPA prepared from this tool, and treat a western blot that showed no band for a membrane or nuclear protein as untrustworthy.

2026-09-24 — Cloning Planner: Golden Gate ignored the chosen enzyme, and gave BsmBI the wrong reaction temperature

A result could have been wrong

What was wrong: The type IIS enzyme you selected was shown in the output but used in no calculation, so parts were never screened for the enzyme's own recognition site — the most common reason a Golden Gate reaction fails. The planner's own default CMV promoter part contains a reverse-strand BsaI site and was reported as a valid, high-fidelity assembly with no warnings. The printed protocol also cycled BsmBI at 55 °C, which is its plain-digest temperature; NEB's Golden Gate protocol for BsmBI-v2 cycles at 42 °C, and T4 ligase is inactive at 55 °C, so nothing would have been ligated.

Who was affected: Any Golden Gate assembly planned here, and anyone who ran the printed BsmBI protocol.

What changed: Every part is scanned on both strands for the selected enzyme's site; a hit invalidates the assembly and names the position. Warnings are now shown at all — the result carried them but nothing rendered them. BsmBI cycling reads 42 °C.

What to do: Re-check any Golden Gate design from this tool for internal sites, and re-run a BsmBI reaction that used the 55 °C protocol.

2026-09-24 — CRISPR Indel Analysis: Reads trimmed differently from the reference were called as edits

A result could have been wrong

What was wrong: The alignment charged a full gap penalty at the ends, so the cheapest way to reconcile two reads that started or finished at different bases was to call the overhang an indel. A read identical to the reference but trimmed 12 bases was reported as a "12 bp deletion, in-frame", and a genuine 3 bp in-frame deletion in a read trimmed at both ends was reported as a 23 bp deletion causing a frameshift. Two Sanger reads almost never start and end at the same base, so this affected ordinary use.

Who was affected: Any indel call where the edited read and the reference were trimmed differently — including negative controls, which could read as deletions.

What changed: Leading and trailing gaps are free, so only indels inside the overlapping region are counted. Trimmed identical reads now report "no edit", and a 3 bp deletion reads as 3 bp and in-frame whether or not the reads are trimmed.

What to do: Re-run any indel call made here, especially one that reported a deletion the size of your trimming.

2026-09-24 — Sanger Trace Viewer, CRISPR Indel Analysis: Trace files were read from the wrong byte offsets, and the quality histogram showed fixed numbers

A result could have been wrong

What was wrong: Both .ab1 readers used directory-entry offsets one field too far. In the indel tool the data offset came from a field that is zero in real files, so every genuine .ab1 was rejected with a message telling the user to re-export it from their sequencing provider. In the trace viewer the element count evaluated to a constant 65,536, so the reader either pulled tens of kilobytes of trace binary in as extra "bases", or skipped the basecall record and fell back to matching the longest run of A/T/G/C in the raw binary and presenting that as the read. The per-base quality histogram was drawn from fixed values (18, 24 or 45) and never used the quality record at all, and the accuracy caption under the mean quality was hardcoded to the Q30 figure, so a mean of Q12 was labelled ">99.9% accuracy" when it is 94%.

Who was affected: Anyone who uploaded a real .ab1: the indel tool refused it, and the trace viewer could show a sequence and a quality profile that were not from the file.

What changed: One shared reader now uses the offsets the ABIF specification defines, cross-checked against the file header's own root entry. It returns the real per-base quality, the histogram and the accuracy figure are computed from it, and the binary-mining fallback is gone. The test fixture was built to the specification rather than to the parser — the previous fixture reproduced the parser's own wrong offsets, which is why this passed unnoticed.

What to do: Re-upload any .ab1 you read here, and re-check any sequence or quality figure taken from one.

2026-09-24 — siRNA Knockdown Designer: The off-target seed screen was reading the wrong strand, and the antisense overhang was on the wrong end

A result could have been wrong

What was wrong: The microRNA seed screen and the seed GC% were computed from the sense (passenger) strand. RISC loads the guide (antisense) strand, and that is the strand whose positions 2-8 cause microRNA-like off-target silencing, so a guide carrying a real oncomir seed was reported as "Low" risk. Update 75 corrected the seed sequences; the check was still pointed at the wrong molecule. Separately, the antisense oligo was written with TT at its 5' end instead of a 3' dTdT, so an oligo ordered from that line was wrong at both ends of the guide.

Who was affected: Every siRNA ranked or ordered from this tool: off-target risk, seed GC% and the antisense sequence.

What changed: Both the seed screen and the seed GC% run on the guide strand, and both strands carry a 3' dTdT. A test now builds the target from the guide, so it would fail on the old behaviour.

What to do: Re-run any siRNA you designed here, and re-check the antisense sequence of anything already ordered.

2026-09-24 — Standard Curve Calculator: On a 4PL curve the top standard interpolated to zero, and mid-curve values were up to 23% out

A result could have been wrong

What was wrong: The curve's two asymptotes were fixed to the lowest and highest observed readings rather than fitted. The top standard's reading therefore equalled the upper asymptote exactly, the inverse divided by zero, and any sample at the top of the curve was reported as 0 — a saturated well read as "analyte absent". Pinning the asymptotes also bent the curve: on noiseless data from a known 4PL, points across the middle came back 11-23% low. Readings above the curve were reported as 1.5x the top standard, a number no measurement produced, which was written into the exported CSV. The below-range flag compared against the blank instead of the lowest calibrator, so readings under the first standard were reported as quantifiable.

Who was affected: Every 4PL interpolation (the default for ELISA): concentrations across the curve, and anything at or above the top standard.

What changed: All four parameters are fitted, so the same known curve now reads every standard back within 5% and the top standard to itself. A reading beyond an asymptote reports "outside the curve" instead of a number, and range flags use the calibrated range.

What to do: Re-interpolate any ELISA or BCA result taken from a 4PL curve here, especially samples near the top standard.

2026-09-24 — Share Tools: Every "Copy Link" button copied a link that 404s

Usability, no effect on results

What was wrong: The eight tool cards carried hand-written paths (/tools/qpcr, /tools/crispr, /tools/ic50 and so on). None matched the tool's real address, so every copied link opened a not-found page for whoever received it, while the toast said "Copied to clipboard". The links also used the non-www domain, which redirects. The "Share Protocol" checkbox copied a /protocol/<id> link to a page that does not exist.

Who was affected: Anyone who sent a tool link or a protocol link from this panel.

What changed: Links are built from each tool's real route, so they are correct by construction, and a test fails if one ever drifts again. The protocol checkbox copies the protocol text instead of a dead link.

What to do: Re-send any SciKeep tool link you shared before today.

2026-09-24 — Team Collaboration: "Invitation Sent!" sent nothing, and "Team access revoked" revoked nothing

A statement on the site was wrong

What was wrong: Adding someone to the lab roster showed "Invitation Sent!" but no email was sent and no account access was granted; removing them said "Team access revoked" while changing only the local list. A "Copy Lab Join Link" button produced a link containing a token that nothing in the app reads.

Who was affected: Anyone who believed colleagues had been invited or that removing a row had revoked their access.

What changed: The roster is described as a list the lab keeps, with an Email invite button that opens your mail client, and the dead join link was removed. Account access is granted through lab membership, which is separate.

What to do: Confirm directly with anyone you believed was invited or removed.

2026-09-24 — 3D Structure Viewer: Alternate conformers were drawn twice, and AlphaFold files lost their confidence colouring

A result could have been wrong

What was wrong: Atoms modelled in more than one position were all kept, so a high-resolution structure showed extra atoms and its backbone trace zig-zagged between conformers: lysozyme 2VB1 was drawn with 2,900 atoms and 70 residues carrying two alpha-carbons. Residues numbered 52, 52A and 52B (antibody numbering) were counted as one residue. An AlphaFold model downloaded and then uploaded as a file was not recognised as a prediction, so pLDDT colouring showed flat grey. The page also said ribbons were derived from backbone geometry; helix and sheet are read from the file's own records, so any file without them — including every AlphaFold model — is drawn entirely as coil with nothing said.

Who was affected: Atom and residue counts, and backbone traces, for structures containing alternate conformers or insertion codes; pLDDT views of uploaded AlphaFold files.

What changed: Only the first conformer is kept (2VB1 now parses to 2,200 atoms and 129 alpha-carbons for its 129 residues), insertion codes are distinguished, AlphaFold files are recognised from their title and open in pLDDT colouring, and the viewer says when a file carries no secondary-structure records.

What to do: Re-check any atom or residue count taken from a structure with alternate conformers.

2026-09-24 — CRISPR gRNA Designer: A fetched or pasted gene froze the tab for seconds to minutes

Usability, no effect on results

What was wrong: The off-target search ran for every candidate guide during the scan, and each run walked both strands of the whole sequence, so the work grew with the square of the sequence length. Measured on a 24 kb sequence, the page was unresponsive for 24 seconds.

Who was affected: Anyone who pasted or fetched more than a few kilobases; no result was wrong, but the tool appeared to hang.

What changed: The off-target search now runs only for the ten guides that are shown, which does not change any reported number. The same 24 kb sequence now blocks the page for 12 ms.

2026-09-24 — CRISPR gRNA Designer: Cloning oligos were labelled by enzyme when the difference is the vector promoter

A statement on the site was wrong

What was wrong: Two oligo pairs were offered, "BsmBI" without an added G and "BsaI" with one. The extra G is what the U6 promoter needs to start transcription when the guide does not begin with G; it has nothing to do with which type IIS enzyme cuts the vector. A guide starting with G was given an extra G anyway, and oligos were shown for SaCas9 and Cas12a guides, which use different vectors.

Who was affected: Anyone who ordered oligos from the row whose label matched their enzyme rather than their vector.

What changed: One pair is shown for SpCas9 20-mers, with the G added only when the guide needs it, and the note names the vector (lentiCRISPRv2 with BsmBI, pX330 with BbsI). No oligos are shown for nucleases that need another vector.

What to do: Check any ordered oligos: the top should read CACC + G only if your guide does not already start with G.

2026-09-24 — Freezer Inventory: Every added sample was saved as a plasmid at 100 ng/µL, 50 µL, added by "Current Researcher"

Showed fixed or invented output as analysis

What was wrong: The Add form had no fields for type, concentration, volume or owner, so every item was stored as type Plasmid with concentration 100 ng/µL, volume 50 µL and owner "Current Researcher", and these values synced to your account. Items could not be deleted. The page described rack and shelf maps and 81- or 100-place boxes that the tool does not have.

Who was affected: Every item added through the form: its type shows as Plasmid regardless of what it is, and exported or synced records carried the invented values.

What changed: The form has type, concentration, volume and owner fields, items can be removed, the invented values are cleared from items saved earlier, and the description matches the tool.

What to do: Check the type of each item you added before today; it was recorded as Plasmid.

2026-09-24 — Oligo & Primer Resuspension Calculator: The G extinction coefficient came from a different reference set, so G-rich oligos got too much water

A result could have been wrong

What was wrong: ε260 used dA 15,200, dC 7,050 and dT 8,400 from one published set but dG 11,500 from another; the matching value is 12,010. ε was therefore too low for G-containing oligos (about 2.5% for a G-rich 25-mer), so the calculated amount and the resuspension volume were too high and stocks came out slightly below their stated concentration. The page also described ε as nearest-neighbour; it is a nucleotide sum.

Who was affected: Resuspension volumes and nmol amounts for G-containing oligos.

What changed: dG is 12,010, the calculator uses the shared tested functions, and the description says what the method is.

What to do: For G-rich primers where a few percent matters (for example qPCR standards), recompute the stock concentration.

2026-09-24 — Dilution Calculator: Impossible dilutions were printed as bench instructions

A result could have been wrong

What was wrong: A target more concentrated than the stock gave a negative diluent volume (for example "add 50000 µL stock to -40 mL buffer"), and a zero stock gave "Infinity µL". The page also described unit conversion and serial dilution, which the tool does not do.

Who was affected: No correct calculation was affected; impossible inputs produced impossible instructions.

What changed: Impossible inputs are refused with the reason, and the description matches the fixed units (stock mM, target µM, volume mL).

What to do: If you made a solution from a negative or infinite volume shown here, remake it from a stock more concentrated than the target.

2026-09-24 — Pre-Submission Manuscript Audit: The audit passed items it had not found and wrote claims into the suggested reviewer response

A result could have been wrong

What was wrong: The MIQE amplification-efficiency item passed for any text containing "R2", so a qPCR paper about HER2 or NR2B passed with no efficiency data, while a real "slope of -3.35 (R2 = 0.99)" did not match. The suggested reviewer response said "Statistical power analysis was performed" whenever an n was given and called any RRID "authenticated"; a found RRID was reported as "compliant with ARRIVE antibody guidelines", though ARRIVE has no antibody item. The page claimed 25 checks against ARRIVE, CONSORT and MIQE (there is no CONSORT check, and several checks apply only to some manuscripts), numbered SAMPL items S1-S5 (SAMPL does not number its items), and described the AI reviewer as a Gemini model trained on journal review rubrics.

Who was affected: MIQE efficiency results for manuscripts containing gene names with "R2", and any reviewer-response text copied from the audit.

What changed: The efficiency pattern requires a slope with its R², an efficiency or a standard curve. The response paragraph states only what the text reports. The counts, guideline names, SAMPL labels and AI description now match what runs.

What to do: Re-read any reviewer response copied from the audit and remove claims about power analysis or antibody authentication your manuscript does not make.

2026-09-24 — Transfection Calculator: Lipofectamine 3000 amounts were wrong for large vessels and left out P3000 Reagent

A result could have been wrong

What was wrong: Amounts came from fixed multipliers that were not the vessels' growth areas: a 10 cm dish got 10 µg DNA in 8 mL, where the manufacturer's table gives 14–28 µg in 10 mL, and a T-75 was under-scaled similarly, while a 96-well well got 0.13 µg instead of 0.1 µg. The steps for Lipofectamine 3000, the default reagent, never mentioned the required P3000 Reagent, and used 2 µL of lipid per µg where the protocol tests 1.5 and 3 µL. Every other reagent was given the same DNA amounts and Opti-MEM steps, and an "expected efficiency" was shown with no source.

Who was affected: Transfections planned with this tool, especially in 10 cm dishes, T-75 flasks or 96-well plates, or with Lipofectamine 3000 without P3000.

What changed: Lipofectamine 3000 amounts now come from Thermo Fisher's protocol and scaling table, including P3000 and both lipid doses. For other reagents, you enter a condition that works and it is scaled by growth area. The efficiency estimate was removed.

What to do: If a Lipofectamine 3000 transfection planned here underperformed, check that P3000 Reagent was included and re-scale from the manufacturer table.

2026-09-24 — Antibody Tracker: RRID lookup filled in the wrong antibody for most RRIDs

Showed fixed or invented output as analysis

What was wrong: The lookup read the registry's reply in a format it never uses, so it never used the live SciCrunch record. It fell back to five stored records, three of which were attached to the wrong RRID (AB_2687626 is an HRP secondary, not Ki-67; AB_10694088 is a biotinylated cleaved caspase-3 antibody, not GAPDH; AB_2178887 is an NF-κB p65 antibody, not cleaved caspase-3), and any other RRID was answered with the Ki-67 record under your RRID. Missing fields were filled with "Rabbit", "Cell Signaling Technology" and WB/IHC/IF.

Who was affected: Any antibody added with the RRID Lookup button: its name, target, host, vendor and catalogue number may belong to a different antibody.

What changed: The live registry record is read correctly; offline, only five records copied from the registry are used, and any other RRID returns "not found". Nothing is filled in that the registry does not state. The SEO page's example RRID was also wrong and has been replaced.

What to do: Check each antibody you added with RRID Lookup against scicrunch.org/resolver before citing it.

2026-09-24 — Methods Section Writer: Template Methods paragraphs stated kits, catalogue numbers and guideline compliance you never entered

Showed fixed or invented output as analysis

What was wrong: Without AI, the writer filled in specific kits, catalogue numbers, software versions, dilutions and cell densities that were not among your inputs, and some templates ended with a compliance statement (ARRIVE, MIQE, "Nature Reporting") that nothing had checked. The example flow-cytometry RRID (AB_314154) is an anti-CD11b antibody, not the CD4/CD8 antibodies named; two other example RRIDs did not match or did not exist.

Who was affected: Any Methods paragraph copied from the template (non-AI) path.

What changed: Anything you have not entered is now a [bracketed placeholder]; no template claims compliance. The page shows an example notice until you change the pre-filled fields, and the wrong example RRIDs were replaced with placeholders.

What to do: Re-read any template-generated Methods text against your records; remove any kit, catalogue number, version or compliance sentence you did not supply.

2026-09-24 — Protocol & SOP Writer: The protocol library was described as generated from your inputs and "Bench Validated"

A statement on the site was wrong

What was wrong: The page and its confidence note said protocols were generated from your inputs with values echoed from what you entered, and a badge said "Bench Validated". The tool is a fixed library of twelve reference protocols that takes no input and was not validated. Its CRISPR protocol also gave Nucleofector program CA-137 for HEK293T; Lonza's own HEK293 protocol uses CM-130.

Who was affected: Anyone who followed a library protocol believing its values were tailored to them or validated.

What changed: The page, badge, description and note now call it a reference library, the copied and downloaded text carries a notice to check each value, and the program code points to Lonza's optimized-protocol list.

What to do: Check any Nucleofector program you took from the library against Lonza's protocol for your cell line.

2026-09-24 — Drug Synergy Matrix: A "published NCI ALMANAC reference" panel showed synergy scores that no published source contains

Showed fixed or invented output as analysis

What was wrong: The tool offered five "verified" drug pairs with Bliss and Loewe scores and record IDs attributed to NCI ALMANAC, DrugCombDB and SynergyFinder, and showed the matching score next to your result as the "expected" value. NCI ALMANAC screened only FDA-approved drugs on the NCI-60 cell panel and reports ComboScores, not Bliss or Loewe scores; three of the five pairs used cell lines outside that panel, one used a drug that is not FDA-approved, and none of the record IDs could be traced. Site pages also claimed a "5,000+" or "2M" pair benchmark database and Loewe, ZIP and HSA models. Only Bliss has ever been computed, and only five pairs were ever stored.

Who was affected: Anyone who compared their result to, or cited, one of the reference scores, or who described the tool as benchmarked against NCI ALMANAC.

What changed: The reference table, the pair selector and the reference panel were removed, and every page now describes the tool as a Bliss independence calculator. The Methods paragraph no longer claims cross-referencing against NCI ALMANAC. Your own Bliss scores were never affected.

What to do: Remove any reference score or NCI ALMANAC comparison taken from this tool from your notes or manuscript.

2026-09-24 — qPCR ΔΔCt Expression: Pfaffl mode ignored the reference gene's own efficiency

A result could have been wrong

What was wrong: In Pfaffl mode, which is the default, the whole ΔΔCt was raised to the target gene's efficiency, so the reference gene's measured efficiency was never used. With a target at 90% and a reference at 110%, a 2-cycle target shift and a 1-cycle reference shift gave 1.90-fold instead of Pfaffl's 1.72-fold. The error grows with the gap between the two efficiencies and with the size of the shift.

Who was affected: Pfaffl fold changes, intervals and p-values where the target and reference efficiencies were not equal. Livak results, and Pfaffl results with equal efficiencies, were correct.

What changed: Each gene's Ct is now weighted by its own efficiency, exactly as in Pfaffl (2001), and a hand-worked test checks the 1.72 case.

What to do: Re-run any Pfaffl analysis where the target and reference efficiencies differed.

2026-09-24 — siRNA Knockdown: Four of the five microRNA seeds in the off-target screen were wrong

A result could have been wrong

What was wrong: The screen that warns when a guide carries a known microRNA seed held the wrong 7-nucleotide seed for miR-21, let-7a, miR-34a and miR-126 (only miR-17 was right), so guides carrying those real seeds were not flagged, and some unrelated guides were.

Who was affected: siRNA designs ranked before this date: a guide could have been shown as low-risk while carrying one of these seeds.

What changed: All five seeds are now positions 2-8 of the miRBase mature sequence (accession cited), checked by a test that derives each seed from the published sequence.

What to do: Re-run the designer on any guide you chose, and check its seed region against miRBase.

2026-09-23 — Reagent Orders & Stock: The order-count and spend tiles showed fixed numbers

Showed fixed or invented output as analysis

What was wrong: The "orders this month" and "spent this month" tiles always showed 12 and $3,450, whatever you had entered, and the add, edit and delete buttons did nothing.

Who was affected: Anyone who read those tiles as a summary of their own orders.

What changed: Both tiles and the spend-by-grant breakdown are computed from your entries, and adding, editing and deleting reagents works.

2026-09-21 — Cell Imaging & Quantification: Public pages quoted out-of-date accuracy figures and presented our deep-learning model as ahead of Cellpose

A statement on the site was wrong

What was wrong: The landing page, the user guide, the deep-learning control and the note at the top of the imaging tool quoted the built-in engine at F1 12.0% on H&E tissue and 60-64% on fluorescence, and presented the tool's own deep-learning model as ahead of Cellpose (67.2% against 63.2%). The engine figures were measured before the 19 August watershed fix. The comparison with Cellpose is an August measurement that has not been re-run.

Who was affected: Visitors and users deciding whether to use the built-in engine, the deep-learning option or Cellpose on H&E tissue, or how far to trust the engine on fluorescence.

What changed: The figures were replaced with the 21 September re-measurement: on two samples of 62 NuInsSeg H&E patches Cellpose scored F1 65% and 69%, StarDist 45% and 52%, and the built-in engine 23% and 26%; on fluorescence the built-in engine scores 87% (BBBC039), 80% (BBBC020) and 71% (BBBC038). The tool now recommends Cellpose for H&E tissue you plan to publish: at the top of the imaging tool, in the user guide, on the landing page, in the deep-learning control and in the upload summary when H&E is recognised. The claim that our model is ahead of Cellpose was removed, because it rests on the August measurement.

What to do: If you used the built-in engine on H&E tissue because of the published figures, re-check those images with Cellpose, and check the overlay before quoting counts.

2026-09-21 — Cell Imaging & Quantification: The default segmentation setting was too permissive on most image types; it has been changed

A result could have been wrong

What was wrong: The classical (non-deep-learning) segmentation grows each detected nucleus outward to a low threshold placed halfway (0.50) between the background and the Otsu level. That value was chosen in August to correct nuclei that came out about 38% too small, which was largely a side effect of the watershed defect fixed on 19 August. On the current code it scored below a plain single threshold on four of five benchmark sets, mostly by merging nuclei and by over-detecting on H&E tissue.

Who was affected: Object outlines, counts and areas from the classical pipeline before this date, on macrophage, H&E and mixed images. On evenly stained fluorescence nuclei (BBBC039) the change makes no difference. On the benchmarks the object count moved by 1 to 4% on most sets and by about 11% on TNBC (H&E: 9,666 objects before, 8,556 after, against 4,028 true).

What changed: The low threshold was raised from 0.50 to 0.85. It was chosen on BBBC039 and BBBC038 alone by a protocol fixed before the run, then confirmed on four sets that took no part in choosing it, including a second NuInsSeg sample of different patches fetched afterwards: F1 rose on every one (BBBC020 78.3% to 79.9%, TNBC 29.5% to 35.1%, NuInsSeg 21.2% to 22.8% and 22.0% to 25.5%; mean +3.1 points). A plain single threshold still scores about as well overall on the held-out sets, so this corrects a value that was too low; it does not show that the two-level mask is better than a single threshold. The deep-learning path and the growth cap are unchanged.

What to do: If you counted or measured nuclei on H&E, macrophage or mixed images with the default settings before this date, re-run them and check the overlay. H&E counts remain a screen, not a measurement.

2026-09-21 — Cell Imaging & Quantification: The H&E accuracy figure across organs was out of date, and the comparison with deep-learning tools overstated the gap

A statement on the site was wrong

What was wrong: The tool's note gave F1 12.0% for H&E tissue across 31 organs (NuInsSeg) and described a 53-point gap to Cellpose. That was measured on the code before the 19 August watershed fix. The note also did not say that the tool's automatic stain detection recognised only a small share of those patches as histology.

Who was affected: Anyone who read the note to judge the tool on H&E tissue, or to decide between it and Cellpose or StarDist.

What changed: Re-run on the same 62 patches (3,248 nuclei, 31 organs): F1 21.2% at the shipped settings (precision 15.6%, recall 32.9%, count 2.1 times too high), 22.7% with the plain single threshold; by organ from 64.4% (cerebellum) to 0.0%. Automatic detection recognised 9 of the 62 as histology. The scoring was cross-checked with an independent scorer, and reverting the single watershed line reproduces the old 12.0% exactly. The gap to Cellpose (64.9%, not re-run) is now about 44 points. Nothing in the tool's settings was changed.

What to do: On H&E tissue treat counts as a rough screen, and use a dedicated histology model (Cellpose, StarDist, HoVer-Net) for anything you intend to publish.

2026-09-21 — Cell Imaging & Quantification: The macrophage and H&E accuracy figures were out of date, and the two-level mask does not help on three of four benchmarks

A statement on the site was wrong

What was wrong: The tool's note gave F1 62.9% on BBBC020 (mouse macrophage nuclei) and 24.6% / 20.5% on TNBC (H&E breast cancer), measured on the code before the 19 August watershed fix. The note also presented the two-level mask as an improvement and the 3-pixel growth cap as best on every set. The two TNBC figures were the same settings on different subsets (the first 25 patches and all 50), which the note did not say.

Who was affected: Anyone who read those figures, or the claim that the mask improves results, to judge the tool on macrophage or histology images.

What changed: All four benchmark sets have now been re-run on the current code. Shipped settings: BBBC039 87.0%, BBBC020 78.3%, BBBC038 67.8%, TNBC 29.5% (all 50 patches; 36.0% on the first 25). The plain single threshold scores higher than the shipped two-level mask on BBBC020 (81.8%), BBBC038 (70.2%) and TNBC (36.6%), and lower only on BBBC039 (85.7%); the problem the mask was added to fix (nuclei about 38% too small) no longer occurs without it. On H&E the tool finds 2.4 times too many objects, so counts there are a screen, not a measurement. Scores were cross-checked with an independent scorer, and reverting the single watershed line reproduces every old figure. Nothing in the tool's settings was changed; the note and the benchmarks now say what was measured.

What to do: On H&E tissue treat counts as a rough screen only. On any images, check the overlay before quoting counts or per-cell measurements.

2026-09-21 — Cell Imaging & Quantification: The "will it work on my images?" figure was out of date, and the note said the two-level mask helped

A statement on the site was wrong

What was wrong: The tool's note gave F1 60.0% on the BBBC038 collection (30-plus experiments, mixed stains and magnifications) as the generalisation figure, and said the two-level mask improved it from 58.3%. Both were measured on the code before the 19 August watershed fix. Re-measured on the current code, the shipped settings score higher than 60.0% but the two-level mask no longer helps on this set.

Who was affected: Anyone who read that figure to judge how the tool would do on images unlike a single tidy fluorescence set.

What changed: Re-run across the whole collection (670 fields, 29,461 nuclei): F1 67.8% at the shipped settings (precision 77.0%, recall 60.6%; 3.1% missed, 0.9% split, 8.2% merged; the count is 21% low). The plain single threshold scores 70.2% there, so on this heterogeneous set the two-level mask lowers F1 by about 2 points, while on the single-preparation BBBC039 it raises it by 1.3. Scores were cross-checked with an independent scorer. The note states all of this.

What to do: On images that are not evenly stained fluorescence nuclei, expect merged nuclei and a count that can run roughly a fifth low across mixed image types. Check the overlay before quoting counts.

2026-09-21 — Cell Imaging & Quantification: The published accuracy figure was out of date and understated

A statement on the site was wrong

What was wrong: The tool's note quoted F1 63.8% on the BBBC039 nuclei benchmark. That was measured on the code as of 17 August. A watershed defect fixed on 19 August (the flood stalled wherever the distance map rose, dropping about two-thirds of each rotated or elongated nucleus) had been lowering it. The note was not updated, and the benchmark scripts that produced it were not in the project.

Who was affected: Anyone who read the note to judge how well the tool finds nuclei. The figure understated it; the defect itself was fixed on 19 August.

What changed: The benchmark scripts are now in the project. Re-run on 21 September across all 200 BBBC039 fields (19,389 nuclei), the segmentation stage scores F1 87.0% at IoU 0.5 (89.9% on the 20-field subset behind the old figure), cross-checked with an independent scorer. The other datasets quoted in the note (BBBC020, BBBC038, TNBC, NuInsSeg) predate the fix and are marked as out of date until they are re-run.

What to do: Check the overlay on your own images before quoting counts or per-cell measurements; the benchmark measures one fluorescence dataset, not your preparation.

2026-09-21 — Kaplan-Meier Survival Curves: Log-rank p-values were wrong

A result could have been wrong

What was wrong: The tool converted the log-rank statistic to a p-value with a normal approximation that is poor at one degree of freedom, and then doubled the tail probability. A chi-squared upper tail already covers both directions, so the doubling was an error on top of the approximation.

Who was affected: Every log-rank p-value the tool displayed or exported before this date. Weak and moderate effects were reported as less significant than they are (a statistic of 3.84 gave p = 0.077 instead of 0.05; a small worked example gave 0.913 instead of 0.433). Strong effects were reported as more significant (a statistic of 10.83 gave about 0.0003 instead of 0.001). Anything below 0.0001 was floored at 0.0001.

What changed: The p-value is now the exact chi-squared tail for one degree of freedom, using the same function that the statistics tool checks against published critical values. Two hand-worked examples are now pinned by tests. The Kaplan-Meier curve itself, the risk table and median survival were not affected.

What to do: Re-run any survival comparison whose p-value you reported or whose conclusion depended on it. If a result crossed 0.05 in either direction, it may change.

2026-09-21 — Site: Verification notes cited tests that were not in the project

A statement on the site was wrong

What was wrong: Notes on several tools said their outputs were "asserted in npm test against published reference values". A review found that many of those tests were not in the repository: the original suite did not survive a platform migration, and the notes kept describing it. Several tools were graded 9 or 10 with no automated check on their own code, and the built-in flow-cytometry spillover values were described as "published" without a source.

Who was affected: Anyone who read a tool's grade or note as evidence that a specific reference check existed.

What changed: The missing checks that could be rebuilt were rebuilt, this time from published values or hand calculation written before the code was run (the "verified against" table below lists every check and the test behind it). Eight tools moved from "verified" to "standard method", and one from 10 to 9, where no check on the tool's own code exists; plain-arithmetic calculators without a test are shown as such. Notes that described tests which do not exist now say so. A new test fails if a tool is graded "verified", or a note says "asserted", without a real test behind it.

2026-09-21 — Site: Pricing and billing statements did not match what exists

A statement on the site was wrong

What was wrong: The FAQ quoted annual prices for plans that no longer exist, the landing page comparison showed a "$79/mo" plan, and the comparison-page cost calculator used $499 a year. The site also claimed a "14-domain automated verification suite", SSO and a dedicated account manager, and wire and ACH billing. None of those exist.

Who was affected: Visitors comparing prices, and institutions asking about procurement.

What changed: The one paid plan (Lab, $49/month or $468/year, per lab) is stated everywhere, and the unsupported claims were removed. Card payment through Stripe is the only billing offered today; purchase orders are handled case by case.

2026-09-21 — Site: Four references pointed to the wrong or a non-existent DOI; the MSA tool was labelled "Clustal"

A statement on the site was wrong

What was wrong: The DOIs for Thompson 1994 (CLUSTAL W), Vincent & Soille 1991 (watersheds) and Sebaugh 2011 (IC50 guidelines) did not resolve, and a citation of 21 CFR Part 11 was shown as if it had a DOI. Separately, the multiple sequence alignment tool was labelled "Clustal-W" or "Clustal Omega" although it runs Needleman-Wunsch alignment with a UPGMA guide tree.

Who was affected: Anyone who followed those reference links, or who described the alignment as Clustal in a methods section.

What changed: The DOIs were corrected and checked against doi.org; the tool now describes what it actually runs. Every DOI in the source is now tested for form, and DOIs are links. The tool still uses a "Clustal-style" conservation notation (* : .) and says so.

What to do: If you cited the MSA tool as Clustal, describe it as progressive Needleman-Wunsch alignment with a UPGMA guide tree, or use Clustal Omega or MAFFT for deep alignments.

2026-09-21 — All tools: Text could not be selected or copied

Usability, no effect on results

What was wrong: A page-wide setting that prevents text selection was applied to the whole app, so no result, table cell or computed value could be selected or copied.

Who was affected: Everyone. No numbers were changed.

What changed: Selection is enabled everywhere except the toolbars and image viewers, where dragging is the interaction. A test prevents it coming back.

2026-09-16 — Sample Size & Power Calculator: Proportion and survival designs used the t-test formula

A result could have been wrong

What was wrong: The Test Type picker offered Proportions (chi-square) and Survival (log-rank), but both ran the same two-sample t-test formula for a standardised mean difference. A proportion or a hazard ratio has no such relationship.

Who was affected: Any sample size or power figure produced for a proportion or survival design.

What changed: Each design now uses its own formula (Fleiss two-proportion z-test; Schoenfeld log-rank events). The t-test sample sizes were also checked against G*Power (d = 0.2 / 0.5 / 0.8 / 1.0 need 394 / 64 / 26 / 17 per group).

What to do: Recompute the sample size for any proportion or survival study you planned with this tool.

2026-09-16 — Standard Curve Engine: qPCR standards were blank-subtracted by default

A result could have been wrong

What was wrong: "Subtract zero-standard blank" was on by default and was not reset when the qPCR preset was chosen. qPCR standards are Ct values, which fall as concentration rises, so the lowest-concentration standard was subtracted from every point, driving later values negative and floored at zero.

Who was affected: A standard curve built from the qPCR preset with the box left checked reported R² = 0, slope 0 and an undefined efficiency.

What changed: Blank subtraction is off and disabled for qPCR standards, and the blank is taken from the standard whose concentration is actually zero.

What to do: If you built a qPCR standard curve with the qPCR preset and the blank box checked, rebuild it and check that R² and the slope are sensible before using the efficiency.

2026-09-16 — PCR Primer Designer: Pasting a new template did not recompute primers

A result could have been wrong

What was wrong: Primer candidates were computed once for the bundled example and not again when a real template was pasted over it.

Who was affected: Primers shown after pasting a template could belong to the example sequence rather than yours.

What changed: Candidates recompute whenever the template changes. The example template is now labelled as an example.

What to do: If you ordered primers from this tool by pasting a template over the default, check the primer sequences against your template.

2026-09-16 — Several exports: CSV files were cut off at the first "#"

A result could have been wrong

What was wrong: Eight CSV exports built the file as a web address, which silently truncates the file at the first # character: the primer order sheet, both CRISPR exports (guide order sheet and indel call), the plate-reader dose-response and 96-well plate exports, the master-mix export, and two administrator-only exports. (The blot export had the same fault and is listed under Western Blot.)

Who was affected: Any of those files that contained a "#" — in a compound name, a comment row or a note. Everything after it was missing, with no warning.

What changed: All exports now use the shared downloader, which encodes the file correctly and also escapes commas and quotes.

What to do: If you exported a CSV from these tools and a name or row looked short, export it again.

2026-09-16 — Wound Healing: Opened with a fixed demo and a typed-in p-value

Showed fixed or invented output as analysis

What was wrong: The tab showed a fixed demo dataset that never changed, with a p-value of 0.0004 typed into the sample data and displayed as if calculated.

Who was affected: Anyone who read the on-screen result as an analysis of their own image.

What changed: Rebuilt on real segmentation of the uploaded scratch image. The example is labelled as an example, and no p-value is shown that was not computed from measured replicates.

2026-09-16 — 3D Spheroid Morphometry: Opened with a fixed demo and a typed-in p-value

Showed fixed or invented output as analysis

What was wrong: Three hand-typed conditions with a typed-in p-value of 0.0002, unaffected by anything uploaded. The volume and sphericity formulas were sound but only ever received hand-typed axis lengths.

Who was affected: Anyone who read the on-screen result as an analysis of their own images.

What changed: The spheroid is now detected in the uploaded image and measured; the example is labelled.

2026-09-16 — Subcellular Puncta & Foci Counter: Foci counts were random numbers

Showed fixed or invented output as analysis

What was wrong: The tab opened with three conditions of Math.random()-generated foci counts and did not respond to uploaded images.

Who was affected: Anyone who read the counts as measured.

What changed: Nuclei and spots are now detected in the uploaded channels; the example is labelled.

2026-09-16 — Neurite Outgrowth & Angiogenesis: Skeleton metrics were hand-typed

Showed fixed or invented output as analysis

What was wrong: Total length, junctions, endpoints and a Sholl curve shaped like a Gaussian were typed in and never changed with the uploaded image.

Who was affected: Anyone who read the metrics as measured.

What changed: The image is skeletonised (Zhang-Suen thinning) and measured; the example is labelled.

2026-09-16 — 2D Scatter Gating: The gated population was regenerated at random on every drag

Showed fixed or invented output as analysis

What was wrong: The tab generated 240 entirely synthetic cells every time it was drawn, including on every gate-slider movement, so moving a gate created a new population instead of re-gating one.

Who was affected: Anyone who read gate percentages as describing their own cells.

What changed: Gates now read the real per-cell measurements already made in the imaging tool.

2026-09-16 — SuperPlots: Replicates were invented from cell ID parity

A result could have been wrong

What was wrong: When no condition label said "control", the tool split one image's cells into fake Control and Treated groups by even and odd cell ID, cut those into thirds and called the thirds biological replicates. The p-value came from a lookup table, not a test.

Who was affected: Any SuperPlot built from a single image, and its p-value.

What changed: Groups come only from real condition labels on separately imaged samples, and a p-value is shown only when every shown condition has at least two real replicates. Otherwise it says why it cannot.

What to do: Do not use a SuperPlot or p-value from before this date as evidence of a biological effect unless it was built from separate samples.

2026-09-16 — Site: Unknown tool addresses returned "found"; two tools did not label their example data

Usability, no effect on results

What was wrong: A mistyped or removed /tools/ address returned a normal page (and opened whichever tool loaded first) instead of a "not found". Separately, the Kaplan-Meier tool and the Primer Designer opened on example data and computed results from it without saying so.

Who was affected: Visitors following old links; search engines indexing non-existent pages.

What changed: Unknown addresses now return not-found, and both tools show the example-data notice until you enter your own.

2026-09-15 — Cell Imaging & Quantification: The multi-image comparison had a silent control and a fabricated ANOVA p-value

A result could have been wrong

What was wrong: Whichever image was loaded first was silently the control, with no way to change it. The one-way ANOVA p-value was read from a hand-picked table of F thresholds that ignored degrees of freedom, and starred results on that number. The help text promised Dunnett testing that was never implemented.

Who was affected: Significance stars and p-values in the "Impact vs Ctrl" comparison before this date.

What changed: You choose the control; the ANOVA p-value is the real F-distribution value; pairwise testing is Tukey-Kramer, as the help now says.

What to do: Re-check any multi-image comparison you reported, especially where the first image was not the control.

2026-09-15 — Western Blot: Six defects in blot quantification

A result could have been wrong

What was wrong: Significance stars were placed on the control lane (a missing p-value compared as less than 0.05). A lane with no loading-control band produced a fold change around 4.5 trillion. Lane IDs could duplicate after removing and adding lanes. Excluded lanes still fed later steps. Molecular weight was measured in the wrong coordinate frame when the band window changed. The exported CSV was cut at the first "#".

Who was affected: Fold changes, stars, molecular-weight estimates and exports from the blot tool before this date.

What changed: A lane without a control band shows "n/a", stars appear only with a real p-value, excluded lanes stay excluded, lane IDs are unique, molecular weight is measured against the ladder, and the export uses the shared downloader. You can also choose the control lane and band order, and error bars show SD, SEM or 95% CI with Holm-adjusted comparisons.

What to do: Re-run any blot quantification you reported, especially where a lane lacked a loading-control band.

2026-09-07 — Methods Generator: The prompt allowed invented experimental values

A result could have been wrong

What was wrong: The AI prompt asked the model to add placeholders "only where critical lab-specific catalogue details are missing", which licensed it to invent centrifuge speeds, incubation times and dilutions for everything else.

Who was affected: Any Methods text generated before this date could contain a plausible-looking value you never supplied.

What changed: The prompt now says to use only the parameters you provided, to mark anything not supplied with a bracketed placeholder, never to substitute a "typical" value, and not to add citations or claim compliance with a reporting standard.

What to do: Re-read any generated Methods paragraph against your own records, sentence by sentence.

2026-09-05 — Drug Synergy Calculator: The heading claimed Loewe and Combination Index; only Bliss is computed

A statement on the site was wrong

What was wrong: The tool was titled "Drug Synergy & Combination Index Suite (Bliss & Loewe)" but computes the Bliss independence score only. Loewe values appeared only as reference numbers.

Who was affected: Anyone who described a Bliss score from this tool as Loewe or Chou-Talalay.

What changed: The heading and FAQ state that it computes Bliss.

2026-09-05 — qPCR ΔΔCt, Statistical Tests, Drug Synergy: Example data produced real-looking results with no label

Showed fixed or invented output as analysis

What was wrong: These tools open on a worked example and calculate fold changes, p-values and synergy scores from it, with nothing on screen saying the data were an example.

Who was affected: Anyone who screenshotted or exported before entering their own data.

What changed: An "example data" notice appears until the first real edit.

2026-09-03 — Cell Imaging & Quantification: Analysis ran with settings from before the image was profiled

A result could have been wrong

What was wrong: The analysis closed over a stale copy of the image profile, so it ran with pre-profiling channel and size settings and hid the zero-result explanation.

Who was affected: On the test image, the same clicks gave 0 objects before the fix and 1,328 after. Counts from before this date may be too low.

What changed: The analysis uses the current profile.

What to do: Re-run any counts you reported, and check the overlay.

2026-09-03 — Cell Imaging & Quantification: The toolbar Upload button silently failed on TIFF files

Usability, no effect on results

What was wrong: The cyan Upload button on the image toolbar read files in a way no browser can decode as TIFF, and had no error handler, so a .tif produced no image and no message.

Who was affected: Anyone who uploaded TIFFs with that button. The separate "Open File" button worked.

What changed: Both buttons use the same loader, and failures say what went wrong.

2026-09-02 — Cell Imaging & Quantification: Missing calibration was assumed; a channel with no signal was segmented

A result could have been wrong

What was wrong: When a file carried no real scale, the tool assumed 0.325 µm per pixel without reading the scale bar burned into the image; on one test file the true value was 0.180, so areas were off by 3.25 times. Separately, segmentation defaulted to the blue channel even on images with no nuclear stain, giving zero objects.

Who was affected: Areas and object counts on images without embedded calibration, or without a nuclear channel.

What changed: The burned-in scale bar is detected and used, with a warning where the scale is assumed; the channel is chosen from a profile of the image and the choice is stated.

What to do: Check the scale on any image whose file had no embedded calibration; re-measure areas if the tool had assumed one.

2026-09-02 — Site: The Terms listed Syria as comprehensively sanctioned

A statement on the site was wrong

What was wrong: The sanctions section of the Terms named Syria as a comprehensively sanctioned country. US sanctions on Syria under those orders were revoked in 2025.

Who was affected: Visitors reading the Terms.

What changed: Syria was removed from the Terms and from the access block list, with a test asserting it stays out.

2026-08-31 — Western Blot: Band detection framed the wrong region on photos with a dark surround

A result could have been wrong

What was wrong: The tool judged whether bands were dark-on-light from the median of the whole image, so a photo with a dark surround was read as bright-on-dark and the blown-out white block became the strongest "signal". Separately, the minimum band height was a fixed constant no caller could change, so faint blots reported no bands.

Who was affected: Lane detection and densities on blots photographed with a dark surround, and faint blots reported as empty.

What changed: Polarity is measured inside the region you draw, with a notice when it changes; the threshold scales with the image's own noise and is deliberately conservative so a blank lane does not produce invented bands.

What to do: Re-run any blot whose lanes looked wrong or whose bands were reported as missing.

Scope and limitations

SciKeep is a research aid. Validate any result you intend to publish against an independent method or a known control. It is not a medical device and is not for clinical or diagnostic use.

All 42 tools · Methods guide · Compare with Prism, ImageJ and Benchling · Scientific disclaimer

SciKeepLoading SciKeep…