Why Your Protein's pI Matters — and How We Calculate It Correctly

How the Henderson-Hasselbalch bisection algorithm computes pI, why different tools disagree, and why a folded protein can sit half a pH unit from the calculated value.

A calculated isoelectric point is a prediction, not a measurement, and it is wrong often enough that you should know which way it fails before you order a column.

**The short answer.** The pI is the pH at which the protein's net charge is zero. You compute it by summing the charge of every ionisable group across a pH range and finding the crossing point. The arithmetic is exact; the pKa values it depends on are not, and a folded protein's real pI can sit half a unit or more from the sequence-derived number.

**What the calculation actually does.** Each ionisable group is either positive when protonated (the N-terminus, Lys, Arg, His) or negative when deprotonated (the C-terminus, Asp, Glu, Cys, Tyr). At a given pH, the fraction of each group carrying charge follows Henderson-Hasselbalch. Sum them:

Q(pH) = Σ (+1)/(1 + 10^(pH − pKa)) − Σ 1/(1 + 10^(pKa − pH))

Q is a smooth, monotonically decreasing function of pH: strongly positive in acid, strongly negative in base, crossing zero exactly once. Because it is monotonic, bisection finds the crossing reliably — halve the interval, check the sign, repeat. SciKeep iterates until the bracket is narrower than 0.01 pH units. There is no fitting and no ambiguity in the root; every disagreement between pI calculators comes from the pKa table, not the algorithm.

**Why different tools give you different answers.** Side-chain pKa values are measured on model compounds, and the published sets disagree. SciKeep uses a textbook set (Asp 3.9, Glu 4.07, His 6.04, Cys 8.14, Tyr 10.46, Lys 10.54, Arg 12.48). ExPASy's Compute pI/Mw uses the Bjellqvist values, derived by fitting to the observed focusing positions of polypeptides in immobilised pH gradients (Bjellqvist et al., *Electrophoresis* 1993) — that is, calibrated against denatured proteins in a gel.

Those two sets will not agree exactly. For most sequences the gap is a few tenths of a pH unit; for proteins with unusual composition it can be larger. If you are comparing a number from one tool against a number from another, or against a published value, check which table each used before concluding anything. SciKeep's validation page states plainly that its pI is not benchmarked against ExPASy, because it is not.

**Why the folded protein disagrees with both.** The bigger error is not the table. It is that this calculation treats every ionisable group as independent and fully solvent-exposed, which a folded protein is not.

- A buried carboxylate with no water around it can shift by several pH units. - A salt bridge stabilises the charged form of both partners, moving both pKa values apart. - Adjacent like charges destabilise each other, moving them together. - Histidine, sitting near physiological pH, is the residue most likely to be shifted and the one most likely to matter.

So a sequence-derived pI describes the *unfolded* chain. For an IEF gel run under denaturing conditions that is a reasonable model — which is precisely why the Bjellqvist set was fitted to that experiment. For a native protein in a chromatography buffer it is an estimate, and proteins with a highly asymmetric surface charge distribution are the ones that surprise you.

**What it is genuinely good for.**

**Choosing an ion exchanger.** Above its pI the protein is net negative and binds an anion exchanger; below, it is net positive and binds a cation exchanger. Pick a buffer pH roughly one unit away from the pI in the direction that gives you the charge you want. Working within about 0.5 units of the pI is where behaviour becomes unpredictable, and it is also where the calculation is least trustworthy — the two problems compound.

**Predicting where it precipitates.** Solubility is at a minimum at the pI, because net repulsion between molecules disappears. If your protein crashes out during dialysis, compare the buffer pH against the pI before you blame anything else. This is also exploitable: isoelectric precipitation is a genuine purification step for robust proteins.

**Reading a 2D gel.** The horizontal position in the first dimension is the focusing pI, which is what Bjellqvist's values were built to predict. This is the application where a computed pI is closest to a measurement.

**What it will not tell you.** Not the surface charge distribution — a protein with a strongly acidic patch and a strongly basic patch can have a neutral pI and behave like neither. Not the behaviour of a glycosylated or phosphorylated protein: each phosphate adds roughly −2 at neutral pH and sialic acids drag the pI down substantially, and neither is in your sequence. Not the pI of a complex, where the interface buries charges that the free subunits expose.

**A note on the companion number, ε280.** The same calculator reports the molar extinction coefficient, and that one is on much firmer ground: ε280 = 5500 × N(Trp) + 1490 × N(Tyr) + 125 × N(cystine), from Pace et al. (*Protein Science* 1995). Note the last term is per disulfide **bond**, not per cysteine — SciKeep counts floor(Cys/2), and gives you both the reduced and oxidised values because you know which state your protein is in and the calculator does not.

That formula is accurate to a few percent for most proteins, and unlike the pI it is measuring something the sequence genuinely determines. A protein with no Trp and no Tyr has essentially no A280, and for those you need a different assay rather than a more careful spectrophotometer.

**How to use the number responsibly.** Treat a calculated pI as a starting hypothesis with about half a pH unit of uncertainty, wider for proteins with unusual composition or heavy post-translational modification. Use it to choose which resin to try first, not to explain away a result that disagrees with it. If the pI matters to your conclusion — a purification that has to be reproducible, a claim about surface charge — measure it by isoelectric focusing. That takes an afternoon and settles the question.

**What the calculator does.** SciKeep computes net charge across pH from the Henderson-Hasselbalch sum, finds the root by bisection, and reports pI alongside molecular weight, ε280 for the reduced and oxidised forms, and Kyte-Doolittle hydropathy. The residue masses and the Pace coefficients are asserted against published values in the test suite. The pKa set is a textbook one and is not benchmarked against ExPASy, which the validation page states rather than glossing over — if you need agreement with a specific published pI, check which table produced it.

All articles · SciKeep tools

SciKeepLoading SciKeep…