method and records
Out-of-Specification Results and What to Do About Them
The difference between out of specification and out of trend, why retesting is never the first move, how to interrogate the method before the material, the averaging trap, and what to do when a supplier certificate contradicts your own result.
A result outside specification is information, and the way most laboratories destroy that information is by running the test again before doing anything else. The second number then displaces the first, the question of why they differ is never asked, and whatever produced the original result stays in place to produce another one later. This page sets out a proportionate sequence for a small laboratory: what to distinguish, what to check first, what retesting may and may not be used for, and how to handle the case where a supplier certificate and your own measurement disagree.

Out of specification, out of trend, out of expectation
An out-of-specification result falls outside a limit that was defined in advance. The definition contains its own requirement: if no limit was written down before the measurement, the result cannot be out of specification, only surprising. Laboratories that have never set acceptance criteria discover during their first anomaly that they have nothing to compare against, and end up arguing about whether a number is acceptable at the worst possible moment to be deciding it.
An out-of-trend result sits inside the limit but away from where the method normally puts it. A system suitability check that has run at 1.2 per cent relative standard deviation for a year and now runs at 1.9 per cent has passed and has also told you something. Out-of-trend findings are usually more informative than out-of-specification ones because they arrive earlier, while the cause is still small and still present, and they are the only warning most drifting systems give before they fail outright.
Out of expectation is the third category and the one small laboratories meet most: a result with no formal limit that is nonetheless clearly wrong — a recovery of 140 per cent, a blank with a peak in it, a duplicate pair that disagree by a factor of two. Handle these the same way. The absence of a written limit changes the paperwork, not the reasoning.
Why the first move is never a retest
The instinct to repeat the measurement is strong and almost always wrong, for a reason that is statistical rather than procedural. Any measurement carries scatter. If a result is near a limit, repeating it will sometimes produce a passing value by chance alone, and a laboratory that retests until it passes has built a machine for generating acceptable results from unacceptable material. This is the reasoning that underlies formal guidance on the subject, which treats the initial result as valid unless an assignable cause is identified, and treats retesting as a step within an investigation rather than as a substitute for one 1.
The practical consequence is a rule that can be stated in one line and is worth writing into a procedure: the original result is never deleted, and it is never overwritten by a subsequent one. It is either explained by an identified cause, in which case the explanation is recorded and the cause corrected, or it stands as a valid result. There is no third outcome in which it simply disappears because a later number was nicer.
Two things do happen immediately, before any investigation. Quarantine the material and the solution, physically and in the record, so that nothing is consumed or discarded while the question is open — the sample is evidence, and half of all investigations stall because it was used up. And preserve everything: the raw instrument file, the notebook page, the preparation record, the balance printout, the reagent lot numbers. An investigation is only as good as what survives the first hour of it.
Investigate the method before the material
The first phase is a laboratory investigation, and its question is narrow: is this result an artefact of how it was produced? Only when that has been answered negatively does the material itself come under suspicion. The order matters because analytical causes are both more common and cheaper to exclude, and because concluding that a batch is bad when the pipette was miscalibrated is an expensive mistake in both directions.
- Check the arithmetic. Recalculate from the raw readings, independently, without looking at the first calculation. Transcription and dilution-factor errors outnumber every other cause.
- Check the system suitability data for the run. If suitability failed or drifted, the run is invalid and no result from it means anything.
- Check the standards. Correct identity, correct concentration, in date, prepared from the right stock, stored correctly since.
- Check the preparation. Weighing, dilution sequence, solvent, volumetric glassware, whether the correct net peptide content correction was applied.
- Check the instrument. Calibration status, column history, detector performance, whether anything was changed between the last passing run and this one.
- Interview the analyst, without blame and while memory is fresh. Ask what was different, not whether a mistake was made.
- Compare against history. Has this method, this analyst or this instrument produced a similar deviation before?
- Only then extend to the material: sampling, homogeneity, storage history, batch identity, and whether an excursion is recorded against it.
An assignable cause is a specific, evidenced failure, not a plausible story. Somebody probably misread the balance is not an assignable cause; a balance printout showing a mass inconsistent with the recorded value is. The distinction is the entire integrity of the process, because a laboratory that accepts plausible stories can invalidate any result it dislikes. Where no assignable cause is found, the result stands and the investigation says so.
One check belongs early and is habitually skipped: ask whether the result is outside the limit by more than the measurement uncertainty of the method. A purity result of 97.8 per cent against a limit of 98.0 per cent, from a method whose expanded uncertainty is plus or minus 1.5 per cent, is not evidence that the material is below limit — it is evidence that the method cannot resolve the question being asked of it 2. That finding is about the specification, not the batch, and it is resolved by fixing the specification or the method rather than by retesting.
Retesting rules and the averaging trap
Retesting has a legitimate place: it tests a hypothesis. If the investigation identifies a candidate cause, a retest designed to confirm or exclude that cause is sound science. What is not sound is a retest with no hypothesis, performed in the hope of a different answer. The difference between the two is entirely a matter of what was decided before the second result was seen, which is why the rules have to be written into a procedure in advance.
| Decision | Set in advance as | Why |
|---|---|---|
| Number of retests | A fixed number, stated in the procedure | Otherwise testing continues until a pass appears |
| Who performs them | A second analyst where possible | Separates analyst technique from material effect |
| What sample is used | The original preparation, or a fresh one from the same sample | A new sample answers a different question and must be labelled as such |
| How results combine | Reported individually; averaging only if the method defines it | Averaging after the event conceals the failing value |
| What decides the outcome | A rule written before the retest is run | A rule written afterwards is a preference |
The averaging trap deserves naming because it looks so reasonable. A failing result and two passing results are averaged, the mean sits inside the limit, the batch is released. The arithmetic is correct and the conclusion is wrong: averaging assumes the values are replicate estimates of one quantity differing only by random error, which is exactly the assumption an out-of-specification result puts in doubt. Averaging is legitimate only where the method specifies replicate determinations and defines the mean as the reportable value in advance 4. Applied afterwards to a set containing a failure, it does not resolve the failure — it dilutes it, and the spread of the values, which is the informative part, is discarded in the process.
Document the investigation as it proceeds, not once it concludes. The record needs the original result and when it was obtained; the material and preparation involved; what was checked and what each check found, including checks that found nothing; the cause identified, with its evidence, or an explicit statement that none was found; what was done about it; and what happens to the material. Investigations that find nothing are worth keeping precisely because they accumulate: three unexplained deviations on the same instrument in a year is a pattern that no single investigation could see, and the ability to see it is one of the arguments for keeping records at all 3.
When your result and the certificate disagree
A supplier certificate states 99.1 per cent purity; your own determination gives 96.4 per cent. Before this is treated as a dispute, establish whether the two figures are measuring the same thing. Very often they are not. Purity by area normalisation at one wavelength on one gradient is not comparable with purity determined on a different column, gradient or detector, and neither is comparable with a figure that includes or excludes water, residual solvent and counterion. The most common resolution to an apparent disagreement is that the two numbers are answers to different questions 4.
Work through it in order. Confirm you are looking at the right batch — certificates travel between lots more easily than anyone expects. Read the method conditions on the certificate and compare them with yours; if the certificate does not state them, the figure cannot be compared with anything and that is itself the finding. Establish what each figure includes: chromatographic purity, net peptide content and mass fraction are three different quantities and are routinely confused. Run your own system suitability and a reference standard to demonstrate that your method is performing before you assert that somebody else is wrong.
If a genuine discrepancy survives all of that, raise it with the supplier in writing, stating your method conditions, your system suitability data and your result, and ask for theirs. Record the exchange against the batch. Meanwhile the material stays quarantined and the practical decision is straightforward: continue with the batch on the basis of your own measured value rather than the certificate value, or stop using it. What you must not do is proceed while quietly using the certificate figure in your calculations, because every result computed from that batch then rests on a number your own laboratory has evidence against.
References
- Guidance for Industry: Investigating Out-of-Specification (OOS) Test Results for Pharmaceutical Production
- Quantifying Uncertainty in Analytical Measurement, Eurachem/CITAC Guide CG 4, third edition
- ISO/IEC 17025:2017 General requirements for the competence of testing and calibration laboratories
- The Fitness for Purpose of Analytical Methods: A Laboratory Guide to Method Validation and Related Topics, second edition