Screening Tests in O&G
Start with one idea, because everything else is built on it: a test does not tell you whether disease is present; it changes how likely disease is. Before the test you already hold a probability — call it the pre-test probability — drawn from how common the condition is in this kind of woman, her age, her history, her symptoms. The test result nudges that probability up or down. A good test moves it a long way; a weak test barely moves it; a useless test leaves it where it was. No single test, however impressive its numbers, abolishes uncertainty. Once you internalise "test = probability shift, not verdict", sensitivity, specificity, predictive values and likelihood ratios stop being formulae to memorise and become the machinery that tells you how far the result should move you.
From that one idea, screening follows directly. Screening is the deliberate management of probability in people who are apparently well, or in broad clinical populations not yet sorted by symptoms. It is not diagnosis. A screening programme identifies the women or fetuses whose post-test probability is now high enough to justify counselling, closer surveillance, confirmatory testing, or preventive treatment — and reassures the rest. The value of screening therefore depends not only on test accuracy but also on disease burden, prevalence, acceptability, access, the confirmatory pathway behind it, treatment benefit, harm, and follow-up.
This is why screening sits at the centre of O&G statistics. Cervical HPV testing, antenatal HIV and syphilis testing, aneuploidy screening, pre-eclampsia risk algorithms, gestational diabetes testing, and obstetric early-warning scores all do the same thing: they turn a measurement into a changed probability and then a decision. The FCOG Primary task is to read the numbers honestly — sensitivity, specificity, PPV, NPV, likelihood ratios, ROC curves — and to keep sight of the denominator logic underneath them. The sections below climb in deliberate order: first the 2 by 2 table that defines the numbers, then how prevalence bends them, then the likelihood-ratio engine that does the probability arithmetic, then thresholds, calibration and net benefit, and only then the whole-programme and bias questions that decide whether a good test is actually a good thing to do.
Screening Versus Diagnosis
Before the numbers, fix the two jobs apart, because they sit at opposite ends of the same probability line. Screening starts from a low or moderate pre-test probability and asks only "is this woman's risk now high enough to take the next step?" Diagnosis starts from an already-raised probability and asks "is disease truly present or absent?" Screening and diagnosis therefore ask different questions and tolerate different errors.
| Feature | Screening | Diagnosis |
|---|---|---|
| Population | Asymptomatic or broad-risk group | Symptomatic, screen-positive, or high pre-test probability |
| Aim | Sort into higher and lower risk | Confirm or exclude disease |
| Threshold | Often favours sensitivity | Often favours specificity and certainty |
| Result | Probability or risk category | Disease classification |
| Harm | Anxiety, false positives, overdiagnosis, false reassurance | Procedure risk, treatment harm, missed disease |
| O&G example | HPV screen, NIPT, syphilis rapid test, pre-eclampsia risk model | Cervical biopsy, fetal karyotype, confirmatory treponemal testing |
A woman with postcoital bleeding and a visibly abnormal cervix does not need "screening"; she needs diagnostic assessment. A woman with a high-risk NIPT result does not have a diagnosis; she needs confirmatory invasive testing before irreversible decisions. Confusing these two steps is a high-risk clinical error.
The 2 by 2 Table
To measure how far a test should move you, you first need truth to compare against. That truth is the reference standard (sometimes loosely called the gold standard): the best available way to say whether the woman or fetus genuinely has the condition — histology for cervical disease, fetal karyotype for aneuploidy, a validated case definition for pre-eclampsia. Apply the test to a group whose true disease state is known by that standard, and every person falls into exactly one of four cells. Two cells are agreements between test and truth (true positive, true negative) and two are disagreements (false positive, false negative). Every screening statistic is just a ratio of these four numbers, so the table is worth drawing every time.
| Disease present | Disease absent | |
|---|---|---|
| Test positive | True positive (TP) | False positive (FP) |
| Test negative | False negative (FN) | True negative (TN) |
| Measure | Formula | Plain meaning | Clinical use |
|---|---|---|---|
| Sensitivity | TP / (TP + FN) | Among diseased, proportion who test positive | High sensitivity helps rule out when negative |
| Specificity | TN / (TN + FP) | Among non-diseased, proportion who test negative | High specificity helps rule in when positive |
| PPV | TP / (TP + FP) | Among test-positive people, proportion truly diseased | What a positive result means to the patient |
| NPV | TN / (TN + FN) | Among test-negative people, proportion truly not diseased | What a negative result means to the patient |
| Accuracy | (TP + TN) / all tested | Proportion of all results that are correct | Can mislead badly when disease is rare |
| False-positive rate | FP / (FP + TN) | 1 - specificity | Burden of unnecessary follow-up |
| False-negative rate | FN / (FN + TP) | 1 - sensitivity | Missed disease |
| LR+ | sensitivity / (1 - specificity) | How much a positive test increases odds | Portable effect of a positive result |
| LR- | (1 - sensitivity) / specificity | How much a negative test decreases odds | Portable effect of a negative result |
Sensitivity and specificity are calculated down the disease columns — you start by knowing who is diseased, then ask how the test behaved. Predictive values are calculated across the test-result rows — you start by knowing the test result, then ask what it means for that person. That direction matters: a clinician almost always lives in the rows (a result is in front of you and you want to know what it means), while sensitivity and specificity describe the columns and are properties of the test, not of the individual result.
Two classroom mnemonics survive into practice because they capture how the extremes are used. A highly Sensitive test, when Negative, helps rule a condition out (SnNout): with few false negatives, a negative result is trustworthy. A highly Specific test, when Positive, helps rule a condition in (SpPin): with few false positives, a positive result is trustworthy. They are shortcuts, not laws — the likelihood ratios below say the same thing more precisely and let you handle the messy middle where most real tests live.