Frequency Distributions, Centre and Spread
Before choosing a statistical test, look at the data. Descriptive statistics are not a warm-up exercise before the "real" analysis; they are the first protection against nonsense. A mean without a measure of spread is nearly useless. A percentage without a denominator is unsafe. A p value without the shape of the data can hide outliers, skew, missingness, changing definitions, and clinically important subgroups.
In O&G, distributional thinking is everywhere. Birthweight may be approximately bell-shaped within a defined gestational-age band but not across all gestations. Estimated blood loss, length of labour, duration of hospital stay, beta-hCG, and theatre waiting time are usually right-skewed. Parity is a count with many zeros and ones. Apgar score is ordinal, not truly continuous. Maternal deaths are rare event counts over live-birth denominators. Perinatal mortality combines stillbirth and early neonatal death over total births. Each requires a different summary.
Data Type Comes First
The variable type determines the display, the summary, and the later test. Many exam errors begin by treating ordinal or skewed data as if they were normally distributed continuous measurements.
| Data type | Meaning | O&G example | Usual summary |
|---|---|---|---|
| Nominal categorical | Categories without natural order | Blood group, mode of birth, HIV status | Count and proportion |
| Binary | Two-category nominal variable | Eclampsia yes/no, live birth yes/no | Risk, odds, proportion |
| Ordinal categorical | Ordered categories, unequal spacing possible | CTG category, pain score bands, cancer stage | Count by category, median sometimes |
| Discrete count | Whole-number counts | Gravidity, previous caesareans, antenatal visits | Median/IQR or count model |
| Continuous | Measured on a scale with meaningful intervals | Haemoglobin, birthweight, blood pressure | Mean/SD if symmetric; median/IQR if skewed |
| Time-to-event | Time until an event, with possible censoring | Time from PPROM to delivery, time to recurrence | Kaplan-Meier, median time, hazard ratio |
| Rate with person-time | Events over variable follow-up time | Infections per catheter-day, admissions per woman-year | Incidence rate |
The same clinical construct can be measured in different ways. Postpartum haemorrhage can be a continuous blood-loss volume, a binary threshold such as >=1000 mL, a transfusion outcome, a hysterectomy outcome, or death due to bleeding. The distribution and clinical meaning change with the measurement choice.
Frequency Distributions
A frequency distribution shows how often values occur. It should be inspected before summarising. The appropriate display depends on the data.
| Display | Best for | What to look for |
|---|---|---|
| Frequency table | Categorical variables | Counts, percentages, missing category |
| Bar chart | Nominal/ordinal categories | Category dominance, rare categories |
| Histogram | Continuous data | Shape, skew, gaps, outliers |
| Box plot | Continuous data by group | Median, IQR, outliers, group comparison |
| Dot plot | Small datasets | Every patient visible |
| Density plot | Larger continuous datasets | Smooth shape, multimodality |
| Line chart | Rates over time | Trends, seasonality, abrupt definition changes |
Distribution shapes carry clinical meaning.
| Shape | Interpretation | O&G example | Summary risk |
|---|---|---|---|
| Approximately symmetric | Values spread fairly evenly around centre | Haemoglobin in a stable antenatal cohort | Mean and SD may be appropriate |
| Right-skewed | Many low/moderate values with long high-value tail | Blood loss, hospital stay, beta-hCG | Mean overstates the typical patient |
| Left-skewed | Long low-value tail | Apgar in a mostly well cohort | Median/category summaries often better |
| Bimodal | Two peaks suggest mixed populations | Birthweight across preterm and term births | Overall mean is misleading |
| Zero-inflated | Many zero values plus positive counts | Number of previous caesareans in primigravidas | Special count thinking may be needed |
| Truncated | Values cut off by eligibility or measurement | Birthweight only among neonatal admissions | Generalisation is limited |
A bimodal birthweight distribution may simply mean preterm and term births are being mixed. The answer is not a clever p value; it is to stratify or model gestational age appropriately.
Counts, Proportions, Ratios and Rates
Descriptive statistics often fail because the denominator is vague. These terms are related but not interchangeable.
| Summary | Structure | O&G example | Trap |
|---|---|---|---|
| Count | Number of events | 12 eclamptic seizures | No denominator, so burden cannot be compared |
| Proportion | Numerator is part of denominator | 18 caesareans among 100 births | Must define denominator exactly |
| Ratio | Numerator not necessarily part of denominator | Maternal deaths per 100,000 live births | Called a ratio even though used as a risk proxy |
| Rate | Events per population-time | Infections per 1,000 catheter-days | Needs time at risk |
| Incidence risk | New events among initially at-risk people | New pre-eclampsia among booked women | Requires follow-up and defined risk period |
| Prevalence | Existing cases at a point/period | Anaemia at booking | Not incidence; no temporality |
South African audit examples are built on this distinction. The institutional maternal mortality ratio is maternal deaths in facilities per 100,000 facility live births. The perinatal mortality rate is stillbirths plus early neonatal deaths per 1,000 total births. Stillbirth rate uses stillbirths over total births; early neonatal mortality rate uses early neonatal deaths over live births. A facility can change its apparent rate by changing inclusion thresholds, referral patterns, or denominator capture without any biological change.