Study Design in Clinical Research
Start with one idea, because everything else in this chapter is built on it: every clinical study is an attempt to make a fair comparison. You want to know whether something is true of a group of women, or whether one group differs from another, or whether a treatment changes an outcome. The honest answer would come from measuring the entire population, but that is impossible, so you measure a sample and reason back to the population. A study design is simply the set of rules you put in place so that the comparison you make in your small sample tells the truth about the much larger population you actually care about.
If a study is a fair comparison, then there are only two ways it can mislead you, and the whole craft of design is built to defend against them. The first is bias: a systematic error built into how participants were chosen, measured, or followed, so the sample is not representative or the groups are not measured the same way. Bias does not shrink as the study gets bigger; a large biased study is just a confidently wrong study. The second is confounding: a third factor that travels with both the exposure and the outcome and creates an association that is not causal. Hold those two failure modes in mind. A "good design for this question" is just the design that best protects against the bias and confounding that threaten this particular comparison.
From there, design is a matter of matching the design to the verb in the question. The design is not chosen because one type of paper sounds more prestigious than another. It is chosen because the question is about frequency, aetiology, prognosis, diagnostic accuracy, intervention effect, implementation, lived experience, or evidence synthesis. A randomised trial is powerful for treatment effects but usually the wrong design for prevalence, rare harms, patient barriers, or national mortality surveillance.
In the FCOG Primary, "choose the design" questions test whether you can match the clinical question to the most defensible method and name the bias that remains. At Intermediate and Final level the same skill becomes critical appraisal: deciding whether an aspirin trial, a pre-eclampsia biomarker study, a stillbirth risk-factor paper, or a perinatal audit report should change practice. This chapter therefore builds in order — first the question, then the two failure modes (bias and confounding), then the designs from simplest to strongest, and finally how each design defends itself.
Start with the Question
The same topic can demand different designs depending on the exact question. Before naming any design, force the question into one of a small number of shapes. Is it asking how common (frequency), what is associated with or what causes (aetiology), what will happen next (prognosis), is the test accurate (diagnosis), does the treatment work (intervention effect), why does the service behave this way (implementation/experience), or what does all the evidence show (synthesis)? The verb fixes the design.
| Clinical topic | Question | Best-fit design | Main estimate |
|---|---|---|---|
| Anaemia at booking | How common is anaemia among women booking before 20 weeks? | Cross-sectional survey | Prevalence |
| Anaemia and PPH | Does booking anaemia increase PPH risk? | Cohort | Risk ratio, risk difference |
| Severe PPH | What prior exposures are associated with severe PPH requiring hysterectomy? | Case-control if rare | Odds ratio |
| PPH treatment | Does tranexamic acid reduce death due to bleeding? | Randomised controlled trial | Risk ratio, absolute risk reduction, NNT |
| PPH pathway | Why are women receiving treatment late? | Qualitative or mixed-methods study | Themes plus pathway measures |
| PPH evidence | What does all trial evidence show? | Systematic review/meta-analysis | Pooled effect, certainty |
The design follows the verb in the question: "how common," "what risk," "what causes," "does treatment work," "is the test accurate," "what predicts," "why does the service fail," and "what does the totality of evidence show."
How Comparisons Go Wrong: Bias and Confounding
Before naming designs, master the two failure modes, because they are the reason any design exists. A design is chosen and judged by how well it defends against these.
Bias: a systematic error in selection or measurement
Bias is any systematic process that makes the sample unrepresentative or measures the groups unequally. It is a property of the method, not of the play of chance, so collecting more data does not fix it. Three forms recur in O&G research and are worth recognising on sight.
Selection bias arises when the participants who enter the study are not a fair cross-section of the population the question is about, or when the comparison groups are drawn from different populations. If you study blood-pressure risk factors by sampling from a tertiary hypertension clinic rather than the community, your sample is sicker and unrepresentative by construction. A subtler version: if you study survival after an operation using archived records, and the notes of women who died are stored or retrieved differently from the notes of survivors, the very availability of a record depends on the outcome — and the comparison is poisoned.
Observer (and responder) bias arises during measurement. If the assessor knows which group a woman is in, their judgement of a soft outcome — "was this pre-eclampsia?", "was the CTG reassuring?" — drifts toward what they expect. The participant's side of the same problem is recall bias: a woman asked about past exposures answers differently depending on whether she already has the outcome. A mother whose baby has a malformation searches her memory far harder for a first-trimester exposure than a mother of a healthy baby, manufacturing an association out of differential remembering rather than differential exposure. Recall bias is most corrosive in designs that ask about the past after the outcome is already known.
Confounding: a third factor that mimics a real effect
Confounding is different from bias. The data may be collected perfectly, with no selection or measurement error, and you can still be fooled. A confounder is a variable that is associated with both the exposure and the outcome, and is not simply a step on the causal pathway between them. It creates a "mixing of effects" — an association in the data that is not the causal effect you think you are measuring.
The clean way to hold this is the confounding triangle: three corners are exposure, outcome, and the candidate confounder, joined by three arrows. To be a confounder, a variable must be linked to both the exposure and the outcome. Age is the archetype, because almost every disease varies with age and so do many exposures: if older women are over-represented among the "exposed" group, age alone can manufacture an outcome difference. Crucially, a variable is only a confounder for a particular exposure-outcome pair. When studying obesity and pre-eclampsia, smoking is a confounder, because smoking is linked to both obesity and pre-eclampsia risk — all three arrows exist. But when studying obesity and gestational diabetes, smoking is not a confounder, because smoking is not meaningfully associated with gestational diabetes — one arrow is missing, the triangle is broken, and adjusting for smoking would be unnecessary.
You handle confounding in one of two ways. In observational studies you can only handle the confounders you thought of and measured — by restriction, matching, stratification, or statistical adjustment (multivariable regression). The honest limitation is permanent: you can never adjust for a confounder you did not measure or did not imagine, so residual confounding always remains and any single observational study supports association more than causation. In randomised studies you handle confounding by design, and this is the central reason randomisation is so powerful — addressed below.
Confounding versus interaction (effect modification)
These two are routinely confused and the distinction is examinable. Confounding is a nuisance you want to remove, because it distorts a single underlying effect. Interaction (effect modification) is a real finding you want to describe, because the effect of one factor genuinely differs across levels of another. If obesity alone trebles pre-eclampsia risk and smoking alone doubles it, you might naively expect both together to multiply to a sixfold risk; if the observed combined risk is tenfold, obesity and smoking are interacting — the joint effect exceeds the product of the separate effects. Confounding asks "is this association spurious?"; interaction asks "does the effect change depending on who you are?" You adjust away a confounder; you report and explore an interaction. Beware, though, that hunting for interactions in subgroups is statistically inefficient and prone to false positives, which is why formal interaction tests are preferred over eyeballing subgroup differences.
Descriptive Designs
With the failure modes in hand, the designs ascend from simplest to strongest. Descriptive designs come first because they only quantify patterns; they do not primarily test causal effects. They are indispensable in public health and audit.