01Core principlesThe concepts and mechanisms needed to understand the subject.
Medical statistics begins with the data-generating question. Prevalence is the proportion with a state at a defined time or period; cumulative incidence is the proportion initially at risk who develop the event over a stated interval; an incidence rate divides new events by person-time. These quantities are not interchangeable. A busy clinic can have high prevalence because disease lasts longer even if incidence is unchanged, and person-time rates permit unequal observation but do not directly state an individual's interval risk.
For binary outcomes, calculate event risk separately in exposed and comparison groups. The risk ratio is exposed risk divided by comparison risk. The risk difference subtracts those risks and preserves baseline context. Relative risk reduction equals one minus the risk ratio when lower risk is beneficial. Number needed to treat is one divided by absolute risk reduction, with direction and time horizon stated; round conventionally upward for benefit. Odds are event probability divided by non-event probability, so odds ratios increasingly diverge from risk ratios as events become common.
Sampling uncertainty surrounds every estimate. A standard error measures expected sampling variation of an estimator under assumptions; a confidence interval combines estimate and uncertainty on an appropriate scale. For ratios, the null is one; for differences, zero. A wide interval may contain important benefit and harm even when the p value exceeds 0.05. Conversely, a huge study may estimate a clinically trivial difference very precisely. Repeated testing, selective outcomes and data-driven subgroups raise false-positive risk unless design and analysis address multiplicity.
Association requires causal caution. Random allocation protects against baseline confounding on average when concealment and follow-up are sound. Observational adjustment controls only measured, correctly modelled confounders and can introduce bias if it conditions on consequences of exposure. Selection into a study or analysis can create collider bias. Measurement error may dilute, inflate or redirect estimates. Statistical modelling cannot repair an inappropriate comparator, informative missingness or systematic outcome misclassification without additional assumptions and evidence.
Key points
- Define the population, unit, time horizon, numerator and denominator before calculating prevalence, risk, rate or an effect measure.
- Risk ratio compares probabilities; odds ratio compares odds; risk difference gives absolute change; number needed to treat is the reciprocal of a non-zero absolute risk reduction.
- A confidence interval expresses precision under model assumptions; it is not the probability that this one fixed interval contains the true value.
- A p value is the probability of data at least as incompatible with the null under the model, not the probability that the null hypothesis is true.
- Statistical significance does not establish clinical importance, absence of bias, correct design or causality.
- Random error decreases with information, while selection bias, measurement bias and confounding can persist or intensify in a very large study.
- Inspect missing data, multiplicity, outcome switching, subgroup credibility and absolute event counts before accepting a headline estimate.
02Mechanisms and patternsImportant relationships and how to distinguish them.
Risk uses people initially at risk, rate uses person-time, and prevalence uses the population assessed; changing the denominator changes the estimand.
Risk ratios and odds ratios compare groups proportionally but conceal how common the outcome is without accompanying absolute risks.
Risk difference translates a relative effect through baseline risk, so the same relative effect can imply very different clinical benefit across populations.
Confidence intervals widen with less information or greater variability; narrow intervals do not protect against systematic bias.
A common cause of exposure and outcome can generate or mask association unless design or analysis adequately accounts for it.
Effect modification means the effect truly differs across a factor; it should be assessed with a direct interaction test rather than separate subgroup p values.
03Interpreting evidenceInformation, measurements and their limitations.
Consider the information, its meaning and its limitations before deciding what follows.
- 01
Two-by-two table - Why
- Expose event and non-event counts in intervention or exposure and comparison groups.
- Interpretation and limitations
- Calculate risks, odds and absolute differences from the cells; sparse cells create unstable estimates and continuity corrections can materially affect results.
- 02
Confidence interval - Why
- Show a range of estimates compatible with data and model at the chosen confidence level.
- Interpretation and limitations
- Assess width and clinically important boundaries, not only null inclusion; bias and model misspecification are not represented automatically.
- 03
Hypothesis test and p value - Why
- Quantify incompatibility between observed data and a specified null model.
- Interpretation and limitations
- The result depends on test assumptions, sample size and analysis plan; it does not measure effect size, clinical value or truth probability.
- 04
Adjusted regression estimate - Why
- Estimate association while conditioning on prespecified covariates.
- Interpretation and limitations
- Check causal rationale, functional form, events per parameter and missing data; adjustment cannot remove unmeasured confounding by declaration.
- 05
Heterogeneity assessment - Why
- Evaluate variation in study effects beyond the point estimates in a synthesis.
- Interpretation and limitations
- Inspect clinical and methodological differences with forest plots and interval estimates; I-squared alone does not explain the cause or importance of heterogeneity.
- 06
Sensitivity analysis - Why
- Test whether conclusions survive plausible alternative assumptions or analytic decisions.
- Interpretation and limitations
- Credibility is stronger when alternatives were prespecified and address a real uncertainty; many selective analyses can instead enable cherry-picking.
04Applied reasoningWorked examples connecting principles to decisions.
01Worked exampleCalculate relative and absolute treatment effectsInputs: over two years, 18 of 300 control participants and 9 of 300 intervention participants experience the outcome; calculate risks, RR, ARR and NNT.+
- 1Calculate control risk: 18 divided by 300 equals 0.06, or 6%; intervention risk: 9 divided by 300 equals 0.03, or 3%.
- 2Calculate risk ratio: 0.03 divided by 0.06 equals 0.50, corresponding to a 50% relative risk reduction.
- 3Calculate absolute risk reduction: 0.06 minus 0.03 equals 0.03, or 3 percentage points over two years.
- 4Calculate NNT: one divided by 0.03 equals 33.33, rounded upward to 34 people treated for two years to prevent one additional outcome, assuming the estimate applies.
- 5State outcome with both scales: risk fell from 6% to 3%, RR 0.50, ARR 3 percentage points and NNT 34 over two years; uncertainty requires confidence intervals not supplied here.
- 6Verify with counts: treating 300 people produced nine fewer events, and 300 divided by nine equals 33.33, agreeing with the reciprocal calculation before upward rounding.
02Study interpretationRead beyond statistical significanceA large trial reports RR 0.98 with 95% CI 0.97 to 0.99 and p below 0.001.+
- 1Confirm outcome definition, time horizon, allocation, missing data and whether the analysis followed the prespecified plan.
- 2Translate the relative effect using representative baseline risk to obtain an absolute difference.
- 3Compare the effect and interval with thresholds for meaningful benefit and known harms or burden.
- 4Conclude on precision, clinical importance and risk of bias separately rather than treating the p value as a verdict.
03Observational appraisalInterrogate an adjusted associationA cohort reports that exposure is associated with an outcome after multivariable adjustment.+
- 1Draw the presumed causal relations among exposure, outcome, common causes and consequences before reading the covariate list.
- 2Check selection, measurement timing, loss to follow-up and whether adjusted variables were measured accurately before exposure.
- 3Compare crude and adjusted estimates with confidence intervals and assess residual or unmeasured confounding.
- 4Describe the result as an association unless design, assumptions and corroborating evidence support a causal interpretation.
05Checking understandingVerify the reasoning, revisit uncertainties and apply feedback.
- Reproduce headline calculations from raw counts whenever they are available and reconcile any discrepancy with the reported analysis population.
- Check that every effect estimate retains its outcome definition, comparator and follow-up period when transferred into notes or decisions.
- Review protocols or registrations for prespecified primary outcomes, subgroups and analysis plans before interpreting multiplicity.
- Track missing participants and missing outcomes by group, and test how plausible departures from missing-at-random assumptions change conclusions.
- When applying an estimate, update absolute effects using the relevant baseline risk and verify that population and care context are sufficiently similar.
06Special situationsVariants, exceptions and circumstances that change the usual approach.
Rare-outcome approximation
An odds ratio approximates a risk ratio when outcomes are rare, but the approximation can exaggerate apparent risk change when outcomes are common.
NNT is context-bound
Number needed to treat changes with baseline risk, follow-up and effect estimate; it must carry all three rather than appearing as a timeless property.
Non-significance is not equivalence
Failure to reject a null may reflect inadequate precision. Equivalence or non-inferiority requires prespecified margins, suitable design and interval-based analysis.
Adjustment can harm
Conditioning on a mediator blocks part of the effect, while conditioning on a collider can create association; more covariates do not guarantee less bias.
Reporting guidance aids appraisal
CONSORT and STROBE improve transparent reporting but adherence does not itself prove that design, conduct or causal inference was valid.
07Common pitfallsFrequent interpretation and management errors.
- 01
Quoting a relative reduction without baseline risk, absolute difference, follow-up period or uncertainty.
- 02
Interpreting a p value as the probability that the null hypothesis or study conclusion is true.
- 03
Calling two subgroup effects different because one is statistically significant and the other is not, without an interaction test.
- 04
Treating a narrow confidence interval as protection from selection bias, confounding or misclassification.
- 05
Using an odds ratio as if it were a risk ratio for a common outcome and overstating the change in probability.
- 06
Calculating NNT from a rounded percentage without preserving direction, time horizon and confidence interval.