Core principle
A case-control study starts with outcome status: cases have the outcome/disease and controls do not. The study then compares prior exposures. The controls must represent the exposure distribution in the population that produced the cases. Poor control selection is one of the most damaging errors in this design.
Approval and governance
Usually needs formal review / authorisation
- Access to identifiable historical records, registries, laboratory systems or imaging commonly requires ethics/governance and data-access authorisation.
- Prospective contact with cases or controls generally requires review and consent unless an authorised body approves otherwise.
- Linkage across datasets requires explicit authorisation and a secure linkage plan.
May follow a lighter or different route
- Fully anonymised secondary datasets may have a different review route depending on provenance and terms.
- Some registry studies may have pre-existing consent/governance frameworks, but the proposed analysis still needs the applicable institutional determination.
Do not do this
- Do not choose controls because their exposure status is convenient or known.
- Do not use a control group that could never have become a case in the source population.
- Do not extract identifiable data before the approval/data-access route is complete.
Mediclinic Middle East publicly states that research projects carried out at MCME are to receive approval from its internal Research and Ethics Committee and applicable local regulatory authorities before initiation. The exact route varies by project, facility and emirate. Dubai projects may involve DSREC depending on applicability. This hub must therefore route users to the Research Office and current local forms rather than declaring a project “ethics exempt.” Institution-specific forms and contacts will be inserted after verification.
Step-by-step workflow
Define the outcome and source population
Write an operational case definition and identify the population/time period from which cases arose. Example: all adult patients treated at specified facilities during 2024–2026 who met validated diagnostic criteria.
Define incident versus prevalent cases
Incident cases are newly diagnosed during the study period and often reduce survival/prevalence bias. Prevalent cases may be easier to identify but can overrepresent survivors. State which is used and why.
Choose the control source before looking at exposure
Controls should be sampled from people who would have been eligible to become cases if they developed the outcome. Hospital controls can be biased if the exposure causes the control diagnoses; community controls may differ in healthcare access. Justify the choice.
Set the case-to-control ratio and matching plan
More than one control per case can improve efficiency up to a point. If matching, choose only necessary factors, define exact matching rules and plan matched analysis. Do not match on a factor that is a consequence of exposure or lies on the causal pathway.
Define exposure measurement identically for cases and controls
Use the same source, time window and coding rules. If reviewers can infer case status while assessing ambiguous exposure, measurement bias may result; blind assessors where feasible.
Identify confounders using subject knowledge
Predefine variables plausibly associated with both exposure and outcome and not on the causal pathway. Avoid automatic variable selection as a substitute for causal reasoning.
Calculate/justify sample size
Specify anticipated exposure prevalence among controls, target odds ratio, case-control ratio, alpha/power and inflation for matching or missing data where relevant.
Obtain approvals and create extraction forms
Submit case/control definitions, record-query logic, matching variables, extraction form, consent/waiver request if applicable, and data-security plan.
Build a screening log
Record candidate cases/controls, eligibility, exclusion reasons and matching status without exposing unnecessary identifiers in the analytic file.
Clean and lock the dataset
Check duplicates, date relationships, impossible exposure windows and matching integrity. Freeze a clean analysis version before the primary analysis.
Describe cases and controls
Report source, numbers, eligibility, demographics, exposure data completeness and relevant clinical characteristics separately.
Estimate the association
The principal measure is usually the odds ratio with confidence interval. Use unconditional logistic regression for unmatched designs and conditional methods for individually matched designs as appropriate. Adjust only for prespecified confounders and clearly label exploratory models.
Test robustness
Consider sensitivity analyses for exposure misclassification, alternative control definitions, missing data or different confounder sets when scientifically justified.
Write using STROBE
Explain how cases were ascertained, how controls were selected, why controls represent the source population, how matching was handled, how exposure was measured, how confounding and missing data were addressed, and the study’s susceptibility to selection and recall bias.
Common errors
- Control selection unrelated to the source population.
- Matching without using matched analysis.
- Overmatching on variables related to exposure.
- Using post-outcome information as if it were a pre-outcome exposure.
- Differential exposure ascertainment between cases and controls.
- Interpreting an odds ratio as a risk ratio when the distinction matters.
- Claiming causality from an observational association.
Final checklist
- Case definition is reproducible.
- Source population defined.
- Control selection justified independently of exposure.
- Matching rules and analysis aligned.
- Exposure window occurs before the outcome when causality is discussed.
- Confounders prespecified.
- Approval/data access documented.
- Odds ratios include confidence intervals.
- STROBE checklist completed.