Use this design when
- You need the prevalence of a condition, behaviour, exposure or characteristic in a defined population.
- You want to examine associations between variables measured at roughly the same time.
- You have a survey or clinical dataset representing a defined period and population.
Do not use it to claim
- That an exposure caused an outcome when temporality is unknown.
- Incidence over time unless the design truly follows participants.
- Risk ratios that imply future risk when only current status was measured.
Approval and governance
Usually needs formal review / authorisation
- Prospective surveys involving patients/staff and collection of research data commonly require ethics/governance review.
- Extraction of identifiable clinical data requires authorised data access and usually ethics/governance review or a documented waiver/determination.
- Collection of sensitive personal information requires a privacy and security plan.
May follow a lighter or different route
- Analysis of genuinely anonymous, openly available aggregate data may not be human-participant research in some systems, but terms of use and institutional policy still apply.
- Some anonymous service-evaluation surveys may be routed as QI/service evaluation rather than research; obtain the official classification.
Do not do this
- Do not email yourself patient spreadsheets or place identifiable data in personal cloud storage.
- Do not state “anonymous” when you collect combinations of variables that allow re-identification.
- Do not begin recruitment or extraction while approval is still pending.
Mediclinic Middle East publicly states that research projects carried out at MCME are to receive approval from its internal Research and Ethics Committee and applicable local regulatory authorities before initiation. The exact route varies by project, facility and emirate. Dubai projects may involve DSREC depending on applicability. This hub must therefore route users to the Research Office and current local forms rather than declaring a project “ethics exempt.” Institution-specific forms and contacts will be inserted after verification.
Step-by-step workflow
Define the primary objective
Choose either a prevalence objective or an association objective as primary. Example prevalence question: “What proportion of adults attending clinic X during period Y meet criterion Z?” Example association question: “Among adults attending clinic X, is exposure A associated with outcome B?”
Define the target population and sampling frame
State who you want to generalise to and the actual list/process from which participants will be selected. Hospital attendees, all registered patients, staff, and community residents are different populations.
Set inclusion/exclusion criteria before sampling
Use criteria tied to the question, not to whether records are complete or results are convenient. Decide how repeat visits and duplicate patients will be handled.
Choose a sampling strategy
- Census: include every eligible person/record in the period.
- Simple/random sampling: every eligible unit has a known chance of selection.
- Systematic sampling: e.g., every kth eligible record after a random start.
- Stratified sampling: sample within predefined groups to ensure representation.
- Convenience sampling: easier, but document the high risk of selection bias.
Calculate or justify the sample size
For prevalence, specify the expected prevalence, desired precision, confidence level and design effect if relevant. For association analyses, base the calculation on the primary effect/outcome and planned model. Avoid “we included whoever was available” unless feasibility is explicitly the design constraint.
Define every variable
Create a data dictionary stating variable name, definition, source, coding, units, allowable range, missing code and whether it is exposure, outcome, confounder or descriptive variable.
Choose validated measurement tools where possible
For questionnaires or scales, record the version, language, scoring rules and evidence of validity/reliability in the target setting. If translating or adapting a tool, plan the validation process rather than silently modifying items.
Plan bias control
- Selection bias: make sampling and non-response visible.
- Information bias: standardise measurement and train data collectors.
- Recall bias: minimise long recall windows when possible.
- Confounding: identify plausible confounders from subject knowledge before analysis.
Obtain approvals and pilot
Submit the protocol, questionnaire/extraction sheet, consent materials and data-security plan. After approval, pilot the workflow on a small number to detect ambiguous questions, impossible fields or coding problems.
Collect data with a screening log
Track eligible, approached, consented/included, declined and excluded participants where applicable. Use unique study IDs and separate identifiers from analytic data.
Perform data quality checks before analysis
Check duplicates, impossible dates, out-of-range values, inconsistent units, missingness and skip logic. Document any corrections and retain the original raw dataset read-only.
Analyse prevalence correctly
Report numerator and denominator, prevalence proportion and confidence interval. If sampling was clustered, stratified or weighted, use analysis that reflects the sampling design.
Analyse associations transparently
Report unadjusted and prespecified adjusted estimates with confidence intervals. Choose models that match the outcome and sampling design. State which confounders were included and why. Do not select confounders solely from univariable p-values.
Address missing data and sensitivity
Report missingness for important variables, explain the primary approach, and perform sensitivity analyses when missing data could materially change conclusions.
Write using STROBE
Report design, setting, dates, participant selection, variables, data sources, bias, study size, statistical methods, participant flow, missing data, estimates with precision, limitations and generalisability.
Common errors
- Calling a convenience sample representative of the whole hospital/community.
- Reporting only p-values without effect estimates and confidence intervals.
- Using “risk” or causal language when temporality cannot be established.
- Not defining the denominator for prevalence.
- Changing questionnaire scoring after seeing results.
- Treating repeated visits from one patient as independent people.
- Ignoring non-response and missing data.
Final checklist
- Target population and sampling frame are defined.
- Primary objective is explicit.
- Eligibility and sampling method prespecified.
- Sample size justified.
- Variables and confounders defined before analysis.
- Approvals/data permissions documented.
- Missing data quantified.
- Effect estimates include precision.
- STROBE checklist completed.