MediclinicResearch Hub
RESEARCH ACADEMY · PROCEDURAL GUIDE

Statistics Planning

A pre-analysis framework for defining outcomes, estimands, models, missing data and sensitivity analyses before results are known.

9 sections · Procedural guidance

Before opening the final dataset

  1. Write the primary research question and primary outcome.
  2. Define the unit of analysis: patient, visit, eye, lesion, admission, clinician, facility, etc.
  3. List all variables with type, coding, units and valid ranges.
  4. Define the primary effect measure/estimand.
  5. Specify descriptive statistics.
  6. Specify the primary model/test and assumptions.
  7. Prespecify confounders/covariates using subject knowledge.
  8. Plan handling of repeated measures, clustering or matching.
  9. Define missing-data strategy.
  10. Define multiplicity/subgroup strategy.
  11. Define sensitivity analyses.
  12. Lock/version the statistical analysis plan before primary outcome analysis.

Core rules

  • Report effect estimates and confidence intervals, not only p-values.
  • Do not choose a statistical test because it gives a smaller p-value.
  • Distinguish statistical significance from clinical importance.
  • Do not categorise continuous variables without a defensible reason; arbitrary cut points can lose information and introduce bias.
  • Account for non-independence when patients contribute multiple observations or data are clustered by facility/clinician.
  • Use models that match the outcome distribution and study design.
  • Document data exclusions and transformations before analysis.

Missing data

For each important variable, quantify missingness and investigate patterns. Complete-case analysis is not automatically unbiased. The appropriate approach depends on the missingness mechanism, model and estimand. Prespecify the primary approach and use sensitivity analysis when assumptions are uncertain.

When you need a statistician

  • Before sample-size calculation for anything beyond a simple descriptive study.
  • Matched/clustered/repeated-measures designs.
  • Survival/time-to-event outcomes.
  • Prediction models or diagnostic models.
  • Multiple outcomes/time points with complex multiplicity.
  • Meta-analysis with uncommon effect measures or dependent data.
  • Missing-data methods beyond simple cases.
  • Any analysis you cannot explain in plain language before running it.

Final checklist

  • Primary outcome/estimand defined.
  • Unit of analysis correct.
  • Sample-size rationale documented.
  • Model chosen before results.
  • Confounders prespecified.
  • Clustering/repeated measures handled.
  • Missing-data plan defined.
  • Sensitivity analyses planned.
  • Effect sizes and CIs will be reported.

Descriptive statistics

  • For categorical variables, report counts and denominators/percentages; make missing values visible.
  • For approximately symmetric continuous variables, mean and standard deviation may be suitable; for skewed variables, median and interquartile range often communicate distribution better.
  • Do not choose summaries solely from a normality test. Inspect distributions and consider the scientific scale.
  • Always state units. Never mix mmol/L and mg/dL, days and months, or percentages and proportions silently.
  • For repeated observations, distinguish number of observations from number of independent participants.

Regression/model planning

  • Specify outcome, link/model family, predictors, interactions and adjustment variables before running the model.
  • Check whether continuous predictors require non-linear terms rather than arbitrary categorisation.
  • Assess model assumptions and influential observations using methods appropriate to the model.
  • Avoid automated stepwise selection as the primary strategy for causal adjustment.
  • Report coefficients/effect estimates with uncertainty and translate them into clinically interpretable quantities when possible.
  • For prediction studies, evaluate calibration and discrimination and separate model development from validation.

Reproducibility

Keep a read-only raw dataset, a scripted cleaning file when possible, an analysis-ready dataset and analysis code/output with version dates. Manual spreadsheet edits should be documented. A colleague should be able to rerun the primary analysis from the preserved files.

Primary standards and sources