Start with the decision, not the dataset
A useful research question states what you want to learn before you decide which convenient variables to analyse. The question should be important, answerable with a defensible design, and narrow enough that the primary outcome and population are unambiguous.
Step-by-step workflow
Write the clinical/research problem in plain language
State what is uncertain and who is affected. Avoid beginning with a method such as “I want to run a regression.”
Search enough to establish the gap
Identify recent systematic reviews, guidelines and key primary studies. Distinguish a true unanswered question from a question that is already well answered.
Choose the question framework
- PICO: population, intervention, comparator, outcome — often useful for intervention questions.
- PECO: population, exposure, comparator, outcome — useful for many observational questions.
- PCC: population, concept, context — commonly used to frame scoping-review questions.
- Qualitative questions may instead focus on experiences, meanings, processes or contexts rather than forcing a PICO structure.
Define the population precisely
Specify setting, age/condition where relevant, time period and major eligibility boundaries. “Adults with diabetes” may still be too broad.
Define one primary outcome
Specify exactly how and when it will be measured. Avoid several competing “primary” outcomes unless the design explicitly supports them.
Check feasibility
Estimate access to eligible participants/records/studies, expected sample size, follow-up, measurement availability, expertise and project timeline.
Choose the design only after the question is stable
Use the design that answers the question with the least avoidable bias and feasible resources. Use the Start Research wizard if the design is unclear.
Write primary and secondary objectives
The primary objective should map directly to the primary outcome and main analysis. Secondary objectives should be explicitly secondary.
Question quality test
- Population is defined.
- Exposure/intervention/phenomenon is defined where relevant.
- Comparator is explicit where relevant.
- Primary outcome is measurable and time-framed.
- The question identifies a genuine knowledge or local evidence gap.
- The proposed study can realistically answer the question.
- The question can be stated in one sentence without vague terms.
Example transformation
Too broad
“Does obesity affect surgery outcomes?”
More answerable
“Among adults undergoing [specified procedure] at [defined setting/time], is [defined obesity exposure] associated with [specified postoperative outcome] within [specified follow-up], compared with [defined comparator]?”
Governance checkpoint
The question itself does not determine the approval route. Once the proposed population, data source, recruitment, intervention and identifiability are defined, route the protocol through the appropriate institutional ethics/governance process before beginning activities that require authorisation.
Feasibility and ethics test before locking the question
Feasible
Can the team access enough eligible participants, records or studies in the available time? Are the required measurements routinely available and sufficiently reliable?
Scientifically useful
Will the answer change understanding, practice, policy, future research or a meaningful local decision? A statistically analysable question is not automatically an important one.
Ethically defensible
Can the question be answered without exposing participants to unjustified risk or collecting unnecessary identifiable information? The ethics committee/institution makes the formal determination where review is required.
From question to protocol: write these sentences
- Gap: “Existing evidence does not adequately establish … in …”
- Primary objective: “To determine/estimate/compare/explore …”
- Primary outcome: define the exact measurement, unit/scale and assessment time.
- Population: define inclusion/exclusion boundaries and setting.
- Main comparison/exposure: define how group membership or exposure will be measured.
- Design: name the design and explain why it answers the question.
If these statements cannot be written consistently, the project is not yet ready for a final protocol or sample-size calculation.
Common question-design failures
- Starting from a convenient dataset: this encourages aim-switching and data-driven hypotheses. Start from the question, then confirm whether the data can answer it.
- Outcome is vague: “clinical improvement” must be converted into a defined measure and time point.
- Comparator is implicit: clarify what the exposed/intervention group is being compared against.
- Population is too broad: heterogeneous settings or disease stages can make interpretation meaningless.
- Several questions are competing for primary status: choose one primary objective and label the rest secondary/exploratory.
- Causal language with a non-causal design: phrase the question so the design can support the inference you plan to make.