A cohort study follows a group of people defined by their exposure status and tracks them forward in time to see who develops the outcome. Because exposure is recorded before the outcome appears, this observational study design can establish temporal order, measure incidence, and estimate relative risk, which is exactly what a single-snapshot design cannot do. That forward direction is the reason cohort studies sit near the top of the observational evidence hierarchy.
Why direction of time is the whole point
In a cohort study you start with people who do not yet have the outcome, classify them as exposed or unexposed, and then wait. Because the exposure is measured first, an association you observe later carries the one thing a cross-sectional study cannot supply: the knowledge that the cause preceded the effect. This is what lets a cohort estimate the incidence rate, the risk ratio, and the hazard ratio, and it is why cohort evidence is so persuasive when a randomized trial is impossible or unethical. You cannot randomize people to smoke; you can follow smokers and non-smokers and count the cancers.
Prospective and retrospective cohorts
The two flavors differ only in when the investigator enters the timeline.
- A prospective cohort study identifies exposure now and follows participants into the future. It gives the cleanest data because the researcher decides in advance what to measure and how, but it is slow and expensive, sometimes running for decades.
- A retrospective cohort study uses existing records to reconstruct an exposed and unexposed group whose outcomes have already occurred. It is faster and cheaper because the follow-up has, in effect, already happened, but it is limited to whatever the historical records captured.
Both are genuine cohorts because both start from exposure and move toward outcome. The retrospective version is not a case-control study; the direction of reasoning is still exposure to outcome, only the calendar is different.
What a cohort study measures that others cannot
The headline metric is incidence: the rate at which new cases appear in the exposed compared with the unexposed. From that you compute a relative risk, the ratio of those incidences, which has a direct and honest interpretation that the odds ratio only approximates. When you need to turn counts into a comparative risk, our relative risk calculator does the arithmetic and shows the confidence interval. Cohorts also support time-to-event analysis, so a single study can report both whether an outcome happened and how quickly.
The cost of following people: loss to follow-up
The signature threat to a cohort study is attrition, the loss of participants before the outcome is observed. If the people who drop out differ systematically from those who remain, and especially if dropout is related to both exposure and outcome, the result is biased in a way no analysis can fully repair. A well-run cohort plans retention from the start, documents every loss, and reports how the lost participants compared with the retained ones. Differential loss to follow-up is to a cohort what poor sampling is to a survey: the quiet failure that undermines everything downstream.
Confounding and how cohorts handle it
Because exposure is not randomized, the exposed and unexposed groups may differ in ways that also affect the outcome. Confounding is the central analytic problem of any observational comparison. Cohort studies address it by measuring potential confounders at baseline and adjusting for them with multivariable regression, and increasingly with methods such as propensity scoring that balance the groups before the outcome is modeled. The credibility of a cohort's causal claim rests on how completely it anticipated and measured the confounders, which is a study-design decision long before it is a statistics decision.




