Today's lesson
You just read four studies and worked out how each one was built. Every one of them answers the same question in a different way: when did you look, relative to the outcome?
1 The picture
Draw a line. Mark a point partway along it — the outcome happens here.
Everything before the mark is the life that led up to it: the exposures. Everything after is what we can observe once it has happened. The four designs are four places to stand on that line.
Same line every time. The only thing that changes is where the researcher stands, and which direction they face.
2 The four
We'll walk the line four times. Same outcome, same mark — the only thing that changes is when you showed up and what you were allowed to do.
You measure exposure and outcome at the same time, in the same people, on one occasion.
Picture yourself standing right at the mark. You look around once and write down everything you see — who exercises, who is depressed, all of it in the same glance. Nothing tells you which of those things showed up first, because you only looked once.
That single glance is the whole design, and it is also the whole limitation.
It lets you say
"People with X also have Y." "Twenty-three percent of adults in this county report frequent distress."
It does not let you say
Which came first, or whether one caused the other.
Why anyone uses it: fast and cheap, and it is the right design when you want to know how common something is right now. Most public health surveillance is cross-sectional.
You start with people who don't have the outcome yet, record their exposure, then follow them forward and see who develops it.
Here you show up early. Everybody you enroll is still fine, and you write down what they are exposed to before anything has happened to them. Then you walk forward alongside them and watch to see who develops the outcome and who doesn't.
You wrote the exposure down before anyone got sick, so you know it came first. It is a receipt. That certainty about the order is the only thing the eight years buy you — and it is the one thing a survey taken once can never have.
It lets you say
"People who did X went on to develop Y more often than people who didn't."
It does not let you say
That X caused Y. The two groups chose themselves, and may differ in a hundred ways nobody measured.
What it costs: years, money, and people who drop out. And it is hopeless for rare outcomes — follow 10,000 people for a disease that strikes 1 in 50,000 and you will end up with nobody.
You start from the outcome — people who have it, plus comparable people who don't — then look backward at what each group was exposed to.
This time you arrive after it has already happened. You gather people who have the outcome, then find comparable people who don't, and you work backward through both groups' histories looking for what was different.
You are reconstructing the past rather than watching it, and everything weak about this design comes from that.
It lets you say
"X appears more often in the histories of people with Y."
It does not let you say
How common the outcome is. You chose how many sick people to include, so any percentage you calculate is a product of your own sampling.
Why anyone uses it: it is the only practical design for rare outcomes. Instead of following 50,000 people for twenty years, you find the 200 who already have it. Where it is weak: recall — you are asking people to remember, and people with a diagnosis search harder for a reason. Choosing fair controls is genuinely difficult too.
You assign the exposure yourself, at random, then follow both groups forward.
You show up early again, but this time you don't just watch. You decide who gets exposed, and you decide it by coin flip. Then you follow both groups forward exactly as you would in a cohort.
Everything that makes this design powerful comes from that one act of assigning.
It lets you say
"X caused Y" — in this population, under these conditions, and only if the trial was run well.
It cannot do
You cannot randomize poverty, smoking, a ZIP code, or a childhood. Most public health questions can never be a trial — which is exactly why the other three designs exist.
Why randomization is the whole trick: because a coin flip decided who got exposed, the two groups differ only by chance — including on things nobody thought to measure. No observational design can buy you that. It still has to be run well: people drop out, and in Study D everyone knew which group they were in.
3 The test
Ask these in order, and stop at the first yes.
Did the researchers assign the exposure themselves?
Yes → randomized controlled trial
Did they start with people who didn't have the outcome, and follow them forward?
Yes → cohort
Did they start from people who already have the outcome?
Yes → case-control
Did they measure everything at once?
Yes → cross-sectional
The one people mix up is cohort versus case-control, and one question settles it: what did you start from — the exposure, or the outcome?
4 Side by side
Does social media use cause depression in teenagers?
| Design | What you'd do | What you could say |
|---|---|---|
| Cross-sectional | Survey 3,000 teens once about usage and mood | Heavy users report more depression. They differ — order unknown. |
| Cohort | Enroll 3,000 teens without depression, record usage, follow 3 years | Heavy users went on to develop depression more often. |
| Case-control | 300 teens in treatment, 300 not, ask about past usage | More usage in the histories of the treated group. |
| RCT | Randomly assign 600 teens to cut usage to 30 minutes for 8 weeks | Cutting usage reduced symptoms. |
Same question. Four studies. Only the last one earns the word cause.
5 Our own file
Our dataset is cross-sectional. Every ZIP code, measured once.
So the semester question — why is frequent mental distress higher in Jamaica than in most of New York City — can be described with this data. It cannot be explained by it.
A row in our file is a ZIP code, not a person. So "ZIP codes with less insurance have more distress" is a fact about ZIP codes — not about uninsured people.
Knowing that isn't a failure. It's the difference between an analyst and someone who just runs the numbers.
6 Leaving with this
What a study is allowed to claim is decided by how it was built — before anybody calculates anything.