CDC PLACES · 168 New York City ZIP codes

The dataset

This is the whole file, all semester. Every question we ask comes back to it. Sort it, search it, and look at how the answer changes depending on which column you rank by.

Rank the city two ways

Same 168 neighborhoods, same file. Press one, then the other, and watch which places move to the top.

What the columns mean

Read this before you quote a number

These are modeled small-area estimates, not a count of people who were actually asked. The CDC uses a national survey to estimate what each ZIP code probably looks like. So say about and roughly, never exactly. We take this apart properly in Week 12.

Identifying the place
zipZIP code. A label, not a quantity — ZIP 11436 is not "more" than 11432.
boroughWhich of the five boroughs the ZIP sits in.
adultsEstimated number of adults living there. This is the "out of whom."
Frequent mental distress
distress_pctPercent of adults reporting 14+ days of poor mental health in the past 30. Our main measure.
distress_casesEstimated number of adults, not a percent. Big neighborhoods have big numbers.
distress_low
distress_high
The range the estimate could plausibly sit in. When two ranges overlap, be careful about claiming a difference is real — that's Week 9.
Depression
depression_pct
depression_cases
depression_low
depression_high
Same four shapes, for diagnosed depression. Our secondary measure.
Everything else — for comparison only
no_insurance_pct
poor_health_pct
no_checkup_pct
inactive_pct
smoking_pct
binge_pct
short_sleep_pct
diabetes_pct
Other health measures for the same ZIP codes. We use these to ask whether something else might explain a pattern — that's confounding, in Week 11.

Getting the file

Download it: nyc_mental_health_by_zip.csv — 168 rows, 19 columns.

To load it in Colab, this is the line:

import pandas as pd
df = pd.read_csv("https://ph320.vercel.app/data/nyc_mental_health_by_zip.csv")
df.head()

You are not required to run anything. Notebooks are prepared and shown in class.