df = pd.read_csv("nyc_mental_health_by_zip.csv")Today's lesson
Last week we asked who answered the question. This week we open the file their answers ended up in — and find out what each column will and won't let us do.
1 Start here
Every row is one thing you measured. Every column is one fact about that thing.
In Lakeview's file, a row is one person who picked up the phone, and every column is one question they answered. In the file we'll open later, a row is a ZIP code and every column is something we know about it.
Nothing mysterious about the shape. The whole question today is what each column will let you do.
2 The two kinds of data
Before you do anything with a column, you have to know which kind it is. That decides what you're allowed to do with it.
Kind 1
The value is an amount. It tells you how much or how many of something there is.
How to know: add two answers together. If the total means something, it's numeric. 25 years plus 60 years is 85 years.
Ratio zero means none
Zero means there is none of it, so "twice as much" makes sense. 10 days is twice 5 days.
Age · poor mental health days · people in a household · percent
Interval zero was chosen
The steps are equal, but zero is just a point somebody picked. 0°F is not "no temperature," so 80°F is not twice as hot as 40°F.
Temperature in °F or °C · calendar year
Kind 2
The value is a label. It puts the row into a group, even when the label is written with digits.
How to know: add two answers together. If the total means nothing, it's categorical. "Good" plus "fair" isn't anything.
Nominal no order
Just names for groups. No group is higher or lower than another.
Borough · ZIP code · insurance type · yes / no
Ordinal has an order
The groups have a built-in order, but nobody measured the distance between them. "Good" to "fair" isn't a set amount.
Excellent / good / fair / poor · mild / moderate / severe
Can you add two answers and get something that means something?
Yes → numeric. No → categorical.
If it's categorical: do the groups have a natural order?
No → nominal. Yes → ordinal.
If it's numeric: does zero mean "none of it"?
Yes → ratio. No → interval.
| Type | Count it / give a percent | Put it in order | Average it |
|---|---|---|---|
| Nominal | Yes | No | No |
| Ordinal | Yes | Yes | No — the one that gets averaged by mistake |
| Interval | Yes | Yes | Yes |
| Ratio | Yes | Yes | Yes |
Example · why ordinal can't be averaged
Four people answer "How is your health?" — excellent, excellent, poor, poor. The survey codes them excellent = 1, good = 2, fair = 3, poor = 4.
(1 + 1 + 4 + 4) ÷ 4 = 2.5
2.5 sits between "good" and "fair," so the report says this group is about average. But nobody said good or fair. Half the group is doing great and half is doing badly. The 1–4 are labels, and nobody measured whether "good" to "fair" is the same distance as "fair" to "poor."
Report it instead: 50% rated their health excellent, 50% rated it poor.
Digits show up in both kinds. A ZIP code is five digits and a name tag. A survey answer coded 1 to 4 is a number and a checked box. Neither one is an amount.
So the question is never "does it look like a number." It's "does this number measure an amount of something."
3 Back to Lakeview
Same table you just read in the opener. One at a time — verdict and reason before we open it.
You can add ages together. 25 years plus 60 years is 85 years — the addition means something, so the average does too.
Numeric · ratio
Same. Two households of 2 and 4 make 6 people. Real addition, real average.
Numeric · ratio
Q14 gave people four boxes: excellent, good, fair, poor. Somebody turned those into 1, 2, 3, 4 and averaged them. But "good" isn't 2 of anything — the numbers were assigned after the fact, so 2.7 is measuring the coding, not anyone's health. Recode the same boxes as 1, 2, 4, 8 and the average changes while every single answer stays identical.
Report it this way instead: "31% of adults rated their general health as fair or poor." Or give all four: 18% excellent, 51% good, 24% fair, 7% poor.
Categorical · ordinal — ordered boxes, no fixed distance between them
3 days plus 5 days is 8 days. A count of days is a real quantity.
Numeric · ratio
Q24 gave people four boxes: not at all, a little, somewhat, a lot. Somebody turned those into 1, 2, 3, 4 and averaged them. But "a little" isn't 2 of anything — the numbers were assigned after the fact, so 2.3 is measuring the coding, not how much anyone's life was disrupted.
Report it this way instead: "Among adults reporting poor mental health days, 38% said it interfered with their usual activities somewhat or a lot." Or give all four: 22% not at all, 40% a little, 26% somewhat, 12% a lot.
Categorical · ordinal — ordered boxes, no fixed distance between them
Adding two ZIP codes gives you nothing. The digits name a place; they don't count anything. And 60426.25 isn't a place.
Publish instead: nothing. Delete the line — there's no corrected version of this one.
Categorical · nominal
The county counted how many people checked "employed full time" — 883 out of 1,840. That's 48%. Counting is the right way to summarize a checkbox question, and that's what they did.
Categorical · nominal
The county counted how many people said yes — 1,306 out of 1,840, or 71%. Correct.
Categorical · nominal
All three treated a checked box as if it were a measured amount. Two of them had the correct version sitting four rows down in the same table.
4 More examples
Run the three questions on each one before you look at which box it's in.
Say the zero test out loud as "zero means none of it." Zero days — none of them, true, so ratio. Year zero — no year? Doesn't make sense, so interval.
Nominal labels, no order
Ordinal labels with an order
Ratio numbers, zero means none
Interval numbers, zero was chosen
Short list on purpose. Almost everything in public health data is ratio.
Watch the two duration/position pairs. Years lived in this city is ratio — zero years means none. The year you moved here is interval — 2019 is a spot on the calendar, not an amount.
5 Your portfolio
Your portfolio lives on GitHub this semester. Today you set up the homepage, and each week you'll add what you made.
Copy the prompt below and paste it into Copilot. Copilot will ask you questions one at a time. Answer in your own words. When you're done, it gives you one file, index.html.
You are helping me, a public health student, build my personal portfolio homepage for GitHub Pages. I don't know how to code, so you'll do all the code. HOW TO WORK WITH ME - Interview me first. Ask ONE question at a time and wait for my answer before asking the next one. - Keep your questions short and friendly. If my answer is very short, you may ask one follow-up to get a little more. - Do not write any code until I've answered every question. - Before you write the code, show me a short summary of my answers and ask "Anything to change?" Only write the code after I say it looks good. ASK ME THESE QUESTIONS, IN THIS ORDER 1. What name do you want shown on your page? 2. What's your major, and when do you expect to graduate? 3. In one sentence, how would you describe yourself? (Example: "Public health student focused on community mental health.") 4. Why did you choose public health? 5. What health issue do you care most about, and why? 6. What career or job are you working toward? 7. What are three skills or strengths you have? 8. What email do you want shown? (School email is fine.) 9. Do you have a LinkedIn link you want to include? (Optional.) 10. Any colors you like? (Optional.) WHEN YOU WRITE THE CODE - Give me ONE complete file called index.html, in a single code block, from <!DOCTYPE html> to </html>. Put all the styling inside the same file. No other files, nothing to install. Do not skip or shorten anything. - Clean, professional, easy to read, and it must look good on a phone and a laptop. - Use ONLY what I told you. Do not invent awards, jobs, certifications, numbers, or experience. - Never include a phone number, home address, or student ID. - After the code block, stop. Do not add instructions or next steps. THE PAGE MUST HAVE 1. A hero section at the top with my name, my one-sentence description, and a photo spot. - Photo spot: show a circle with my initials. Right above it, put a clear HTML comment that says exactly which line to change to show my photo instead, using the file name photo.jpg. 2. An "About me" section using my answers about why I chose public health, the issue I care about, and my career goal. 3. A "Skills" section. 4. A section titled "PH 320 Portfolio — Applied Biostatistics in Public Health, York College CUNY" with cards for Week 3, Week 4, and Week 5. Each card says "Coming soon." Put an HTML comment above each card showing me where to add a link later. 5. A contact section with only my email and LinkedIn (if I gave one). Start now by asking me question 1.
Read your page before you upload it. If Copilot added anything you didn't say, delete it. Never put your phone number, home address, or student ID on it.
6 The tool
Python is a language for giving a computer instructions. Built in 1991, named after Monty Python — the comedy group, not the snake. The name is a joke about not being intimidating.
A notebook is a document where you write those instructions in small chunks, and the answer appears directly underneath each chunk. Like a lab notebook: do a step, record what happened, next step. Not a wall of code — a conversation with the data.
A spreadsheet shows you the data. A notebook shows you what you did to the data.
The steps stay visible. Somebody else can open it and see exactly how you got your number — including how you got a 2.7.
You are not going to write Python in this course. I write the steps; you read them and decide whether the result makes sense.
It's like a recipe. You're learning to read the recipe and judge the dish — not to cook.
And the catch that makes this a statistics course: in math, you show your steps and you produce the answer. In a notebook the computer produces the answer — so the wrong steps give you a confidently wrong number that still looks perfectly tidy. 2.7 looked perfectly tidy.
7 On screen
You'll see this line at the top of every notebook in this course.
df = pd.read_csv("lakeview_survey.csv")
"Use pandas to open the survey file as a table, and call that table df."
| pd | Our short name for pandas — a toolbox that knows how to work with tables. Python on its own doesn't. The dot means "reach into that toolbox." |
| read_csv(…) | The actual instruction: open a CSV file — rows and columns — and turn it into a table. |
| "…csv" | Which file to open. |
| df = | Give that table a name so the next line can use it. Short for data frame. |
The distinction worth holding onto: pd is the tool, df is your table.
Neither name is magic. df is a nickname somebody chose, and I could have chosen a clearer one:
survey = pd.read_csv("lakeview_survey.csv")
Exactly the same instruction. We'd just write survey.head() afterward instead of df.head(). But df is what nearly everyone uses, so you'll see it in every notebook you ever open — including ours.
Why give it a name at all: without one, Python reads the file, shows it once, and lets it go. Like pulling a book off the shelf, reading it aloud, and putting it straight back. The name keeps it on the desk so the next line can pick it up.
headhead shows you the top of the file — the first five rows. Head as in the head of a list. Nobody wants 168 rows on a projector.
| A | B | C | D | |
| 1 | zip | borough | adults | distress_pct |
| 2 | 10001 | Manhattan | 29742 | 14.7 |
| 3 | 10002 | Manhattan | 70736 | 15.9 |
| 4 | 10003 | Manhattan | 53076 | 15.7 |
| 5 | 10005 | Manhattan | 9477 | 15.0 |
| 6 | 10009 | Manhattan | 55073 | 15.7 |
| 7 | 10010 | Manhattan | 30139 | 14.8 |
| 8 | 10011 | Manhattan | 48298 | 13.7 |
| 9 | ⋮ | ⋮ | ⋮ | ⋮ |
df = pd.read_csv("nyc_mental_health_by_zip.csv")df.head()| zip | borough | adults | distress_pct | |
| 0 | 10001 | Manhattan | 29742 | 14.7 |
| 1 | 10002 | Manhattan | 70736 | 15.9 |
| 2 | 10003 | Manhattan | 53076 | 15.7 |
| 3 | 10005 | Manhattan | 9477 | 15.0 |
| 4 | 10009 | Manhattan | 55073 | 15.7 |
Same file, same five rows. The only difference is that on the right you can see what I did to get there.
8 Our own file
| Column | Looks like | Type | Why |
|---|---|---|---|
| zip | 10001 | Nominal | Digits, but a name. Adding two of them gives you nothing. |
| borough | Manhattan | Nominal | A label with no order. Queens isn't more than Brooklyn. |
| adults | 29742 | Ratio | A count of people. Zero adults means none. |
| distress_pct | 14.7 | Ratio | Zero percent means nobody. 14 percent is twice 7 percent. |
| year collected | 2022 | Interval | A position on the calendar, not an amount. Somebody chose where the counting starts. |
This file has no ordinal column — but one is hiding in it. poor_health_pct came from the same excellent / good / fair / poor question you just saw. Somebody counted the fair-and-poor answers and published a percentage, which is the correct move, done before the file ever reached us.
9 Why it matters
Lakeview's analyst wasn't careless. He ran the averages across every column at once, and the software handed him all of them.
df["distress_pct"].mean()df["zip"].mean()Both answers came back instantly, in the same font, with no warning. Only one of them means anything — and knowing which is your job, not the computer's.
10 Leaving with this
What you're allowed to do with a column depends on what kind of thing is in it — and digits on the screen don't make a column numeric.