๐Ÿ“Š Exploratory Data Analysis

explained like you're 12 ยท College dataset (777 US universities)

โญ The big idea: this assignment is all about getting to know a dataset before doing any math on it โ€” peek at rows, check for junk (missing values, duplicates), slice out pieces you care about, and draw pictures. EDA = the "look before you leap" of data science. Here the data is clean (no missing, no duplicates), so the fun is in the slicing and the plotting.

๐ŸŸฆ The Setup

The College.csv dataset (from the ISLR textbook) holds 777 U.S. universities with 18 variables: applications received (Apps), accepted (Accept), enrolled (Enroll), tuition (Outstate), whether it's Private, graduation rate, and more. The first column holds each university's name, so we make that the index โ€” every row is now labelled with its school.

๐Ÿ—‚๏ธ Index = name tag: instead of calling a row "row 42", we call it "Princeton University". Python won't treat the name as a number, which is exactly what we want.

1๏ธโƒฃ Question 1 โ€” Peek at the data

Use head() to view the first few rows.
โœ… Answer: college.head() โ†’ shows the first 5 rows (Abilene Christian, Adelphi, Adrian, Agnes Scott, Alaska Pacific).
1
After loading, the very first column is the university name (labelled Unnamed: 0 in the raw CSV).
2
We rename it to Names, then set it as the index with college.set_index('Names', inplace=True).
3
Now college.head() prints each school's name at the left edge instead of a row number.

2๏ธโƒฃ Question 2 โ€” The info() summary

Use info() to produce a numerical summary of the variables.
โœ… Answer: college.info() โ†’ 777 rows, all no missing values, 1 categorical (Private: Yes/No) + 17 numeric columns.
1
info() prints each column, its data type, and how many non-null (non-missing) entries it has.
2
Every column shows 777 non-null โ†’ the dataset is complete. No missing values to worry about.
3
Only Private is object (text). Everything else is a number, so most columns are ready for math.
๐Ÿ” Health check: info() is the doctor's checkup โ€” it tells you if any variable is "sick" (missing data) before you start analysing.

3๏ธโƒฃ Question 3 โ€” Check for duplicates

Examine for duplicates and drop them in place.
โœ… Answer: college.drop_duplicates(inplace=True) โ†’ 0 duplicates found, so nothing changes. (Still good practice to check!)
1
college.duplicated().sum() counts how many duplicate rows exist.
2
Here the count is 0 โ€” every row is unique.
3
The inplace=True argument tells pandas to modify college directly instead of returning a copy.
๐Ÿงน Tidiness: duplicates often come from merging messy data or typing errors. Always sweep for them โ€” even if the floor is already clean.

4๏ธโƒฃ Question 4 โ€” Replace missing Apps with 0

Replace any missing values in the Apps column with 0.
โœ… Answer: college['Apps'] = college['Apps'].fillna(0) โ†’ there were already no missing values, so nothing changes (and the test confirms no row has 0 apps).
1
fillna(0) replaces every NaN (missing) in Apps with 0.
2
Since Apps had no NaN, the column is unchanged.
3
The assertion college[college['Apps']==0].empty checks that no real school has 0 applications โ€” every school got some applications.
โ˜‘๏ธ Just in case: even clean data gets a safety check. Filling missing numbers with 0 is a simple default so downstream math doesn't crash.

5๏ธโƒฃ Question 5 โ€” Least out-of-state tuition

Find the college with the least out-of-state tuition (return the name).
โœ… Answer: college_least_tuition = college['Outstate'].idxmin() โ†’ "Brigham Young University at Provo".
1
idxmin() returns the index label (the school name) of the row with the smallest value in that column.
2
That's why we set the index to university names โ€” now idxmin() gives us the name directly, not a row number.
๐Ÿท๏ธ Name, not number: idxmin() is like asking "who has the cheapest tuition?" and getting "Brigham Young" back instead of "row 87".

6๏ธโƒฃ Question 6 โ€” Select the PhD column

Select the PhD column as a dataframe; find its length.
โœ… Answer: phd_column = college[['PhD']] โ†’ phd_column_length = len(phd_column) = 777.
1
Use double brackets [['PhD']] to get a DataFrame (neat table, keeps row names, 2-D).
2
Use single brackets ['PhD'] to get a Series (a 1-D column, like a numpy array).
3
Since there's one PhD value per university, its length equals the number of rows = 777.
๐Ÿ“‘ Table vs list: a DataFrame is a spreadsheet; a Series is one column of it. Double brackets = "give me the whole table with just this column".

7๏ธโƒฃ Question 7 โ€” Slice rows 15 & 16

Select Private and Top10perc, keeping only rows 15 and 16 (0-indexed).
โœ… Answer: private_top10 = college[['Private','Top10perc']].iloc[15:17] โ†’ rows are American International College (Top10perc 9) and Amherst College (Top10perc 83). Length = 2.
1
First pick the two columns with double brackets.
2
.iloc[15:17] slices by position โ€” rows 15 and 16 (17 is exclusive), giving exactly 2 rows.
3
So len(private_top10[['Private']]) = 2.
๐Ÿ”ข Position vs label: iloc counts rows like an index card (0,1,2โ€ฆ); loc looks them up by name. Both are useful โ€” know which one you're using.

8๏ธโƒฃ Question 8 โ€” All schools with "Penn"

Select all rows whose index contains "Penn" (capital P).
โœ… Answer: many_penns = college[college.index.str.contains("Penn")] โ†’ 8 schools.
1
.index.str.contains("Penn") checks each school name for the substring "Penn".
2
It's case-sensitive, so only "Penn" (capital P) matches โ€” no lowercase "penn".
3
Result: 8 schools including University of Pennsylvania, Penn State, and several "of Pennsylvania" schools.
๐Ÿ”Ž Search: this is a text filter โ€” like Ctrl+F for "Penn" across all the name tags, grabbing every row that has it.

9๏ธโƒฃ Question 9 โ€” Pairplot of the first 10 columns

Use seaborn.pairplot() on the first 10 columns.
โœ… Answer: sns.pairplot(college[college.columns[:10]]) โ†’ a scatterplot matrix of every pair.
pairplot
๐ŸŽจ How to read this: each square is one pair of variables. The diagonal shows each variable alone (histogram). The key takeaway โ€” Apps, Accept, Enroll, and F_Undergrad all grow together: big schools get more applications, accept more, and enroll more.
๐Ÿ‘€ Eyes on everything: a pairplot shows every two-variable relationship at once. The clear diagonal "clusters" between size variables tell you they're strongly correlated.

๐Ÿ”Ÿ Question 10 โ€” Tuition: private vs public

Boxplot of Outstate vs Private. Find average out-of-state tuition for private and for public.
โœ… Answer: private โ‰ˆ $11,802 ยท public โ‰ˆ $6,813.
boxplot
๐ŸŽจ How to read this: the blue box = private schools, amber box = public. Private schools charge far more out-of-state (median well above $10,000), and they spread higher too. The dashed lines mark each average.
1
Compute the means by filtering on the Private column: college.loc[college['Private']=='Yes','Outstate'].mean().
2
Private โ‰ˆ $11,802 โ€” about 73% more than public ($6,813).
๐Ÿ’ธ Price gap: private universities charge out-of-state students roughly double what public schools do โ€” the classic sticker-price difference everyone knows.

1๏ธโƒฃ1๏ธโƒฃ Question 11 โ€” Build the "large university" mask

Create large_university: True where Enroll exceeds the mean of all enrollments.
โœ… Answer: large_university = college['Enroll'] > college['Enroll'].mean() (mean โ‰ˆ 780).
1
This creates a boolean Series (a "mask") โ€” one True/False per university.
2
A school is "large" if it enrolls more than the average of ~780 new students.
๐ŸŽš๏ธ The divider: a mask is a row of switches โ€” True keeps a row, False drops it. It's the tool we use to slice data by a condition.

1๏ธโƒฃ2๏ธโƒฃ Question 12 โ€” The large_universities dataframe

Create a dataframe of only the large universities.
โœ… Answer: large_universities = college[large_university] โ†’ 218 rows (out of 777, so ~28%).
1
Passing the mask inside college[...] keeps only the rows where the mask is True.
2
218 universities have enroll > mean (~780).
๐ŸŽฃ Filtering: the mask is like a net that catches only "big" fish. Out of 777 schools, 218 are big by this rule.

1๏ธโƒฃ3๏ธโƒฃ Question 13 โ€” 75th percentile of Enroll

Use describe(include='all') on large_universities; find the 75th percentile of Enroll.
โœ… Answer: enroll = large_universities.describe(include='all').loc['75%','Enroll'] = 2,408.
1
describe(include='all') gives a full summary (count, mean, min, quartiles, max) for every column.
2
Grab the 75% row and the Enroll column โ†’ 2,408.
3
Meaning: the top quarter of "large" universities enroll more than 2,408 new students.
๐Ÿ“ถ Upper rung: even among the "big" schools, the biggest 25% are enormous โ€” reinforcing that enrollment is right-skewed (a few giant flagships dominate).

1๏ธโƒฃ4๏ธโƒฃ Question 14 โ€” Applications histogram

Histogram of Apps with bin size 200. Is it true more universities get 0โ€“10,000 apps than 10,000โ€“20,000?
โœ… Answer: applications = True โ€” 732 schools get 0โ€“10,000 apps, only 42 get 10,000โ€“20,000.
histogram
๐ŸŽจ How to read this: nearly all the bars are crammed against the left edge (the green 0โ€“10,000 band). Only a tiny handful reach the amber 10,000โ€“20,000 band. Distribution is heavily right-skewed.
1
Count apps in 0โ€“10,000: ((college['Apps']>=0)&(college['Apps']<10000)).sum() = 732.
2
Count apps in 10,000โ€“20,000: = 42.
3
732 > 42, so applications = True.
๐Ÿ“‰ Long tail: most U.S. colleges are small or regional and get modest applications. Only elite/large flagships pass 10,000. Application volume is a long-tail distribution, not a bell curve.

1๏ธโƒฃ5๏ธโƒฃ Question 15 โ€” Acceptance rate

Add acceptance_rate = Accept / Apps to a copy. Which college has the lowest rate? Highest?
โœ… Answer: lowest = college_copy['acceptance_rate'].idxmin() โ†’ "Princeton University" (~15.5%). Highest = idxmax() โ†’ "Emporia State University" (~100%).
acceptance rate
๐ŸŽจ How to read this: red bars = most selective (Princeton ~0.15, Harvard, Yale, Columbia). Green bars = most open (Emporia State, MidAmerica, Wayne State, Southwest Baptist all accept essentially everyone).
1
Always copy first: college_copy = college.copy() so we don't touch the original dataframe.
2
Add the derived column: college_copy['acceptance_rate'] = college_copy['Accept'] / college_copy['Apps'].
3
idxmin()/idxmax() return the school name with the lowest/highest rate.
๐ŸŽฏ Selectivity: acceptance rate = how picky a school is. Princeton is hyper-competitive (only ~15% in); Emporia State is open-admission (accepts basically everyone). The range spans 15% โ†’ 100%.

๐ŸŽฎ Play with it yourself

โš–๏ธ Tuition Gap Calculator (Question 10)

Drag the private average and public average. Watch how big the gap is and the % difference.

โ€”

๐ŸŽ“ Large-University Splitter (Question 11โ€“13)

Drag the enrollment threshold. Watch how many schools count as "large" and where the 75th percentile lands. (The real assignment uses the mean โ‰ˆ 780 โ†’ 218 schools, 75th percentile = 2,408.)

โ€”

๐ŸŽฏ Acceptance-Rate Explorer (Question 15)

Drag the number of applications and accepted. Watch the acceptance rate and how selective it feels.

โ€”

๐Ÿ“‹ Quick Answer Sheet

#QuestionAnswer
1head()First 5 rows (5 schools)
2info()777 rows, no missing, 1 categorical + 17 numeric
3Duplicates0 found โ€” drop changes nothing
4Missing Apps โ†’ 0No missing already; fillna(0) no-op
5Least out-of-state tuitionBrigham Young University at Provo
6Length of PhD column777
7Rows 15โ€“16 (Private, Top10perc)Length = 2 (Am. International, Amherst)
8Colleges with "Penn"8 schools
9Pairplot first 10 colsSize variables all correlate
10Avg outstate private / public$11,802 / $6,813
11large_university maskEnroll > mean (~780)
12large_universities218 schools
1375th pct of Enroll2,408
14More 0โ€“10k than 10kโ€“20k?True (732 vs 42)
15Lowest / highest acceptancePrinceton / Emporia State

๐Ÿง  Checklist (did you get it?)