explained like you're 12 ยท College dataset (777 US universities)
The College.csv dataset (from the ISLR textbook) holds 777 U.S. universities with 18 variables: applications received (Apps), accepted (Accept), enrolled (Enroll), tuition (Outstate), whether it's Private, graduation rate, and more. The first column holds each university's name, so we make that the index โ every row is now labelled with its school.
head() to view the first few rows.college.head() โ shows the first 5 rows (Abilene Christian, Adelphi, Adrian, Agnes Scott, Alaska Pacific).Unnamed: 0 in the raw CSV).Names, then set it as the index with college.set_index('Names', inplace=True).college.head() prints each school's name at the left edge instead of a row number.info() summaryinfo() to produce a numerical summary of the variables.college.info() โ 777 rows, all no missing values, 1 categorical (Private: Yes/No) + 17 numeric columns.info() prints each column, its data type, and how many non-null (non-missing) entries it has.Private is object (text). Everything else is a number, so most columns are ready for math.info() is the doctor's checkup โ it tells you if any variable is "sick" (missing data) before you start analysing.college.drop_duplicates(inplace=True) โ 0 duplicates found, so nothing changes. (Still good practice to check!)college.duplicated().sum() counts how many duplicate rows exist.inplace=True argument tells pandas to modify college directly instead of returning a copy.Apps with 0Apps column with 0.college['Apps'] = college['Apps'].fillna(0) โ there were already no missing values, so nothing changes (and the test confirms no row has 0 apps).fillna(0) replaces every NaN (missing) in Apps with 0.Apps had no NaN, the column is unchanged.college[college['Apps']==0].empty checks that no real school has 0 applications โ every school got some applications.college_least_tuition = college['Outstate'].idxmin() โ "Brigham Young University at Provo".idxmin() returns the index label (the school name) of the row with the smallest value in that column.idxmin() gives us the name directly, not a row number.idxmin() is like asking "who has the cheapest tuition?" and getting "Brigham Young" back instead of "row 87".PhD columnPhD column as a dataframe; find its length.phd_column = college[['PhD']] โ phd_column_length = len(phd_column) = 777.[['PhD']] to get a DataFrame (neat table, keeps row names, 2-D).['PhD'] to get a Series (a 1-D column, like a numpy array).PhD value per university, its length equals the number of rows = 777.Private and Top10perc, keeping only rows 15 and 16 (0-indexed).private_top10 = college[['Private','Top10perc']].iloc[15:17] โ rows are American International College (Top10perc 9) and Amherst College (Top10perc 83). Length = 2..iloc[15:17] slices by position โ rows 15 and 16 (17 is exclusive), giving exactly 2 rows.len(private_top10[['Private']]) = 2.iloc counts rows like an index card (0,1,2โฆ); loc looks them up by name. Both are useful โ know which one you're using.many_penns = college[college.index.str.contains("Penn")] โ 8 schools..index.str.contains("Penn") checks each school name for the substring "Penn".seaborn.pairplot() on the first 10 columns.sns.pairplot(college[college.columns[:10]]) โ a scatterplot matrix of every pair.Outstate vs Private. Find average out-of-state tuition for private and for public.private โ $11,802 ยท public โ $6,813.Private column: college.loc[college['Private']=='Yes','Outstate'].mean().large_university: True where Enroll exceeds the mean of all enrollments.large_university = college['Enroll'] > college['Enroll'].mean() (mean โ 780).large_universities dataframelarge_universities = college[large_university] โ 218 rows (out of 777, so ~28%).college[...] keeps only the rows where the mask is True.Enrolldescribe(include='all') on large_universities; find the 75th percentile of Enroll.enroll = large_universities.describe(include='all').loc['75%','Enroll'] = 2,408.describe(include='all') gives a full summary (count, mean, min, quartiles, max) for every column.75% row and the Enroll column โ 2,408.Apps with bin size 200. Is it true more universities get 0โ10,000 apps than 10,000โ20,000?applications = True โ 732 schools get 0โ10,000 apps, only 42 get 10,000โ20,000.((college['Apps']>=0)&(college['Apps']<10000)).sum() = 732.applications = True.acceptance_rate = Accept / Apps to a copy. Which college has the lowest rate? Highest?college_copy['acceptance_rate'].idxmin() โ "Princeton University" (~15.5%). Highest = idxmax() โ "Emporia State University" (~100%).college_copy = college.copy() so we don't touch the original dataframe.college_copy['acceptance_rate'] = college_copy['Accept'] / college_copy['Apps'].idxmin()/idxmax() return the school name with the lowest/highest rate.| # | Question | Answer |
|---|---|---|
| 1 | head() | First 5 rows (5 schools) |
| 2 | info() | 777 rows, no missing, 1 categorical + 17 numeric |
| 3 | Duplicates | 0 found โ drop changes nothing |
| 4 | Missing Apps โ 0 | No missing already; fillna(0) no-op |
| 5 | Least out-of-state tuition | Brigham Young University at Provo |
| 6 | Length of PhD column | 777 |
| 7 | Rows 15โ16 (Private, Top10perc) | Length = 2 (Am. International, Amherst) |
| 8 | Colleges with "Penn" | 8 schools |
| 9 | Pairplot first 10 cols | Size variables all correlate |
| 10 | Avg outstate private / public | $11,802 / $6,813 |
| 11 | large_university mask | Enroll > mean (~780) |
| 12 | large_universities | 218 schools |
| 13 | 75th pct of Enroll | 2,408 |
| 14 | More 0โ10k than 10kโ20k? | True (732 vs 42) |
| 15 | Lowest / highest acceptance | Princeton / Emporia State |
head() peeks; info() checks health (missing values, types).[['col']] โ DataFrame; single ['col'] โ Series.iloc slices by position; loc looks up by index label.idxmin()/idxmax() return the name of the min/max row.df['col'] > value filters rows.