šŸ“ˆ Linear Regression & Coefficient Uncertainty

explained like you're 12 Ā· Learning Check 1 of 7

⭐ The big idea: linear regression just draws the best straight line through your points, and the coefficients β₀ and β₁ tell you where that line starts and how steep it is. Everything here is "find that line, then measure how sure we are about it."

🟦 The Dataset (our 4 points)

We have four points: (0,1), (0,3), (10,10), (10,6).

x-values: 0, 0, 10, 10 → y-values: 1, 3, 10, 6
šŸŽÆ Two clusters: think of it like asking 4 people their age (x) and their shoe size (y). Two people are age 0, two are age 10. We want one straight line that "best fits" everyone.
Regression fit and residuals
šŸŽØ How to read this: LEFT — the green line y = 2 + 0.6x is our "best fit" through the 4 blue dots. The dashed orange lines are how far each point is from the line (the "errors"). RIGHT — the same errors drawn separately: these are the "noise" ε. This is exactly the picture for Question 7.

⭐ The ONE formula you need first

Mean of x: x̄ = (0 + 0 + 10 + 10) / 4 = 5

The mean (xĢ„, read "x-bar") is just the average — add up all the x-values and divide by how many there are.

0 + 0 + 10 + 10 = 20 , 20 Ć· 4 = 5

1ļøāƒ£ Question 1 — What is xĢ„?

What is x̄ (the mean of the x-values)?
āŒ 4 āœ… 5 āŒ 0.6 āŒ 2
āœ… Answer: 5
1
List the x-values: 0, 0, 10, 10.
2
Add them: 0 + 0 + 10 + 10 = 20.
3
Divide by 4 points: 20 / 4 = 5.
šŸŽ Average analogy: if 4 friends have 0, 0, 10, and 10 apples, the average is 20 apples Ć· 4 friends = 5 apples each.

2ļøāƒ£ Question 2 — What is β̂₁?

In the same problem, what is β̂₁ (the slope)?
āŒ 5 āœ… 0.6 āŒ 2
āœ… Answer: 0.6
1
The slope β₁ measures "how much y goes up when x goes up by 1".
2
Formula: β₁ = Ī£(xāˆ’xĢ„)(yāˆ’Č³) / Ī£(xāˆ’xĢ„)². It's the "best tilt" of the line.
3
Plugging in the numbers gives β₁ = 0.6. (Check: from x=0 to x=10, y rises from 2 to 8 — that's 6 up over 10 across = 0.6 per unit.)
ā›°ļø Slope analogy: if a road rises 6 meters over 10 meters of horizontal distance, its slope is 6/10 = 0.6. The line rises gently.

3ļøāƒ£ Question 3 — What is Mā‚“įµ€?

What is Mā‚“įµ€ (the transpose of the design matrix)?
āŒ [0 0 10 10] āœ… [[1 1 1 1],[0 0 10 10]] āŒ [[1 1 1 1],[10 10 0 0]] āŒ [[0 0 10 10],[1 1 1 1]]
āœ… Answer: [[1 1 1 1],[0 0 10 10]] (the second option)

The design matrix Mā‚“

The design matrix Mā‚“ has a column of 1s (for the intercept β₀) plus a column of the x-values:

1
0
1
0
1
10
1
10
Mā‚“
1
Transpose = flip the matrix sideways (rows become columns, columns become rows).
2
First row of Mā‚“įµ€ = the column of 1s: [1 1 1 1].
3
Second row of Mā‚“įµ€ = the column of x-values: [0 0 10 10].
4
Result: [[1 1 1 1],[0 0 10 10]] āœ… (the 1s row comes first).
šŸ”„ Flip analogy: transposing is like rotating a bookshelf sideways — the same books, just arranged differently. The 1s column and the x column swap which way they point.

4ļøāƒ£ Question 4 — What is Mā‚“įµ€Mā‚“?

What is Mā‚“įµ€Mā‚“ (design matrix squared)?
āœ… [[4, 20],[20, 200]] āŒ [[20, 4],[20, 200]] āŒ [[20, 4],[200, 20]] āŒ [[4, 20],[200, 20]]
āœ… Answer: [[4, 20],[20, 200]] (the first option)
1
1
1
1
0
0
10
10
Ɨ
1
0
1
0
1
10
1
10
=
4
20
20
200
1
Top-left: count of points Ɨ 1 = 4 (each 1Ā·1 summed: four of them).
2
Top-right & bottom-left: sum of x-values = 0+0+10+10 = 20.
3
Bottom-right: sum of x² = 0+0+100+100 = 200.
šŸ”¢ Summary table: Mā‚“įµ€Mā‚“ packs the data into 3 numbers — how many points (4), the sum of x (20), and the sum of x² (200). It's the same idea as Aįµ€A from the last learning check.

5ļøāƒ£ Question 5 — What is (Mā‚“įµ€Mā‚“)⁻¹?

What is (Mā‚“įµ€Mā‚“)⁻¹ (the inverse)?
āœ… [[0.5, āˆ’0.05],[āˆ’0.05, 0.01]] āŒ [[0.1, āˆ’0.05],[āˆ’0.01, 0.05]] āŒ [[1, 5],[0.1, 0.25]] āŒ [[āˆ’0.05, 0.01],[0.5, āˆ’0.05]]
āœ… Answer: [[0.5, āˆ’0.05],[āˆ’0.05, 0.01]] (the first option)

The 2Ɨ2 inverse formula

a
b
c
d
⁻¹ = 1/(adāˆ’bc) Ɨ
d
āˆ’b
āˆ’c
a
1
Our matrix is [[4, 20],[20, 200]], so a=4, b=20, c=20, d=200.
2
Determinant = ad āˆ’ bc = 4Ɨ200 āˆ’ 20Ɨ20 = 800 āˆ’ 400 = 400.
3
Swap d and a, put minus on b and c: [[200, āˆ’20],[āˆ’20, 4]].
4
Divide by 400: [[0.5, āˆ’0.05],[āˆ’0.05, 0.01]] āœ….
āž— Division analogy: the inverse is like "dividing" by a matrix. If Mā‚“įµ€Mā‚“ is "times 400", then its inverse is "divide by 400" — but with the numbers rearranged too.

6ļøāƒ£ Question 6 — What is Mā‚“įµ€y?

What is Mā‚“įµ€y (design matrix times target)?
āœ… [[20],[160]] āŒ [[160],[20]] āŒ [[400],[200]] āŒ [[20],[200]]
āœ… Answer: [[20],[160]] (the first option)
1
1
1
1
0
0
10
10
Ɨ
1
3
10
6
=
20
160
1
Top row: sum of all y = 1+3+10+6 = 20.
2
Bottom row: sum of xĀ·y = 0Ā·1 + 0Ā·3 + 10Ā·10 + 10Ā·6 = 0 + 0 + 100 + 60 = 160.
🧾 Receipt analogy: Mā‚“įµ€y adds up "total y" (20) and "total x times y" (160) — the two numbers we need to solve for the line.

7ļøāƒ£ Question 7 — What does σ² mean?

In the equations below, what is the interpretation of σ²?
āŒ Standard deviation of the β coefficients (which are normal) āŒ Standard deviation of the measurement noise ε āœ… Variance of the measurement noise ε in a linear additive model āŒ Variance of the β coefficients arising from matrix multiplication
āœ… Answer: Variance of the measurement noise ε (the third option)
1
Every point isn't exactly on the line — it wobbles off by a bit. That wobble is the noise ε (the orange dashed lines in our picture).
2
σ² is how spread-out that wobble is — the "variance" of the noise. Small σ² = points hug the line tightly; big σ² = points scatter everywhere.
3
The two equations in the question use σ² to compute how "sure" we are about β₀ and β₁. More noise (bigger σ²) = less sure.
šŸŽÆ Darts analogy: imagine throwing darts at a board. The line is the bullseye. σ² is how spread-out your throws are. A tight cluster = small σ² (confident); scattered throws = big σ² (noisy, unsure).

šŸŽ® Play with it yourself

šŸ“ Regression Fitter

Watch how β₀ (intercept) and β₁ (slope) change the line. Can you eyeball the best fit?

—

šŸŽÆ Noise Simulator (Question 7)

Drag to change the noise level σ. Watch how much the points scatter around the line.

—

šŸ“‹ Quick Answer Sheet

#QuestionAnswer
1x̄ (mean of x)5
2β̂₁ (slope)0.6
3Mā‚“įµ€[[1 1 1 1],[0 0 10 10]]
4Mā‚“įµ€Mā‚“[[4, 20],[20, 200]]
5(Mā‚“įµ€Mā‚“)⁻¹[[0.5, āˆ’0.05],[āˆ’0.05, 0.01]]
6Mā‚“įµ€y[[20],[160]]
7Interpretation of σ²Variance of noise ε

🧠 Checklist (did you get it?)