Statistics & Regression

Lesson 3 of 9

Least Squares as a Projection

The linear model y = Xβ + ε, the normal equations as a perpendicular drop onto the column space, R² as Pythagoras, and why the simple-regression slope is ρ times sd(y)/sd(x).

The linear model

You observe nn pairs (xi,yi)(x_i, y_i) and model them as yi=β0+β1xi+εiy_i = \beta_0 + \beta_1 x_i + \varepsilon_i, where the εi\varepsilon_i are noise. Stack the rows and it becomes

y=Xβ+εy = X\beta + \varepsilon

where yy is an nn-vector, XX is the n×pn \times p design matrix (first column all ones for the intercept, one column per predictor) and β\beta holds the pp coefficients. Least squares picks the β^\hat\beta that minimizes the sum of squared residuals ∥y−Xβ∥2=∑i(yi−xiTβ)2\|y - X\beta\|^2 = \sum_i (y_i - x_i^T \beta)^2, where xiTx_i^T is row ii of XX. Squares make the problem smooth with a closed-form answer, and when the noise is normal the same β^\hat\beta comes out of maximum likelihood, as the Maximum Likelihood lesson shows.

Take five points: x=(1,2,3,4,5)x = (1,2,3,4,5), y=(2,4,5,4,5)y = (2,4,5,4,5). The design matrix has rows (1,xi)(1, x_i). The useful move is to stop picturing five points in the plane and treat yy as a single vector in R5\mathbb{R}^5. Every choice of β\beta produces a vector XβX\beta, and the set of all of them is the column space of XX, a 2-dimensional plane sitting inside 5-dimensional space. No line passes through all five points, so yy lies off that plane, and finding the best line becomes a geometry question: which point on the plane is closest to yy?