Statistics & Regression

Lesson 8 of 9

Bias, Variance and Regularization

Why extra predictors make out-of-sample fit worse, how ridge and lasso trade a little bias for less variance, and how cross-validation picks the penalty.

More predictors than observations

You have 100 months of returns and 1000 candidate predictors. OLS solves the normal equations XTXβ^=XTyX^TX\hat\beta = X^Ty, and XTXX^TX is a 1000×10001000 \times 1000 matrix of rank at most 100. It has no inverse. Worse, with more unknowns than equations there are infinitely many exact solutions: you can fit every observation perfectly, R2=1R^2 = 1, on pure noise.

The trouble starts well before p>np > n. Suppose the true signal is zero, the noise has variance σ2\sigma^2, and you fit OLS with pp parameters on nn points. The fitted values are HyHy, where the hat matrix HH has trace pp. Then

E[training MSE]=σ2(1−pn),E[MSE on fresh noise]=σ2(1+pn).E[\text{training MSE}] = \sigma^2\left(1 - \frac{p}{n}\right), \qquad E[\text{MSE on fresh noise}] = \sigma^2\left(1 + \frac{p}{n}\right).

With n=100n = 100 and p=50p = 50, the training error is 0.5σ20.5\sigma^2 and the error on new data at the same inputs is 1.5σ21.5\sigma^2. In sample the model looks twice as good as the truth; out of sample it is 50% worse than predicting zero. Each extra parameter bends the fit toward the particular noise it happened to see, and that noise doesn't repeat.

The rest of this lesson accepts a little bias in exchange for much less variance, so that a fit holds up on data it hasn't seen.