More predictors than observations
You have 100 months of returns and 1000 candidate predictors. OLS solves the normal equations , and is a matrix of rank at most 100. It has no inverse. Worse, with more unknowns than equations there are infinitely many exact solutions: you can fit every observation perfectly, , on pure noise.
The trouble starts well before . Suppose the true signal is zero, the noise has variance , and you fit OLS with parameters on points. The fitted values are , where the hat matrix has trace . Then
With and , the training error is and the error on new data at the same inputs is . In sample the model looks twice as good as the truth; out of sample it is 50% worse than predicting zero. Each extra parameter bends the fit toward the particular noise it happened to see, and that noise doesn't repeat.
The rest of this lesson accepts a little bias in exchange for much less variance, so that a fit holds up on data it hasn't seen.