Problem bank

Problem 236 of 333MediumStatisticsP236

Training error is optimistic

  1. Let y=Xβ+εy = X\beta + \varepsilon with XX a fixed n×pn \times p matrix of full column rank and ε\varepsilon having mean zero and covariance σ2I\sigma^2 I. Fit OLS, β^=(XTX)−1XTy\hat\beta = (X^TX)^{-1}X^Ty.

    (a) Show E[∥y−Xβ^∥2]=σ2(n−p)E\left[\|y - X\hat\beta\|^2\right] = \sigma^2(n - p).

    (b) Let y∗=Xβ+ε∗y^* = X\beta + \varepsilon^* be fresh data at the same inputs, with ε∗\varepsilon^* independent of ε\varepsilon and distributed the same way. Show E[∥y∗−Xβ^∥2]=σ2(n+p)E\left[\|y^* - X\hat\beta\|^2\right] = \sigma^2(n + p).

    (c) With n=250n = 250 daily observations and p=25p = 25 predictors, by what factor does average training MSE understate average test MSE?