Statistics & Regression

Lesson 5 of 9

Regression Inference: Standard Errors, t and F

How precise a regression coefficient is: standard errors, t-tests and intervals for one coefficient, F-tests for several at once, and the residual checks that decide whether to trust any of them.

The slope is a random variable

A regression reports β^1=2.5\hat\beta_1 = 2.5. Is the true slope different from zero, or would another sample have given 0.3? To answer, treat the estimate as what it is: a function of noisy data. Write the model as y=Xβ+εy = X\beta + \varepsilon with E[ε∣X]=0E[\varepsilon \mid X] = 0 and Var(ε∣X)=σ2I\text{Var}(\varepsilon \mid X) = \sigma^2 I. Substituting into the OLS formula from Least Squares as a Projection,

β^=(XTX)−1XTy=β+(XTX)−1XTε,\hat\beta = (X^TX)^{-1}X^Ty = \beta + (X^TX)^{-1}X^T\varepsilon,

so β^\hat\beta is unbiased and

Var(β^∣X)=σ2(XTX)−1.\text{Var}(\hat\beta \mid X) = \sigma^2 (X^TX)^{-1}.

The standard error of coefficient jj is the square root of the jjth diagonal entry, with σ2\sigma^2 replaced by its estimate s2=RSS/(n−p)s^2 = \text{RSS}/(n-p), where pp counts every coefficient including the intercept. Dividing by n−pn-p fixes the same bias as the n−1n-1 in the sample variance: the fit used up pp degrees of freedom.

For a simple regression the diagonal entry has a closed form:

SE(β^1)=s∑i(xi−xˉ)2≈ssxn.\text{SE}(\hat\beta_1) = \frac{s}{\sqrt{\sum_i (x_i - \bar x)^2}} \approx \frac{s}{s_x\sqrt{n}}.

That gives three levers: less residual noise, more spread in xx, more data. Because of the n\sqrt n, halving the standard error costs four times the data.

Worked example: n=27n = 27, residual standard error s=4s = 4, and ∑(xi−xˉ)2=16\sum (x_i - \bar x)^2 = 16. Then SE(β^1)=4/16=1\text{SE}(\hat\beta_1) = 4/\sqrt{16} = 1.