Statistics & Regression

Lesson 7 of 9

Logistic Regression

Why a line fails on yes/no outcomes, how logistic regression models the log-odds, reading coefficients as odds ratios, fitting by Newton's method, and judging a classifier with ROC and AUC.

Why a straight line fails

A credit desk wants the probability that a borrower defaults, given the debt-to-income ratio xx in percent. The outcome is y∈{0,1}y \in \{0, 1\}. The quickest thing to try is ordinary least squares on the 0/1 labels, the linear probability model. Suppose it returns p^=−0.15+0.012x\hat p = -0.15 + 0.012x. At x=10x = 10 it predicts −0.03-0.03, a negative probability. At x=100x = 100 it predicts 1.051.05. A line has no reason to stay inside [0,1][0, 1], and its errors can't be homoscedastic either: for a 0/1 outcome, Var(y∣x)=p(1−p)\text{Var}(y \mid x) = p(1-p), which changes with xx.

The fix is to model a quantity that lives on the whole real line. The odds p/(1−p)p/(1-p) run from 0 to ∞\infty, and the log-odds, or logit, runs from −∞-\infty to ∞\infty. Logistic regression sets the logit equal to a linear function:

log⁡p1−p=β0+β1x⟺p=11+e−(β0+β1x).\log\frac{p}{1-p} = \beta_0 + \beta_1 x \quad\Longleftrightarrow\quad p = \frac{1}{1 + e^{-(\beta_0 + \beta_1 x)}}.

The right-hand side is the sigmoid σ(z)=1/(1+e−z)\sigma(z) = 1/(1 + e^{-z}), an S-curve that flattens near 0 and 1 and is steepest at p=1/2p = 1/2. Written as a model, y∼Bernoulli(p)y \sim \text{Bernoulli}(p) with logit(p)=x⊤β\text{logit}(p) = x^\top \beta. It is a generalized linear model with the logit as its link.

Swap the sigmoid for the standard normal CDF Φ\Phi and you get probit. After rescaling the two curves are nearly indistinguishable. Logit coefficients come out roughly 1.6 to 1.8 times the probit ones, because the logistic distribution has standard deviation π/3≈1.81\pi/\sqrt{3} \approx 1.81 where the normal has 1. Logit is the usual default because its coefficients have the odds reading in the next section.