Statistics & Regression

Lesson 4 of 9

Maximum Likelihood

Pick the parameter that makes the observed data most probable, read its standard error off the curvature of the log-likelihood, and see how a prior turns MLE into MAP.

The likelihood function

A coin lands heads 7 times in 10 flips. What is your best guess for its heads probability pp? Most people say 0.7 without thinking. Maximum likelihood is the method that makes that instinct precise and extends it to models where the answer is less obvious.

Fix the data and treat the probability of seeing it as a function of the parameter. That function is the likelihood:

L(θ)=∏i=1nf(xi;θ)L(\theta) = \prod_{i=1}^n f(x_i; \theta)

For the coin (in the order observed), L(p)=p7(1−p)3L(p) = p^7(1-p)^3. Plug in candidates:

L(0.5)=0.510≈0.000977,L(0.7)=0.77⋅0.33≈0.002224L(0.5) = 0.5^{10} \approx 0.000977, \qquad L(0.7) = 0.7^7 \cdot 0.3^3 \approx 0.002224

The data are about 2.3 times more probable if p=0.7p = 0.7 than if the coin is fair. The maximum likelihood estimate (MLE) θ^\hat\theta is the parameter value where LL peaks: the explanation under which what you saw was most likely to happen.

One warning that interviewers like to probe: L(p)L(p) is not a probability distribution over pp. It does not integrate to 1 in pp, and L(0.7)L(0.7) is not "the probability that p=0.7p = 0.7". It is the probability of the data, given pp. Turning it into a statement about pp itself needs a prior, which is the last section.