Statistics & Regression

Lesson 1 of 9

Sampling and Standard Error

The sample mean as a random variable, its standard error σ/√n, and why the sample variance divides by n − 1.

The sample mean is a random variable

You poll 1000 voters and 52% say they back candidate A. Ask a different 1000 people and you might get 50.7% or 53.4%. The 52% you report is one draw from a distribution, and that distribution is called the sampling distribution of the estimate.

The setup: a population has mean μ\mu and variance σ2\sigma^2. You draw X1,…,XnX_1, \dots, X_n independently from it. The sample mean is Xˉ=1n∑i=1nXi\bar{X} = \frac{1}{n}\sum_{i=1}^n X_i. Since Xˉ\bar{X} is a function of random variables, it is a random variable too, with its own mean and variance:

E[Xˉ]=μ,Var(Xˉ)=1n2∑i=1nVar(Xi)=σ2nE[\bar{X}] = \mu, \qquad \text{Var}(\bar{X}) = \frac{1}{n^2}\sum_{i=1}^n \text{Var}(X_i) = \frac{\sigma^2}{n}

The first fact says Xˉ\bar{X} is unbiased: averaged over many repeated samples, it lands on μ\mu. The second says how tightly it clusters. The variance step uses independence to drop every covariance term. Remember that, because independence is the assumption that most often fails in real data. The LLN & Central Limit Theorem lesson adds the shape: for large nn, Xˉ\bar{X} is close to normal.

For a poll, each XiX_i is 1 if the voter says yes and 0 otherwise, a Bernoulli(pp) variable with σ2=p(1−p)\sigma^2 = p(1-p). With p=0.52p = 0.52 that is 0.52×0.48=0.24960.52 \times 0.48 = 0.2496. The sample proportion p^\hat{p} is just Xˉ\bar{X} for these 0/1 variables, so everything above applies to it directly.