Statistics & Regression

Lesson 2 of 9

Hypothesis Tests, p-values and Confidence Intervals

How to decide whether a result is real or luck: null hypotheses, p-values, confidence intervals, the two kinds of error, power, and why the best of many traders is usually lucky.

The null hypothesis and the p-value

A trader tells you 116 of her last 200 trades made money. Is she good, or could a coin have done that? Hypothesis testing makes the question precise. The null hypothesis H0H_0 is the boring explanation: each trade wins with probability p=0.5p = 0.5. The alternative H1H_1 is her claim, p>0.5p > 0.5.

Next you need a test statistic, a number that measures how far the data sit from what H0H_0 predicts. Under H0H_0 the win count has mean 100100 and standard deviation 200⋅0.25≈7.07\sqrt{200 \cdot 0.25} \approx 7.07, so

z=116−1007.07≈2.26z = \frac{116 - 100}{7.07} \approx 2.26

The p-value is the probability, computed assuming H0H_0 is true, of a statistic at least as extreme as the one observed. Here P(Z≥2.26)≈0.012P(Z \ge 2.26) \approx 0.012 (the exact binomial tail is 0.014). If you fixed a significance level α=0.05\alpha = 0.05 in advance, then 0.012<0.050.012 < 0.05 and you reject H0H_0.

This is a one-tailed test, because only a high win rate supports her claim. If you were asking whether a coin is biased in either direction, you would count both tails: the two-tailed p-value is 2×0.012≈0.0242 \times 0.012 \approx 0.024. Pick the tail before looking at the data. Choosing it afterwards quietly halves your p-value.

The interview trap is the interpretation. A p-value of 0.012 does not mean there is a 1.2% chance she has no skill. It is P(data this extreme∣H0)P(\text{data this extreme} \mid H_0). Turning that into P(H0∣data)P(H_0 \mid \text{data}) needs a prior on how common skilled traders are, which is the reversal covered in Bayes' Theorem.