Hypothesis testing – p-values, critical regions and Type I and II errors
MathematicsStatisticsAges 17–18
Loading…
Sign in to playSet the null hypothesis H₀, a one- or two-tailed alternative H₁ and the significance level α for an exact binomial or Poisson test, a one-proportion z-test or a test for a mean (z or t), and see the null distribution with the critical region and the p-value shaded. The Errors and power screen adds the distribution for a chosen true value, showing the Type I error, the Type II error, the power and the power curve as n and the effect size change; simulation repeats the test to estimate the error rates.
Lesson: Hypothesis testing
What it shows
A hypothesis test asks whether data give enough evidence against a null hypothesis H₀. Assuming H₀ is true, the p-value is the probability of a result at least as extreme as the one observed, in the direction of H₁. If the p-value is at most the significance level α, the result lies in the critical region and H₀ is rejected. Binomial and Poisson tests are exact, so the actual significance level is usually below α. The proportion test uses the normal approximation; the mean test uses z with known σ or t with n − 1 degrees of freedom.
How to use
On Test, choose the Test, the Alternative and Significance α, set n and the value in H₀, then drag the orange triangle or move the observed value; press Random result when H₀ is true to see how often chance alone gives a small p-value. On Errors and power, set the True value, read α, β and power, drag the power curve, and press +100 tests, H₀ true or +100 tests, true value to simulate.
Parameters you can change
- Screen Test, Errors and power
- Test Binomial (exact), Poisson (exact), Proportion (z, normal approximation), Mean (z or t)
- Alternative hypothesis H₁ One-tailed: less than (<), One-tailed: greater than (>), Two-tailed (≠)
- Significance level α 0.1 (10%), 0.05 (5%), 0.025 (2.5%), 0.01 (1%), 0.001 (0.1%)
- Sample size n (number of trials) 1–500
- Proportion p₀ in H₀ 0.01–0.99
- Mean λ₀ in H₀ (Poisson) 0.1–50
- Observed number of successes x 0–1000
- Observed number of events x (Poisson) 0–1000
- Mean μ₀ in H₀ -1000–1000
- Observed sample mean x̄ -1000–1000
- Standard deviation σ (or sample s) 0.01–1000
- Population standard deviation σ known: z-test, σ unknown, use s: t-test
- True value of p (Errors screen) 0.01–0.99
- True value of λ (Errors screen) 0.1–50
- True value of μ (Errors screen) -1000–1000
- Random seed (same number, same simulation) 1–999
Questions to explore
- If H₀ is true, how often does a test at the 5% level wrongly reject it?
- Why is the actual significance level of a binomial test often smaller than α?
- How can you increase the power of a test without changing α?