Sampling distribution of a sample proportion – repeated samples and the normal approximation
MathematicsStatisticsAges 17–18
Loading…
Sign in to playSet the population proportion p and the sample size n, draw thousands of random samples and build the histogram of the sample proportion p̂: it is centred on p, has standard deviation √(p(1 − p)/n) and looks normal when np ≥ 10 and n(1 − p) ≥ 10. Find the probability that p̂ passes a cut-off by simulation, from the exact distribution and with the normal approximation; a finite population (the 10% condition) and the difference of two sample proportions p̂₁ − p̂₂ are included.
Lesson: Sampling distribution of a sample proportion
What it shows
A sample proportion p̂ = X/n changes from sample to sample, and the distribution of all its possible values is its sampling distribution. With independent draws the count X is binomial B(n, p), so p̂ has mean p (it is an unbiased estimator) and standard deviation √(p(1 − p)/n). When np and n(1 − p) are both at least 10 the distribution is close to normal, which is the basis of confidence intervals and tests for a proportion. Sampling without replacement from a small population makes X hypergeometric and the spread smaller; the formula is fine while n is at most 10% of N.
How to use
Set Proportion p and Sample size n, then press Take 1 sample, +100 samples, +1000 samples or Run and watch the histogram grow. Compare it with the Exact distribution and the Normal curve. Type a cut-off c or drag the red line to get the probability three ways. Choose Population N, drag along the lower graph to change n, untick Show p to estimate p, or switch Mode to a difference.
Parameters you can change
- Mode One sample proportion p̂, Difference of two sample proportions p̂₁ − p̂₂
- Population proportion p (or p₁) 0.01–0.99
- Sample size n (or n₁) 1–2000
- Population proportion p₂ 0.01–0.99
- Sample size n₂ 1–2000
- Population size N Very large (independent draws), 100, 200, 500, 1000, 10000
- Event for the probability p̂ ≥ c, p̂ ≤ c
- Cut-off c for p̂ 0–1
- Cut-off c for p̂₁ − p̂₂ -1–1
- Show the population proportion
- Show the exact distribution
- Show the normal curve
- Random seed (same number, same samples) 1–999
Questions to explore
- How does the standard deviation of p̂ change when the sample size goes from 25 to 100?
- With p = 0.05 and n = 50, why is the histogram of p̂ skewed and the normal approximation poor?
- Why is p̂ less spread out than √(p(1 − p)/n) predicts when you sample 400 from a population of 1000?