Sampling distribution of a sample proportion – repeated samples and the normal approximation

MathematicsStatisticsAges 17–18

Loading…

Use with my class ✨ Customize with AI Report a problem

Set the population proportion p and the sample size n, draw thousands of random samples and build the histogram of the sample proportion p̂: it is centred on p, has standard deviation √(p(1 − p)/n) and looks normal when np ≥ 10 and n(1 − p) ≥ 10. Find the probability that p̂ passes a cut-off by simulation, from the exact distribution and with the normal approximation; a finite population (the 10% condition) and the difference of two sample proportions p̂₁ − p̂₂ are included.

Lesson: Sampling distribution of a sample proportion

What it shows

A sample proportion p̂ = X/n changes from sample to sample, and the distribution of all its possible values is its sampling distribution. With independent draws the count X is binomial B(n, p), so p̂ has mean p (it is an unbiased estimator) and standard deviation √(p(1 − p)/n). When np and n(1 − p) are both at least 10 the distribution is close to normal, which is the basis of confidence intervals and tests for a proportion. Sampling without replacement from a small population makes X hypergeometric and the spread smaller; the formula is fine while n is at most 10% of N.

How to use

Set Proportion p and Sample size n, then press Take 1 sample, +100 samples, +1000 samples or Run and watch the histogram grow. Compare it with the Exact distribution and the Normal curve. Type a cut-off c or drag the red line to get the probability three ways. Choose Population N, drag along the lower graph to change n, untick Show p to estimate p, or switch Mode to a difference.

Parameters you can change

  • Mode One sample proportion p̂, Difference of two sample proportions p̂₁ − p̂₂
  • Population proportion p (or p₁) 0.01–0.99
  • Sample size n (or n₁) 1–2000
  • Population proportion p₂ 0.01–0.99
  • Sample size n₂ 1–2000
  • Population size N Very large (independent draws), 100, 200, 500, 1000, 10000
  • Event for the probability p̂ ≥ c, p̂ ≤ c
  • Cut-off c for p̂ 0–1
  • Cut-off c for p̂₁ − p̂₂ -1–1
  • Show the population proportion
  • Show the exact distribution
  • Show the normal curve
  • Random seed (same number, same samples) 1–999

Questions to explore

  1. How does the standard deviation of p̂ change when the sample size goes from 25 to 100?
  2. With p = 0.05 and n = 50, why is the histogram of p̂ skewed and the normal approximation poor?
  3. Why is p̂ less spread out than √(p(1 − p)/n) predicts when you sample 400 from a population of 1000?