Unit 3 · Topic 3.9 · about 20 minutes

Sampling Distributions for the Difference Between Sample Proportions

Describe how the difference between two sample proportions behaves from sample to sample, and use that to find and interpret probabilities about two populations.

Predict first

Suppose 64% of high school seniors in County A have a driver's license and 52% of seniors in County B do. A researcher takes a random sample of 100 seniors from each county. From sample to sample, p^A varies with a standard deviation of about 0.048, and p^B with a standard deviation of about 0.050. How much will the difference p^A−p^B vary?

Two samples, one difference

Plenty of questions compare two groups: two counties, or two treatments in an experiment. Take a random sample from each population, find each sample proportion, and subtract. That difference, p^1−p^2, is a statistic, so it changes from one pair of samples to the next.

Picture repeating the whole process many times: a new sample from population 1, a new sample from population 2, a new difference. The distribution of all those differences is the sampling distribution of p^1−p^2. When the two samples are independent of each other and the values within each sample are independent, its center and spread come straight from the population proportions p1 and p2 and the sample sizes n1 and n2:

μp^1−p^2=p1−p2
σp^1−p^2=p1(1−p1)n1+p2(1−p2)n2

The mean says the sample difference is on target on average. It does not systematically overestimate or underestimate the true difference, which makes it an unbiased estimator of p1−p2 (see 3.1 Estimators).

Where the 0.069 in the license example comes from
StatisticStandard deviationValue
p^A alone0.64(0.36)1000.048
p^B alone0.52(0.48)1000.050
p^A−p^B0.64(0.36)100+0.52(0.48)1000.069

Each fraction under the square root is the square of one sample proportion's standard deviation from topic 3.2. Add the two squared values, then take the square root. That is why 0.048 and 0.050 combine to 0.069 and not to 0.098. The plus sign matters: the statistic is a difference, but the two sources of chance variation still add.

When the formulas apply

The mean and standard deviation formulas rely on independence, so check two conditions:

  • Randomization condition: the data come from two independent random samples, one from each population.
  • 10% condition: when sampling without replacement, each sample is no more than 10% of its population, so n1≤0.10N1 and n2≤0.10N2.

If the data come from an experiment, the 10% condition does not apply. The randomization condition is met when the treatments are randomly assigned to the experimental units.

The shape needs its own check. The sampling distribution of p^1−p^2 is approximately normal when both samples are large enough, and the normality condition says what large enough means: the expected numbers of successes and failures, n1p1, n1(1−p1), n2p2 and n2(1−p2), must all be at least 10. In this topic the problem hands you p1 and p2, so use them. For the license example the four expected counts are 64, 36, 52 and 48.

Worked exampleHow often do the samples point the wrong way?

Health department records show that 46% of adults in Lake County and 38% of adults in Ridge County got a flu shot last season. A survey firm will take a random sample of 150 adults from Lake County and a separate random sample of 200 adults from Ridge County, then compute p^L−p^R. Find the probability that the samples show a higher flu shot rate in Ridge County, that is, p^L−p^R<0.

  1. Name the statistic and its parameters. The statistic is p^L−p^R, the proportion of sampled Lake County adults who got a flu shot minus the proportion of sampled Ridge County adults who did. The true difference is pL−pR=0.46−0.38=0.08.

  2. Check conditions. Randomization: two independent random samples, one from each county. 10%: 150 and 200 are far less than 10% of the adults in either county. Normality: 150(0.46)=69, 150(0.54)=81, 200(0.38)=76 and 200(0.62)=124 are all at least 10, so the sampling distribution is approximately normal.

  3. Calculate. The mean is 0.08 and the standard deviation is 0.46(0.54)150+0.38(0.62)200≈0.0532. Standardize: z=0−0.080.0532≈−1.50, so P(p^L−p^R<0)≈0.0668 from the table, or 0.0665 with technology and the unrounded z.

  4. Interpret in context. If the firm repeated these two surveys many times, about 6.6% of the pairs of samples would show Ridge County with a higher flu shot rate than Lake County, even though Lake County's true rate is 8 percentage points higher.

Answer.

About 0.066. The differences center at 0.08 and typically miss it by about 0.053, so a difference below 0 is uncommon but far from impossible.

Sampling distribution of p^L−p^R

−0.079−0.0260.0270.080.1330.1860.239Difference in sample proportions, Lake minus Ridgez = −3z = −2z = −1z = 0z = 1z = 2z = 3

Mean 0.08, standard deviation about 0.053. The shaded area to the left of 0, about 0.066, is the chance that the samples rank the counties backwards.

Say what the numbers mean

A mean, a standard deviation or a probability from this distribution only means something with both populations named. For the flu shot surveys:

  • Mean: across all possible pairs of random samples, the difference between the Lake County and Ridge County sample proportions averages 0.08, the true difference.
  • Standard deviation: in repeated random sampling, the difference p^L−p^R typically varies from 0.08 by about 0.053.

The probability works the same way. On its own, 0.0665 is just a number. "About 6.6% of pairs of random samples would show a higher flu shot rate for Ridge County than for Lake County" tells a reader what it means.

Lab

Sampling Distribution Machine

The machine draws samples from one population at a time, so it cannot simulate a difference, but it can show you each piece of the formula. In proportion mode, run one population and write down the standard deviation of the simulated sample proportions. Then run a second population with a different proportion or sample size and write that one down too. Square both, add, and take the square root. The formula says the result is how much the difference between those two sample proportions would vary, and it is always bigger than either one alone.

Open the full Sampling Distribution Machine lab

Check your understanding

1

At a large university, 30% of students in the engineering college and 22% of students in the business college are from out of state. Independent random samples of 120 engineering students and 150 business students are selected. What is the standard deviation of the sampling distribution of p^E−p^B?

2

In which situation is a condition for using the mean and standard deviation formulas for p^1−p^2 NOT met?

3

Random samples of 200 riders are taken from each of two bus routes. The sampling distribution of p^1−p^2, the proportion of riders who pay with a phone app on Route 1 minus the proportion on Route 2, has mean 0.05 and standard deviation 0.046. Which is the best interpretation of 0.046?

4

At one large company, 35% of employees work fully remote. At a second large company, 25% do. Random samples of 100 employees are taken from each company. What is the probability that p^1−p^2, the first company's sample proportion minus the second's, is 0.20 or more?

5

A researcher randomly assigns 150 volunteers to take a new allergy medication and 150 to take a placebo, then records whether each volunteer's symptoms improve. To model p^M−p^P with a normal distribution, which conditions must be met?

Course alignment, for teachers

AP Statistics topic 3.9, Unit 3: Inference for Categorical Data: Proportions.

  • Skill 3.D: Calculate means, standard deviations, and parameters for probability distributions.
  • Skill 4.D: Interpret statistical calculations and results to assess meaning or a claim.
  • Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.