Unit 3 · Topic 3.12 · about 20 minutes
Setting Up a Test for the Difference Between Two Population Proportions
Recognize when a two-sample z-test fits, write its hypotheses in context, and verify its conditions before any calculation.
Predict first
A college surveys a random sample of 100 first-year students and a separate random sample of 300 upper-level students. Of the first-year students, 45 use the campus gym at least once a week; of the upper-level students, 90 do. Suppose the two groups really have the same proportion of weekly gym users. What is the best estimate of that shared proportion?
Comparing two proportions with a test
Topics 3.10 and 3.11 estimated how far apart two proportions are. A significance test asks a yes-or-no question instead: do the data give convincing evidence that the two population proportions differ, or that a particular one is larger? The procedure is the two-sample z-test for the difference between two population proportions. It works with data from two independent random samples or from a randomized experiment with two treatments.
Define the parameters before you write anything else. Each definition needs the response variable and its population in context. " is the proportion of all licensed drivers in the state under 25 who texted while driving in the past month" does the job. " is the proportion for group 1" does not.
Writing the hypotheses
The null hypothesis says there is no difference:
The alternative hypothesis comes from the question you are trying to answer, and you choose it before looking at the data. Each hypothesis can be written either way, and both forms mean the same thing, so use whichever reads more naturally.
| The question asks whether | Written as a difference | |
|---|---|---|
| is larger than | ||
| is smaller than | ||
| the proportions differ, in either direction |
Conditions, with one change
The randomization condition and the 10% condition read exactly as they did for intervals:
- Randomization condition: the data come from two independent random samples or a randomized experiment.
- 10% condition: when sampling without replacement, and . It is unnecessary when the data come from a randomized experiment.
The normality condition is where the test differs. The test is carried out assuming is true, so it uses the combined, or pooled, proportion of successes:
Then , , and must all be at least 10. For the gym survey, gives 33.75, 66.25, 101.25 and 198.75, so the condition is met.
Sort it
Which procedure fits each study? Tap a card, then tap its bin.
One-sample z-test for a proportion
Two-sample z-test for a difference
Two-sample z-interval for a difference
Worked exampleSetting up a test about texting and driving
A state transportation agency selects a random sample of 250 licensed drivers under age 25 and a separate random sample of 400 licensed drivers age 25 or older from its license records. In anonymous surveys, 70 of the younger drivers and 76 of the older drivers admit to texting while driving in the past month. The agency wants to know whether the data give convincing evidence that the proportion who text while driving is higher for drivers under 25. Set up the appropriate test, but do not carry it out.
Identify the procedure and parameters. A two-sample z-test for the difference between two population proportions. Let be the proportion of all licensed drivers in the state under 25 who texted while driving in the past month, and the proportion of all licensed drivers in the state 25 or older who did.
State the hypotheses. The agency asks whether the younger drivers' proportion is higher, so the test is one-sided: and . Written as a difference, and .
Check randomization and 10%. The data come from two independent random samples. The state has far more than 2,500 licensed drivers under 25 and far more than 4,000 drivers 25 or older, so each sample is no more than 10% of its population.
Check normality with the pooled proportion. . Then , , and , all at least 10.
All three conditions are met, so the two-sample z-test with is appropriate. The next lesson carries out a test like this one.
Check your understanding
A researcher wants to know whether the proportion of adults who get a flu shot is lower in a state's rural counties than in its urban counties. She takes a random sample of adults from the rural counties and a separate random sample from the urban counties. Let and be the proportions of all adults in the rural and urban counties who get a flu shot. Which hypotheses fit her question?
In an experiment, 24 of 120 seedlings given a new fertilizer and 15 of 100 seedlings given the standard fertilizer wilt within a week. A two-sample z-test will test whether the proportions that wilt differ. Which is the correct check of the normality condition for this test?
A city wants to know whether the proportion of residents who support a new bike lane differs between residents who own a car and residents who do not. It takes a random sample from each group. Which significance test should the city use?
A researcher randomly assigns 200 volunteers to use either a new sleep app or a standard alarm for two weeks, then records whether each volunteer reports feeling rested. Before running a two-sample z-test, which statement about the conditions is correct?
Why does the normality condition for a two-sample z-test use the pooled proportion instead of the two separate sample proportions?
Course alignment, for teachers
AP Statistics topic 3.12, Unit 3: Inference for Categorical Data: Proportions.
- Skill 2.C: Identify appropriate statistical inference methods.
- Skill 2.E: Identify the null and alternative hypotheses.
- Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.