Unit 4 · Topic 4.7 · about 25 minutes
Constructing a Confidence Interval for the Difference Between Two Population Means
Estimate the difference between two population means from independent samples or a randomized experiment, with conditions checked and the interval interpreted.
Predict first
Three studies compare reaction times. Which one calls for a two-sample t-interval rather than a paired analysis?
Two independent groups
When the data come from two separate groups, either two independent random samples or two groups formed by random assignment in an experiment, there is no pairing to use. The point estimate of is the difference in sample means, , and the procedure is the two-sample t-interval for a difference between two population means:
The square root is the standard error of the difference: the formula from Topic 4.6 with the sample standard deviations and in place of and . The margin of error is times that standard error.
The parameter is , and defining it means naming both populations, the response variable and which group is subtracted from which. For an experiment on volunteers, the populations are people like the ones in the study under each treatment.
Degrees of freedom
The critical value comes from a t-distribution, but the degrees of freedom are no longer just . Technology computes them from both sample sizes and both standard deviations, and the result is usually not a whole number. Whatever it is, it falls between the smaller of and at the low end and at the high end.
Use the technology value when you can; on a TI-84, 2-SampTInt reports it (answer No to Pooled). If you are working from the printed table, the smaller of and , called the conservative degrees of freedom, gives a slightly larger and so a slightly wider interval, which errs on the safe side. AP scoring guidelines have accepted either choice when the work shows which one was used.
Three conditions
- Randomization condition: the data come from two independent random samples or from a randomized experiment.
- 10% condition: when sampling without replacement, each sample is at most 10% of its population, and . This condition is unnecessary when the data come from a randomized experiment.
- Sample data condition: both samples have at least 30 observations, or both population distributions are approximately normal. If either sample has fewer than 30, both sample distributions should be free from strong skewness and outliers.
The last rule catches people: one small sample means you graph both samples, not just the small one.
Recall quiz scores (out of 30) by note-taking method
Both groups are roughly symmetric with no outliers, so the sample data condition is met even though each group has only 16 students.
| Group | Mean | Standard deviation | |
|---|---|---|---|
| By hand | 16 | 20.875 | 3.649 |
| Laptop | 16 | 17.750 | 3.804 |
Worked exampleHand notes versus laptop notes
A psychology teacher recruits 32 student volunteers and randomly assigns 16 to take notes by hand and 16 to take notes on a laptop during the same 15-minute video lecture. A week later, all 32 take a 30-point recall quiz without their notes. The boxplots and summary statistics are above. Construct and interpret a 95% confidence interval for the difference in mean recall score.
Name it. Two-sample t-interval for , where is the true mean recall score for students like these who take notes by hand and is the true mean recall score for students like these who take notes on a laptop.
Check conditions. Random: the volunteers were randomly assigned to the two methods. 10%: not needed, because this is a randomized experiment. Sample data: both groups have fewer than 30 students, so both boxplots must be free of strong skewness and outliers, and they are.
Calculate. and . Technology gives and , so the margin of error is about 2.69 and the interval is about points. With the conservative , and the interval is about .
Interpret in context. We are 95% confident that the interval from 0.43 to 5.82 points contains the true difference in mean recall score (by hand minus laptop) for students like those in this study.
About 0.43 to 5.82 points, for the mean recall score by hand minus the mean recall score on a laptop. Topic 4.8 takes up what this interval says about the claim that the method matters.
Check your understanding
A nutrition researcher selects independent random samples of 35 adults from City A and 40 adults from City B and records each adult's daily sugar intake in grams. She wants to estimate how much the mean daily sugar intake differs between the two cities. Which procedure is appropriate?
Independent random samples give , and , . Find the standard error of . Round to three decimal places.
A two-sample t-interval is built from independent random samples of sizes and . Which value could be the degrees of freedom reported by technology?
In an experiment, 38 volunteers are randomly assigned to a new allergy medicine (18 people) or a standard one (20 people), and the hours of symptom relief are recorded. A boxplot for the new-medicine group is roughly symmetric, but the standard group's boxplot is strongly skewed to the right with two high outliers. Which statement about the conditions for a two-sample t-interval is correct?
A consumer group tests independent random samples of Brand X and Brand Y AA batteries from store shelves and records how many hours each one lasts in a digital camera. It builds a two-sample t-interval for . Which is the best definition of the parameter?
Course alignment, for teachers
AP Statistics topic 4.7, Unit 4: Inference for Quantitative Data: Means.
- Skill 2.C: Identify appropriate statistical inference methods.
- Skill 3.E: Calculate appropriate statistical inference method results.
- Skill 4.E: Justify the use of a chosen statistical inference method by verifying conditions.