Unit 4 · Topic 4.10 · about 30 minutes
Carrying Out a Test for the Difference Between Two Population Means
Calculate and interpret the test statistic and p-value for a two-sample t-test, and turn the result into a justified conclusion about the two populations.
Predict first
Two studies each find a 4-point difference between the mean scores of two groups, and the spread within the groups is about the same in both studies. Study A has 15 people in each group. Study B has 150 in each group. Which study's two-sample t-test gives the smaller p-value?
The test statistic
The two-sample t statistic measures how far the observed difference in sample means is from the difference the null hypothesis claims, which is 0, in standard errors:
When the null hypothesis is true, this statistic has a t-distribution. As with the two-sample interval in Topic 4.7, technology finds its degrees of freedom, and they fall between the smaller of and and . The p-value is the area in the direction of the alternative, found with technology (2-SampTTest on a TI-84) or bracketed with the t table. Using the smaller of and , called the conservative degrees of freedom, gives a slightly larger p-value, which is the cautious direction.
Reaction times after chewing gum, in milliseconds
The same data as in Topic 4.9: both groups of 15 are roughly symmetric with no outliers.
| Group | Mean | Standard deviation | |
|---|---|---|---|
| Caffeinated gum | 15 | 278.67 | 20.96 |
| Gum without caffeine | 15 | 295.67 | 23.14 |
Worked exampleDoes caffeinated gum speed up reactions?
In the experiment from Topic 4.9, 30 volunteers were randomly assigned to chew caffeinated gum or gum without caffeine, 15 in each group, and their reaction times were recorded. The boxplots and summary statistics are above. Do the data give convincing statistical evidence, at , that caffeinated gum lowers mean reaction time for people like these volunteers?
Procedure and hypotheses. Two-sample t-test for a difference between two population means. and are the true mean reaction times, in ms, for volunteers like these who chew caffeinated gum and gum without caffeine. and .
Conditions. Random: the volunteers were randomly assigned to the two kinds of gum. 10%: not needed for a randomized experiment. Sample data: each group has only 15 volunteers, and both boxplots are free of strong skewness and outliers.
Calculate. ms and ms, so . Technology gives and a p-value, the area to the left of , of about 0.022. With the conservative the p-value is about 0.027, and the decision is the same.
Conclude. Because the p-value of 0.022 is less than , reject . There is convincing statistical evidence that the true mean reaction time for volunteers like these is lower with caffeinated gum than with gum without caffeine.
, p-value about 0.022, so reject . Because the gum was randomly assigned, the lower mean reaction time can be credited to the caffeine.
Reading the p-value and the decision
Interpret the p-value by stating the assumption it rests on: that the population means are equal. Assuming the true mean reaction time is the same with both kinds of gum, there is about a 0.022 probability of getting a difference in sample means of ms or less, which is a test statistic of or less, just from the random assignment of volunteers to groups.
The formal decision compares that p-value with . If the p-value is less than or equal to , reject ; if it is greater, fail to reject . The conclusion goes back to the question in context, in terms of the alternative hypothesis, naming both parameters and both populations, and it never claims to have proved anything. That conclusion is the statistical reasoning behind the answer to the investigative question, here "Does caffeinated gum lower reaction times for people like these?"
Worked exampleWhen the difference is not convincing
A large school district, with more than 1,000 students in each grade, wants to know whether mean daily screen time differs between its 9th graders and its 12th graders. It selects independent random samples of 40 ninth graders and 35 twelfth graders and collects their phones' screen-time reports. Ninth graders: minutes, . Twelfth graders: minutes, . Both samples are moderately skewed to the right. Test at .
Procedure and hypotheses. Two-sample t-test for a difference between two population means, where and are the mean daily screen times, in minutes, of all 9th graders and all 12th graders in the district. and .
Conditions. Random: independent random samples from each grade. 10%: 40 and 35 are each less than 10% of the more than 1,000 students in each grade. Sample data: both samples have at least 30 students, so the skew does not stop the test.
Calculate. minutes and . With , the two-sided p-value is about 0.218.
Conclude. Because the p-value of 0.218 is greater than , fail to reject . There is not convincing statistical evidence that the mean daily screen time of all 9th graders in the district differs from the mean daily screen time of all 12th graders in the district.
Fail to reject . A 36-minute gap between the sample means is the kind of difference random sampling would often produce even if the two grades had the same mean screen time.
Check your understanding
Independent random samples of hybrid-car owners and gas-car owners in a large city report their weekly driving distance in miles. Hybrid owners: , , . Gas-car owners: , , . Find the two-sample t statistic for . Round to two decimal places.
In a randomized experiment, patients recovering from knee surgery receive a new or a standard physical therapy program. For against , where each is the true mean gain in range of motion, the p-value is 0.04. Which is a correct interpretation of the p-value?
A large grocery store takes independent random samples of 36 Saturday shoppers and 40 Tuesday shoppers and tests against , where each is the mean amount spent per visit. The p-value is 0.003. Which conclusion is appropriate at ?
A student carries out a two-sample t-test with sample sizes and but has no technology, so she uses the conservative degrees of freedom. How many degrees of freedom does she use?
A high school asks, "Do students who eat breakfast have a higher mean score on the state reading assessment than students who do not?" It selects independent random samples of both kinds of students from its records, and a two-sample t-test of gives a p-value of 0.004. At , which answer to the question is best supported?
Course alignment, for teachers
AP Statistics topic 4.10, Unit 4: Inference for Quantitative Data: Means.
- Skill 3.E: Calculate appropriate statistical inference method results.
- Skill 4.F: Interpret results of statistical inference methods.
- Skill 4.G: Justify a claim based on statistical inference method results.