Unit 4 · Topic 4.10 · about 30 minutes

Carrying Out a Test for the Difference Between Two Population Means

Calculate and interpret the test statistic and p-value for a two-sample t-test, and turn the result into a justified conclusion about the two populations.

Predict first

Two studies each find a 4-point difference between the mean scores of two groups, and the spread within the groups is about the same in both studies. Study A has 15 people in each group. Study B has 150 in each group. Which study's two-sample t-test gives the smaller p-value?

The test statistic

The two-sample t statistic measures how far the observed difference in sample means is from the difference the null hypothesis claims, which is 0, in standard errors:

t=(x‾1−x‾2)−0s12n1+s22n2

When the null hypothesis is true, this statistic has a t-distribution. As with the two-sample interval in Topic 4.7, technology finds its degrees of freedom, and they fall between the smaller of n1−1 and n2−1 and n1+n2−2. The p-value is the area in the direction of the alternative, found with technology (2-SampTTest on a TI-84) or bracketed with the t table. Using the smaller of n1−1 and n2−1, called the conservative degrees of freedom, gives a slightly larger p-value, which is the cautious direction.

Reaction times after chewing gum, in milliseconds

CaffeineNo caffeine260280300320Reaction time (ms)

The same data as in Topic 4.9: both groups of 15 are roughly symmetric with no outliers.

Reaction times, in milliseconds, in the caffeinated gum experiment
GroupnMeanStandard deviation
Caffeinated gum15278.6720.96
Gum without caffeine15295.6723.14

Worked exampleDoes caffeinated gum speed up reactions?

In the experiment from Topic 4.9, 30 volunteers were randomly assigned to chew caffeinated gum or gum without caffeine, 15 in each group, and their reaction times were recorded. The boxplots and summary statistics are above. Do the data give convincing statistical evidence, at α=0.05, that caffeinated gum lowers mean reaction time for people like these volunteers?

  1. Procedure and hypotheses. Two-sample t-test for a difference between two population means. μC and μN are the true mean reaction times, in ms, for volunteers like these who chew caffeinated gum and gum without caffeine. H0:μC−μN=0 and Ha:μC−μN<0.

  2. Conditions. Random: the volunteers were randomly assigned to the two kinds of gum. 10%: not needed for a randomized experiment. Sample data: each group has only 15 volunteers, and both boxplots are free of strong skewness and outliers.

  3. Calculate. x‾C−x‾N=278.67−295.67=−17.00 ms and SE=20.96215+23.14215≈8.06 ms, so t=−17.00−08.06≈−2.11. Technology gives df≈27.7 and a p-value, the area to the left of −2.11, of about 0.022. With the conservative df=14 the p-value is about 0.027, and the decision is the same.

  4. Conclude. Because the p-value of 0.022 is less than α=0.05, reject H0. There is convincing statistical evidence that the true mean reaction time for volunteers like these is lower with caffeinated gum than with gum without caffeine.

Answer.

t≈−2.11, p-value about 0.022, so reject H0. Because the gum was randomly assigned, the lower mean reaction time can be credited to the caffeine.

Reading the p-value and the decision

Interpret the p-value by stating the assumption it rests on: that the population means are equal. Assuming the true mean reaction time is the same with both kinds of gum, there is about a 0.022 probability of getting a difference in sample means of −17.00 ms or less, which is a test statistic of −2.11 or less, just from the random assignment of volunteers to groups.

The formal decision compares that p-value with α. If the p-value is less than or equal to α, reject H0; if it is greater, fail to reject H0. The conclusion goes back to the question in context, in terms of the alternative hypothesis, naming both parameters and both populations, and it never claims to have proved anything. That conclusion is the statistical reasoning behind the answer to the investigative question, here "Does caffeinated gum lower reaction times for people like these?"

Worked exampleWhen the difference is not convincing

A large school district, with more than 1,000 students in each grade, wants to know whether mean daily screen time differs between its 9th graders and its 12th graders. It selects independent random samples of 40 ninth graders and 35 twelfth graders and collects their phones' screen-time reports. Ninth graders: x‾9=412 minutes, s9=118. Twelfth graders: x‾12=448 minutes, s12=131. Both samples are moderately skewed to the right. Test at α=0.05.

  1. Procedure and hypotheses. Two-sample t-test for a difference between two population means, where μ9 and μ12 are the mean daily screen times, in minutes, of all 9th graders and all 12th graders in the district. H0:μ9=μ12 and Ha:μ9≠μ12.

  2. Conditions. Random: independent random samples from each grade. 10%: 40 and 35 are each less than 10% of the more than 1,000 students in each grade. Sample data: both samples have at least 30 students, so the skew does not stop the test.

  3. Calculate. SE=118240+131235≈28.96 minutes and t=412−44828.96≈−1.24. With df≈69.1, the two-sided p-value is about 0.218.

  4. Conclude. Because the p-value of 0.218 is greater than α=0.05, fail to reject H0. There is not convincing statistical evidence that the mean daily screen time of all 9th graders in the district differs from the mean daily screen time of all 12th graders in the district.

Answer.

Fail to reject H0. A 36-minute gap between the sample means is the kind of difference random sampling would often produce even if the two grades had the same mean screen time.

Check your understanding

1

Independent random samples of hybrid-car owners and gas-car owners in a large city report their weekly driving distance in miles. Hybrid owners: n1=32, x‾1=68.2, s1=9.5. Gas-car owners: n2=30, x‾2=63.9, s2=11.2. Find the two-sample t statistic for H0:μ1=μ2. Round to two decimal places.

2

In a randomized experiment, patients recovering from knee surgery receive a new or a standard physical therapy program. For H0:μnew=μstd against Ha:μnew>μstd, where each μ is the true mean gain in range of motion, the p-value is 0.04. Which is a correct interpretation of the p-value?

3

A large grocery store takes independent random samples of 36 Saturday shoppers and 40 Tuesday shoppers and tests H0:μSat=μTue against Ha:μSat>μTue, where each μ is the mean amount spent per visit. The p-value is 0.003. Which conclusion is appropriate at α=0.01?

4

A student carries out a two-sample t-test with sample sizes n1=22 and n2=25 but has no technology, so she uses the conservative degrees of freedom. How many degrees of freedom does she use?

5

A high school asks, "Do students who eat breakfast have a higher mean score on the state reading assessment than students who do not?" It selects independent random samples of both kinds of students from its records, and a two-sample t-test of Ha:μB>μN gives a p-value of 0.004. At α=0.05, which answer to the question is best supported?

Practice

Practice until it is automatic

New numbers every time. Each one is checked the moment you answer, with the full working shown.

Comparing two means practice page

Course alignment, for teachers

AP Statistics topic 4.10, Unit 4: Inference for Quantitative Data: Means.

  • Skill 3.E: Calculate appropriate statistical inference method results.
  • Skill 4.F: Interpret results of statistical inference methods.
  • Skill 4.G: Justify a claim based on statistical inference method results.