Unit 1 · Topic 1.13 · about 30 minutes

Experimental Design

Design an experiment with random assignment, choose a design that fits the situation, and say what its results can and cannot show.

Predict first

In a study of a new acne cream, half the participants used the real cream and half used a lookalike cream with no active ingredient. Nobody knew which cream they had. Predict: did the skin of the people using the lookalike cream improve?

Four parts of a well-designed experiment

Suppose a teacher wants to know whether a 10-minute walk just before a math quiz raises scores. Sixty students volunteer. A well-designed experiment has four features.

  1. Comparison of at least two treatment groups. Here, students either walk or sit quietly for the same 10 minutes. The sitting group is the control group, a group created for comparison.
  2. Random assignment of treatments to experimental units. Number the students from 1 to 60 and use a random number generator to pick 30 different numbers. Those students walk, and the other 30 sit.
  3. Replication: more than one experimental unit gets each treatment. With 30 students per group, a difference is much less likely to come down to one or two unusual students.
  4. Direct control of other possible sources of variation in the response, by keeping them the same for every unit: the same quiz, the same room, the same time of day, the same walking route.

Control groups, placebos and blinding

A control group may receive a treatment that differs from the one being tested, such as a placebo, an inactive treatment that looks like the real thing. Then both groups believe they were treated, and believing alone cannot explain a difference between them. The placebo effect is the difference between the average response to a placebo and the average response to no treatment at all.

Blinding keeps expectations out of the results. In a single-blind, or single-masked, experiment, the participants do not know which treatment they received but the researchers who interact with them do, or the other way around. In a double-blind, or double-masked, experiment, neither the participants nor the researchers who interact with them know. The walking study cannot blind the students, who know whether they walked, but a second teacher, who gives the quiz and grades it without being told who walked, can be kept unaware, which makes it single-blind.

Why random assignment matters

An extraneous variable, or extraneous source of variation, is known or believed to affect the response but is not one of the explanatory variables being studied. For the quiz, the amount of sleep the night before, earlier math grades and test anxiety all qualify.

The purpose of random assignment is to make the treatment groups as similar as possible on every extraneous variable, including the ones nobody thought to measure. When it works, the distribution of each extraneous variable is about the same in every group. Here is one random split of the 60 volunteers, compared on the hours they slept the night before.

Hours of sleep in the two randomly assigned groups

WalkSit56789Hours of sleep the night before

Random assignment produced two groups with nearly the same distribution of sleep: medians of 7.25 and about 7.1 hours, and the same third quartile. A difference in quiz scores would be hard to blame on sleep.

Confounding in an experiment

A confounding variable in an experiment is related to the explanatory variable in a way that makes it hard to tell which of the two is changing the response. Suppose the teacher let students choose whether to walk. The walkers might be the energetic morning people who would have scored higher anyway, and energy would be tangled up with walking. Or suppose every walker took the quiz at 8 a.m. and every sitter at 2 p.m.; then time of day is confounded with the treatment. Random assignment and direct control are what keep the potential for confounding low in a well-designed experiment.

Three designs

In a completely randomized design, the treatments are assigned to the experimental units completely at random. The walking study is one. The groups are often the same size, but they do not have to be.

Sometimes you know in advance that a particular variable will affect the response. A blocking variable is such a source of extraneous variation. In a randomized block design, the units are first sorted into groups called blocks, so that the units within a block are alike on the blocking variable. Then the treatments are randomly assigned within each block, so that every treatment appears in every block. Blocking separates the variation caused by the blocking variable from the rest of the variation in the response, which makes the comparison of treatments more precise: inside a block, the treatments are compared on similar units.

A matched pairs design is a randomized block design with only two treatments. Units are paired by matching them on variables that affect the response, and a random choice decides which member of each pair gets which treatment. Alternatively, each unit receives both treatments in a randomly chosen order, such as a runner testing two brands of running shoes, one brand each week, with a coin flip deciding which brand comes first.

Which design fits depends on the goal of the study, the experimental units and the variables involved. If nothing known in advance divides the units, a completely randomized design is enough. If a known variable affects the response, block on it. If there are only two treatments and the units can be matched, or can each receive both treatments, a matched pairs design compares the treatments within pairs of similar units, or within the same unit.

Worked exampleDesigning a randomized block experiment

A school wants to compare two online algebra programs, A and B, by how much students' test scores improve over six weeks. The 80 volunteers include 40 students in honors algebra and 40 in regular algebra, and the school expects the two courses to improve at different rates no matter which program they use. Describe a design.

  1. Choose the blocking variable. Course level, honors or regular, is expected to affect improvement and is not what the study is about, so block on it.

  2. Form the blocks. Block 1 is the 40 honors students. Block 2 is the 40 regular students.

  3. Randomly assign within each block. In each block, number the students from 1 to 40 and use a random number generator to choose 20 different numbers. Those students use Program A, and the other 20 use Program B.

  4. Compare within blocks. Compare the mean improvement for A and B among the honors students, and separately among the regular students. Differences between the two course levels no longer blur the comparison of the programs.

Answer.

A randomized block design, blocked by course level, with programs randomly assigned within each block.

What the results can show

Random assignment of treatments to experimental units is what allows a cause-and-effect conclusion, because it reduces the potential for confounding variables. If the treatment groups differ clearly in their response, the treatments can be credited with causing the difference.

Who the conclusion applies to is a separate question, answered by how the units were selected. It is often unethical or impractical to randomly select people for an experiment, so most experiments use volunteers. Their results then apply to units similar to the volunteers. The walking study could show that walking caused a change in quiz scores, but only for students like those 60.

What a study's design allows you to conclude
Treatments randomly assignedTreatments not randomly assigned
Units randomly selectedCause and effect, for the whole populationGeneralize to the population, but no cause and effect
Units not randomly selectedCause and effect, for units like those in the studyNo cause and effect, and only units like these

Lab

Who Can We Generalize To? What Caused It?

Put the table above to work. For each study, decide whether the units were randomly selected and whether treatments were randomly assigned, then answer the lab's two questions: can the results be generalized to the whole population the individuals came from, and can the study show cause and effect?

Open the full Who Can We Generalize To? What Caused It? lab

Check your understanding

1

To compare two sunscreens, a researcher has each of 30 volunteers apply Sunscreen A to one forearm and Sunscreen B to the other, with a coin flip deciding which arm gets A. Which design is this?

2

A study of a sleep aid has three groups: the real pill, a placebo pill and no pill. The mean extra sleep was 42 minutes with the real pill, 18 minutes with the placebo and 5 minutes with no pill. What is the placebo effect in this study?

3

In a vaccine trial, the participants do not know whether they received the vaccine or a saltwater shot, and neither do the nurses who give the shots and check on the participants afterward. What is this feature of the design called?

4

Researchers randomly assigned 200 volunteers to a high-fiber diet or their regular diet. After 12 weeks, the high-fiber group had a much lower mean cholesterol level. Which conclusion is justified?

5

A company will compare two pain relievers on 120 volunteers with headaches. It expects the pills to work differently for people with migraines than for people with ordinary headaches. What is the best reason to block by headache type?

Course alignment, for teachers

AP Statistics topic 1.13, Unit 1: Exploring One-Variable Data and Collecting Data.

  • Skill 2.A: Identify information to answer a question or solve a problem.
  • Skill 2.B: Justify an appropriate method for ethically gathering and representing data.