Evaluating statistical claims
Reading position saved on this browser; this is not a completion record.
Start here
Match the strength and scope of a claim to the study design.
The central idea: two different kinds of randomization answer two different questions
A study’s mathematical calculations can be correct while its conclusion is unjustified. Before interpreting a result, separate how units entered the study from how treatments were assigned. College Board explicitly distinguishes these roles in this domain.[3]
Random sampling supports generalization
Random sampling uses a chance-based selection process to choose a sample from a defined population. It supports inference to that population when the sampling and response conditions are appropriate. It does not automatically support conclusions about a different city, age group, species, or future setting.
A simple random sample gives all samples of a specified size the same selection probability. “Ask whoever is available” and “let interested people volunteer” are not equivalent designs, even if many people participate. “Random” in an everyday sentence is not a substitute for a described chance process.
Random assignment supports causal comparisons
Random assignment allocates experimental units to treatments by chance. In a well-run comparison, it reduces systematic confounding and supports attributing a statistically convincing difference to the assigned treatment. It balances other characteristics on average over assignments, not perfectly in every realized experiment.[10]
Random assignment does not make a volunteer sample representative of everyone. A volunteer experiment can support a causal comparison for its participants while leaving broad population generalization uncertain.
No random assignment | Random assignment | |
|---|---|---|
No random sample | Describe observed participants; generalization and causation need caution. | Causal comparison may be supported for participants; broad generalization is limited. |
Random sample | Population-level association may be supported; causation is not established by sampling alone. | Population-level causal inference may be supported when execution and evidence are adequate. |
The matrix describes what a design can support, not a promise that every observed difference is real. You must still consider response quality, fair implementation, and whether the result is convincing rather than readily attributable to chance.
Observational studies, experiments, and confounding
What makes a study an experiment?
In an observational study, researchers observe exposures or behaviors without imposing treatments. In an experiment, researchers impose a treatment or condition. An experiment can be nonrandomized. For example, deliberately giving one teaching method to high-scoring students and another to low-scoring students imposes a treatment but does not randomly assign it.
Do not classify every nonrandomized study as observational. First ask whether the researchers imposed the condition; then ask whether assignment was random. Those are separate design features.
A confounder is an alternative explanation
A variable is confounding a comparison when its relationship with treatment or exposure and the outcome makes the treatment’s separate effect difficult to identify. If motivated students choose optional tutoring, a later score advantage may reflect tutoring, motivation, prior achievement, or a combination. The observed association does not isolate one cause.
This does not prove tutoring has no effect. It shows that this design cannot distinguish its causal effect cleanly. Watch answer choices that turn “not established” into “false.”
Fair comparisons need more than two labels
A comparison or control group helps reveal what might have happened without the treatment. A treatment group’s before-and-after improvement can reflect time, practice, seasonal conditions, or another change. Holding other conditions comparable reduces these alternative explanations.
Replication across enough independent units improves the information in a comparison. Measuring one plant many times is not the same as randomly assigning many plants. Blinding, when practical, can reduce differences in expectations or outcome measurement; it does not replace random assignment.
Blocking addresses a known source of variation
When growing conditions differ between greenhouses, randomize both fertilizer treatments within each green- house. Then each location contains a comparison, rather than location being completely tied to fertilizer. This is the idea of a randomized block design.[11] The name is supporting vocabulary; the key reasoning is preventing an important background factor from becoming an alternative explanation for the treatment difference.
Evaluate the entire path from population to claim
Sampling problems and measurement problems
Undercoverage occurs when the sampling process omits or poorly represents part of the target population. A cafeteria survey of people currently buying lunch may miss students who avoid the cafeteria. A voluntary-response survey can attract people with unusually strong opinions. Nonresponse can create bias when responders systematically differ from nonresponders in relevant ways, even after an initially random selection.
A leading or confusing question can change answers; this is a measurement problem, not something a larger sample automatically fixes. A good survey question is neutral, specific, and understandable. The existence of a possible bias does not reveal its direction or exact size without additional information.
Separate four questions about a conclusion
Description: What do the observed data actually show? A sample mean difference is a descriptive fact about those data.
Generalization: To whom can an estimate reasonably extend? Identify the population that the sampling frame represents, not the broadest group mentioned in an answer choice.
Causation: Was a treatment assigned in a way that supports a fair causal comparison? A random sample of self-selected behaviors does not answer this question.
Evidence strength: Is the observed difference convincing beyond chance under the reported analysis? Random assignment makes a causal design possible; it does not guarantee a nonzero causal effect in every experiment.
Average effects are not universal individual effects
Even a well-supported average treatment benefit does not imply that every participant improves, that no one is harmed, or that the same effect occurs in all settings. Keep “average,” the study population, the treatment details, and the measured outcome in the conclusion.
A precise conclusion template
“The [design] supports [an estimate / an association / a causal average comparison] for [the defensible population or participant group], provided [the relevant design and evidence conditions].” Delete any stronger claim that the study did not earn. If the difference could readily be due to chance, say the evidence does not establish an effect; do not conclude the treatments are identical.
What to practice next
Examples 7.01–7.06 separate observation, sampling, assignment, and scope. Examples 7.07–7.10 address confounding and controls. Examples 7.11–7.15 evaluate bias, better designs, and limits on conclusions.
15 worked examples
Try the question before reading the solution. The examples progress from Foundation to Challenge.
7.01. Recognize an observational study
Researchers record the usual weekly exercise time and usual nightly sleep time of 300 adults. They do not ask anyone to change either behavior. Adults who exercise more tend to report more sleep. Which statement is supported?
A. The study is an experiment because researchers collected data.
B. The study is observational and shows an association, not by itself a causal effect.
C. The study proves that exercise causes longer sleep.
D. The study proves that sleep causes more exercise.
Show worked solutionHide worked solution for example 7.01
Recognize the structure. Ask whether researchers imposed a treatment, not whether they measured variables.
Work it through. The researchers observe existing behavior without assigning an exercise program or sleep schedule. This makes the study observational. The association could reflect an effect in either direction or other variables, such as work schedules. The design alone does not isolate a causal effect.
Answer: B.
Check. No step in the description randomly assigns the behavior being compared. Merely collecting a large data set does not supply that missing intervention.
Avoid the trap. The word “study” does not mean “experiment.” Also, lack of causal evidence is not proof that no causal relationship exists.
7.02. Recognize a randomized experiment
Eighty similar seedlings are randomly assigned, 40 each, to fertilizer A or fertilizer B. They are grown under otherwise comparable conditions. A sound analysis finds convincing evidence of a difference in mean growth favoring A. What conclusion does this design support?
Show worked solutionHide worked solution for example 7.02
Recognize the structure. The researchers assign a treatment at random and compare outcomes under controlled conditions.
Work it through. Random assignment makes the treatment groups comparable in expectation with respect to other characteristics. With a well-conducted comparison and convincing statistical evidence, the difference supports a causal effect of the fertilizer treatment on mean growth for the seedlings and conditions studied. The claim concerns average growth, not a guarantee for every seedling.
Answer: Evidence that fertilizer A causes greater mean growth than B under the studied conditions.
Check. The manipulated variable is fertilizer type; the outcome is growth. The random assignment addresses alternative explanations from systematic group differences.
Avoid the trap. Do not automatically extend the result to every species, soil, or climate. Random assignment supports causation; broader population generalization is a separate design question.
7.03. Name the population that was actually sampled
A simple random sample of 200 eleventh-grade students is selected from a complete list of all eleventh-grade students in one district. All selected students respond to a survey about course preferences. To which population does this sampling design directly support generalizing the results?
A. All students in the country.
B. All high-school students in the state.
C. All eleventh-grade students in that district.
D. Only the 200 students, because every sample is smaller than a population.
Show worked solutionHide worked solution for example 7.03
Recognize the structure. Identify the complete group from which random selection occurred.
Work it through. The sampling frame contains all eleventh-grade students in the district, and the random sample comes from that frame. The design therefore supports estimating course preferences for that district’s eleventh graders. It does not include other grades or districts.
Answer: C.
Check. Every member of the stated target population was eligible for selection. Other students were outside the sampling frame.
Avoid the trap. A sample need not include everyone to support inference. But random selection from one well-defined group does not automatically represent a broader group.
7.04. Do not confuse random sampling with random assignment
Researchers randomly sample 500 households in a town and measure each household’s insulation level and energy use. Households with more insulation use less energy on average. No insulation treatment is assigned. Which conclusion is best supported, assuming complete and accurate responses?
A. The association can be generalized to the town’s households, but the study alone does not establish a causal insulation effect.
B. The random sample proves that insulation caused the lower use.
C. Neither association nor generalization is possible.
D. The study proves that energy use determines insulation.
Show worked solutionHide worked solution for example 7.04
Recognize the structure. Random selection answers who is represented. It does not decide who receives a treatment.
Work it through. Random sampling supports population-level description of the association for the town’s households. Since insulation was not assigned randomly, other factors, such as home size or occupant behavior, may differ with insulation level. Thus causation is not established by this study alone.
Answer: A.
Check. The study contains random selection but no randomized intervention. Those are two different uses of randomness.
Avoid the trap. The word “random” is not a universal permission to claim causation. Locate the exact step that was randomized.
7.05. Allow a causal comparison without overgeneralizing volunteers
One hundred students volunteer for a study. They are randomly assigned to use a new reading app or an existing practice method. The study is well conducted, and a sound analysis finds convincing evidence of higher mean improvement for the app group. Which conclusion is justified?
A. The app is guaranteed to improve every student’s score.
B. There is evidence of a causal average benefit in the studied participants, but the volunteer sample alone does not justify automatic generalization to all students.
C. No causal conclusion is possible because the students volunteered.
D. The volunteers form a simple random sample of all students.
Show worked solutionHide worked solution for example 7.05
Recognize the structure. Volunteer recruitment limits population representation; subsequent random assignment can still support an internal causal comparison.
Work it through. The students were not randomly sampled from all students, so population generalization requires caution. However, assignment to app or comparison method was random. Under the stated study quality and evidence, the treatment comparison supports a causal effect on average improvement for the study participants.
Answer: B.
Check. Recruitment and treatment assignment occur at different stages. Weakness at the sampling stage does not erase the random treatment assignment.
Avoid the trap. Do not use an all-or-nothing rule that volunteers make every conclusion invalid. Separate the causal question from the population-scope question.
7.06. Recognize when both kinds of inference are supported
A simple random sample of 160 students is selected from all tenth graders in a district. All participate. The students are then randomly assigned to one of two practice programs. Other study conditions are comparable, and a sound analysis finds convincing evidence that Program A produces greater mean improvement. What is the strongest justified conclusion?
Show worked solutionHide worked solution for example 7.06
Recognize the structure. There is random sampling from a defined population and random assignment between treatments.
Work it through. Random sampling supports extending the comparison to the district’s tenth graders. Random assignment, together with the well-conducted study and convincing result, supports a causal comparison between the programs. The conclusion should be limited to the measured outcome and comparable study conditions.
Answer: Program A has a causal average advantage over the comparison program for the district’s tenth graders under comparable conditions.
Check. The conclusion names the treatment, comparison, outcome, and sampled population. Each part corresponds to information in the design.
Avoid the trap. Neither randomness step guarantees improvement for every individual, and neither makes the sample representative of every grade or district.
7.07. Identify a plausible confounding variable
Students who choose to attend an optional tutoring program have higher average course grades than those who do not. Which additional fact would most clearly provide a possible alternative explanation for the difference?
A. The tutoring room has chairs.
B. Students who attend tutoring were already more motivated to study before the program began.
C. Grades are recorded numerically.
D. Both groups attend the same school.
Show worked solutionHide worked solution for example 7.07
Recognize the structure. A confounder is related to both the group difference and the outcome, making the treatment effect hard to separate.
Work it through. Prior motivation can influence the choice to attend tutoring and can also improve grades through other study behavior. It could therefore help explain the grade difference without that entire difference being caused by tutoring. This does not show that tutoring has no effect; it shows that the observational comparison does not isolate its effect.
Answer: B.
Check. The candidate variable predates participation and is plausibly connected to both participation and grades.
Avoid the trap. Naming any extra variable is not enough. A useful confounding explanation must connect to the treatment or exposure and to the outcome.
7.08. Recognize an experiment that was not randomized
A teacher assigns the students with the highest prior test scores to a new practice method and assigns the remaining students to the old method. The new-method group later scores higher. Which statement is most accurate?
A. This is a randomized experiment because the teacher assigned methods.
B. The new method’s causal benefit is established because it was assigned by a teacher.
C. Treatments were imposed, so this is an experiment, but the nonrandom assignment confounds method with prior achievement.
D. There can be no experiment unless a chemical treatment is used.
Show worked solutionHide worked solution for example 7.08
Recognize the structure. Treatment assignment makes a study experimental; random treatment assignment is an additional property.
Work it through. The teacher intervenes by assigning methods, so the study is an experiment. But assignment is based on prior achievement. The new-method group was already academically stronger, so its later advantage cannot be cleanly attributed to the method from this comparison alone.
Answer: C.
Check. The group difference existed before the intervention. A randomized comparison would avoid deliberately making treatment group membership depend on that difference.
Avoid the trap. Do not classify every nonrandomized study as observational. The important causal limitation here is confounded treatment assignment, not the absence of any intervention.
7.09. Explain what random assignment accomplishes
Why does random assignment strengthen a comparison between two treatments?
A. It guarantees that the treatment groups are exactly equal on every characteristic.
B. It tends to balance other characteristics across groups and reduces systematic confounding.
C. It guarantees that the treatment will work.
D. It makes a volunteer sample representative of the entire population.
Show worked solutionHide worked solution for example 7.09
Recognize the structure. Random assignment prevents a systematic rule from deciding which kinds of participants receive each treatment.
Work it through. Chance allocation tends to balance both measured and unmeasured characteristics between groups. Actual groups can still differ by chance, especially in small samples, so statistical analysis is needed to judge the observed outcome difference. Random assignment creates a fairer comparison; it does not guarantee identical groups or a successful treatment.
Answer: B.
Check. If allocation is not based on motivation, age, or prior performance, those characteristics are not deliberately tied to treatment membership.
Avoid the trap. “Balanced on average” is not “exactly balanced in every sample.” Random assignment also does not replace random sampling.
7.10. Choose a useful control comparison
A gardener gives a new fertilizer to 30 plants and observes that they grow over four weeks. Which design change would most improve the ability to assess the fertilizer’s causal effect?
A. Measure the treated plants more often, with no comparison group.
B. Use only the fastest-growing plants in the final analysis.
C. Randomly assign comparable plants to the new fertilizer or a specified control treatment and keep other conditions comparable.
D. Ask whether the gardener expected the fertilizer to work.
Show worked solutionHide worked solution for example 7.10
Recognize the structure. Growth after a treatment does not reveal how much growth would have occurred without that treatment.
Work it through. A randomized control comparison provides evidence about the counterfactual: what similar plants would do under the comparison condition. Holding light, water, species, and other conditions comparable helps isolate the fertilizer difference. Measuring treated plants alone cannot separate treatment effect from ordinary growth over time.
Answer: C.
Check. The proposed change creates groups that differ deliberately in the treatment being tested, rather than merely comparing later growth with earlier size.
Avoid the trap. A before-after change is not automatically a treatment effect. Time itself and other changes can produce improvement.
7.11. Identify undercoverage in a convenient sampling location
A school wants to estimate all students’ opinions about lunch options. It surveys the first 100 students buying lunch in the cafeteria on one day. Which concern most directly limits generalization to all students?
A. A sample must contain at least half the population.
B. Students who do not buy cafeteria lunch have no chance to enter this sample.
C. One hundred responses can never be analyzed.
D. Collecting opinions is automatically an experiment.
Show worked solutionHide worked solution for example 7.11
Recognize the structure. Compare the target population with the group that could actually be selected.
Work it through. The target is all students, but the sampling procedure covers only cafeteria purchasers present at that time. Students bringing lunch or eating elsewhere are excluded, and their opinions may differ. Taking a larger sample from the same restricted line does not necessarily fix the missing coverage.
Answer: B.
Check. The excluded students are directly relevant to the lunch question; their absence is a plausible source of systematic bias.
Avoid the trap. Sample size does not repair a sampling frame that omits part of the target population. A convenient location is not a simple random sample of all students.
7.12. Notice nonresponse after an initially random sample
A town randomly selects 500 residents for a survey. Only 80 respond. Which statement is most accurate?
A. The responses are automatically representative because the initial selection was random.
B. Nonresponse may bias the results if responders and nonresponders differ on the topic.
C. Nonresponse proves that every reported percentage is false.
D. Increasing the number of decimal places removes nonresponse bias.
Show worked solutionHide worked solution for example 7.12
Recognize the structure. An initially valid selection procedure can be weakened by who ultimately supplies data.
Work it through. The 80 respondents may not preserve the representativeness of the initial 500 if response is related to the opinion or behavior being measured. For example, people with strong views may respond more often. Follow-up efforts and careful analysis can address this concern, but initial random selection alone does not eliminate it.
Answer: B.
Check. A different opinion distribution among the 420 nonrespondents could shift the population estimate substantially.
Avoid the trap. The concern is possible bias, not proof that the estimate is wrong. Uncertainty about nonresponse should weaken confidence in generalization without inventing the direction of the bias.
7.13. Distinguish question wording from sampling quality
A school randomly samples students but asks, “Do you support the excellent new schedule that will improve learning?” Which revision most directly reduces a potential source of response bias?
A. Ask, “Do you support or oppose the proposed new schedule?”
B. Ask only students who already support the schedule.
C. Keep the question but collect more responses.
D. Replace percentages with counts in the report.
Show worked solutionHide worked solution for example 7.13
Recognize the structure. Random sampling cannot neutralize a question that pushes respondents toward a preferred answer.
Work it through. The adjectives and asserted benefit in the original wording encourage support. A neutral question asks about the proposal without presenting its quality or effect as settled. This addresses measurement or response bias at the question stage, which is different from selection bias at the sampling stage.
Answer: A.
Check. The revision preserves the intended topic while removing language that suggests the desired response.
Avoid the trap. More responses to a leading question can produce a more precise summary of biased answers. Improve the measurement, not merely the sample size.
7.14. Prevent treatment from being confounded with location
A researcher has 24 plants in each of two greenhouses. The greenhouses differ in sunlight. To compare two fertilizers, which design most effectively prevents fertilizer type from being confounded with greenhouse?
A. Use fertilizer A on all plants in greenhouse 1 and B on all plants in greenhouse 2.
B. Within each greenhouse, randomly assign 12 plants to A and 12 to B.
C. Give A to the taller plants and B to the shorter plants.
D. Use A this year and B next year without another comparison.
Show worked solutionHide worked solution for example 7.14
Recognize the structure. Each greenhouse should contain both treatments so location does not determine treatment.
Work it through. Random assignment within each greenhouse balances fertilizer use across the two sunlight conditions. Comparing the fertilizers within greenhouse helps separate treatment from location. This is an example of blocking: grouping by a relevant characteristic before randomizing within groups. The term is less important than the design logic.
Answer: B.
Check. Each greenhouse has 12 plants under each fertilizer, and each fertilizer is used on 24 plants overall.
Avoid the trap. Randomly deciding which greenhouse gets A would not create a strong replicated comparison with only one greenhouse per treatment. Individual plants within one greenhouse do not remove the shared-location difference.
7.15. Audit the design, result, and scope separately
A district randomly samples 120 ninth graders, all of whom participate. It randomly assigns 60 to a new practice program and 60 to an existing program under comparable conditions. The new program’s mean improvement is 0.3 points higher. The analysis finds that a difference this small could reasonably arise from random assignment alone. Which conclusion is most appropriate?
A. Random assignment proves that the new program is better.
B. The programs are known to have exactly equal effects.
C. The design supports a population-relevant causal comparison, but this result does not provide convincing evidence that the new program is better.
D. Because only 120 students participated, no population inference is possible.
Show worked solutionHide worked solution for example 7.15
Recognize the structure. A design can be capable of supporting a causal conclusion even when the observed result is not convincing evidence of a benefit.
Work it through. Random sampling supports the scope of district ninth graders; random assignment supports a causal comparison. But chance can create small mean differences even without a treatment advantage. The stated analysis says the 0.3-point difference is not convincing evidence of superiority. It also does not establish that the true effects are identical.
Answer: C.
Check. This conclusion keeps three questions separate: who the study represents, whether the treatment comparison is fair, and whether the measured difference is strong enough evidence.
Avoid the trap. Do not treat random assignment as a guarantee of a positive effect. Do not turn “not enough evidence of a difference” into a proof of no difference.
Another route. A useful conclusion template is: “For [population], this design permits [type of comparison], but the observed evidence [does or does not] support [specific claim].” Fill every slot before accepting a broad statement.
Section 7: exit questions
Try the five section-exit questions before checking the explanations. Use scratch paper for your reasoning and enter the requested numerical value for a question without choices.