Command of Evidence — Quantitative

Use tables and graphs to support the exact claim the question asks you to evaluate.

By the end, you should be able to:

4.1 Read the display before interpreting it

Quantitative evidence questions ask you to connect a claim with numerical evidence. Before comparing values, identify what was measured, the units, the groups, and the time period. A table cell might report total visitors, visitors per event, or the percentage of visitors who returned. Those quantities answer different questions.

For a bar chart, read the category labels, vertical-axis units, and legend before comparing heights. For a line graph, identify what each point represents and follow one labeled series at a time. A rising line means the measured quantity rose; it does not automatically mean conditions improved. On tables, trace the row and column into their intersection. Then translate one value into a complete sentence.

  • Check whether values are counts, averages, percentages, or changes.
  • Read notes for denominators, sample definitions, and measurement periods.
  • Use the axis scale and labels; visual steepness alone does not measure the size of an effect.

W4.1: Attendance per event

A library compared attendance at hypothetical afternoon and evening programs. Each cell reports the mean number of attendees per event, calculated across six events of that type. A planner claims that evening scheduling was associated with higher attendance for poetry readings, but not for storytelling sessions.

Mean attendance by program

Program

Afternoon

Evening

Poetry

18

30

Storytelling

24

24

Mean attendees per event; six events in each cell.

Which statement most directly supports the planner's claim?

A. Evening poetry readings averaged 30 attendees, exceeding the 24 attendees at evening storytelling sessions and 18 at afternoon poetry readings.

B. Poetry averaged 18 attendees in afternoons and 30 in evenings; storytelling averaged 24 at both times.

C. The library held 30 evening poetry readings and 24 evening storytelling sessions.

D. Afternoon storytelling sessions averaged six more attendees than afternoon poetry readings.

Practice this lesson: 4.01, 4.05.

Show worked solutionHide worked solution for W4.1

Answer: B

  1. The claim compares afternoon with evening within each program type.
  2. Poetry rises from 18 to 30 attendees per event; storytelling stays at 24.
  3. B supplies both comparisons and preserves the per-event unit.

Why A fails: These values show poetry's higher evening attendance but omit the afternoon storytelling value needed for the other half of the claim.

Why C fails: The values count mean attendees, not the number of events.

Why D fails: The afternoon comparison is true but cannot establish what changed with evening scheduling.

Key takeaway: Name the unit before you use a number as evidence.

4.2 Translate the claim into the comparison it needs

Before looking for an attractive number, identify the relationship the claim requires. X is larger than

Y requires a comparison of levels. X increased requires two observations of X. An intervention

helped X more than Y requires two changes, followed by a comparison of those changes.

Write a short prediction in words: The increase for unfamiliar works should exceed the increase for familiar works. This separates evidence that addresses the claim from facts that merely sound favorable. The highest final score may belong to the group that benefited less. When a choice offers several correct numbers, ask whether those numbers establish the requested relationship.

W4.2: Higher score or larger improvement?

In a hypothetical experiment, visitors were randomly assigned to view paintings with or without a short recorded introduction. Researchers predicted that the introduction would improve recall more for unfamiliar paintings than for familiar paintings.

Recall with and without an introduction

Mean recall score (%). Categories: Unfamiliar, Familiar. No introduction: 40, 60. Introduction: 60, 70.No introductionIntroduction020406080100Mean recall score (%)Unfamiliar4060Familiar6070

Scores measure the percentage of tested details recalled.

Which statement most directly supports the prediction?

A. With an introduction, recall was higher for familiar paintings than for unfamiliar paintings.

B. Without an introduction, familiar-painting recall exceeded unfamiliar-painting recall by 20 points, 60% compared with 40%.

C. An introduction raised recall for familiar paintings from 60% to 70%.

D. Recall gains were 20 points for unfamiliar paintings but 10 points for familiar paintings.

Practice this lesson: 4.03, 4.07.

Show worked solutionHide worked solution for W4.2

Answer: D

  1. The prediction concerns the size of improvement, not the highest final score.
  2. Unfamiliar works gain 60 - 40 = 20 points; familiar works gain 70 - 60 = 10.
  3. D compares the two improvements directly.

Why A fails: The final-score comparison is true but does not measure either group's improvement.

Why B fails: The baseline difference is true but says nothing about the introduction's effect.

Why C fails: The familiar group improved, but this alone cannot establish which group benefited more.

Key takeaway: A claim about a larger benefit requires a comparison of changes.

4.3 Track multiple groups without switching comparisons

With multiple series, distinguish within-group change from a between-group difference. Follow each group across time first. Then compare the groups at matching times. If the claim concerns convergence, calculate or estimate the gap at the beginning and end.

Two groups can both improve while their gap widens, narrows, or stays constant. A group can also close a gap without catching up. Preserve the claim's exact scope: the gap narrowed does not mean the groups became equal. Do not compare one group at the beginning with a different group at the end unless the question specifically calls for that comparison.

W4.3: Closing a gap without erasing it

Two hypothetical student ensembles practiced the same timing exercise over three weeks. A teacher argues that both ensembles improved, while Ensemble R's initial advantage over Ensemble S became smaller.

Timing accuracy during practice

Correctly timed notes (%). Week: 1, 2, 3. Ensemble R: 70, 75, 80. Ensemble S: 50, 60, 70.Ensemble REnsemble S0102030405060708090100Correctly timed notes (%)Week12375

Each ensemble completed the same-length timing exercise weekly.

Which statement best supports the teacher's argument?

A. Both improved; R's lead shrank from 20 points in week 1 to 10 in week 3.

B. R scored 80% in week 3, the highest score shown.

C. S reached 70% in week 3, matching R's week 1 score and exceeding S's score at every earlier observation.

D. R and S both improved by 10 percentage points between weeks 1 and 3.

Practice this lesson: 4.08, 4.14.

Show worked solutionHide worked solution for W4.3

Answer: A

  1. Follow each ensemble from week 1 to week 3: both scores rise.
  2. Compare matching weeks: the gap is initially 20 points and finally 10.
  3. A captures improvement and convergence without claiming equality.

Why B fails: A high final score does not establish both groups' improvement or a shrinking gap.

Why C fails: This cross-time comparison is true but does not measure the gap at a common time.

Why D fails: R improves by 10 points, but S improves by 20.

Key takeaway: Compare groups at the same time when evaluating a changing gap.

4.4 Distinguish counts from rates and proportions

A count answers How many? A rate or proportion answers How many relative to what? A large sample can produce more successes even when its success rate is lower. Before comparing a percentage, name its denominator: all applicants, admitted applicants, surveyed residents, or something else.

To compare litter density across differently sized areas, divide litter count by surveyed area. To compare success rates, divide successes by attempts. You usually need only simple arithmetic. The more important reading decision is choosing the denominator that matches the claim. Do not replace a measured rate with a total, or assume equal denominators when none are stated.

  • Percentage = relevant part / relevant whole x 100.
  • Rate per 100 units = count / units surveyed x 100.
  • A percentage of a group does not tell you the group's size.

W4.4: Most litter, lowest density

Volunteers surveyed three hypothetical beach zones. A coordinator claims that the South zone contained the most litter items in total but had the lowest litter density per 100 meters surveyed.

Litter survey

Zone

Meters surveyed

Litter items

North

200

40

Central

100

30

South

400

60

Each litter item was counted once within the surveyed length.

Which statement best supports the claim?

A. South had 60 items, compared with 40 in North and 30 in Central.

B. Central had 30 items along 100 meters, while North had 40 along 200 meters.

C. South had the highest count; its density was 15 items per 100 meters, compared with North's 20 and Central's 30.

D. South had twice as many litter items as Central and therefore twice Central's litter density.

Practice this lesson: 4.06, 4.09.

Show worked solutionHide worked solution for W4.4

Answer: C

  1. The count comparison gives South the highest total: 60 items.
  2. Per 100 meters, North has 20 items, Central 30, and South 15.
  3. C supports both the total-count and density parts of the claim.

Why A fails: It establishes the largest count but leaves out the required density comparison.

Why B fails: These figures are accurate but provide no information about South.

Why D fails: South's surveyed length is four times Central's, so twice the count does not mean twice the density.

Key takeaway: Choose the denominator required by the claim, then compare like units.

4.5 Separate absolute change, relative change, and trend

Absolute change is the ending value minus the starting value. Relative change compares that difference with the starting value. A gain of 40 members doubles a club that began with 40, but increases a club that began with 100 by only 40%.

When the values are already percentages, subtraction gives percentage points. A rise from 20% to 30% is 10 percentage points, or a 50% relative increase. Read words such as more, faster, and larger increase carefully to identify the intended quantity. Also check intermediate observations: equal endpoints do not prove a steady trend, and higher cumulative totals do not necessarily mean faster current output.

W4.5: Same gain, different growth rate

Two hypothetical film clubs tracked membership over three years. A student argues that the smaller club grew more in proportional terms, even though both clubs added the same number of members.

Film club membership

Members. Year: 1, 2, 3. Smaller club: 40, 60, 80. Larger club: 100, 120, 140.Smaller clubLarger club020406080100120140160MembersYear123

Membership totals at the end of each year.

Which statement best supports the student's argument?

A. The larger club remained ahead throughout, with membership increasing from 100 in the first year to 140 in the third.

B. The smaller club added 40 members, exceeding the larger club's gain.

C. Both gained 40 members: a 100% increase for the smaller club and 40% for the larger.

D. Both clubs gained 20 members between years 2 and 3.

Practice this lesson: 4.02, 4.13, 4.16.

Show worked solutionHide worked solution for W4.5

Answer: C

  1. Each club gains 40 members over the full interval.
  2. The smaller club gains 40/40 = 100%; the larger gains 40/100 = 40%.
  3. C distinguishes equal absolute gains from unequal proportional gains.

Why A fails: The comparison of membership levels is true but does not compare growth rates.

Why B fails: The larger club also gains 40 members.

Why D fails: This equal one-year gain is true but omits starting sizes and proportional growth.

Key takeaway: An equal numerical increase represents greater proportional growth from a smaller starting value.

4.6 Connect the text's explanation to the graphic

Sometimes the passage supplies a mechanism or prediction, and the display supplies the observations. Translate the explanation into an expected pattern before reading the choices. If shade helps seeds mainly by reducing drying, its benefit should be greater under dry conditions than under wet conditions.

Look for the comparison that distinguishes the proposed explanation from alternatives. Showing that shaded seeds germinated is weaker than showing how shade's benefit changes with water conditions. A choice can quote the graph perfectly and still fail to test the passage's explanation. Prefer the relevant pattern over the most impressive isolated value.

W4.6: Testing a proposed mechanism

In a hypothetical controlled trial, identical seed batches were randomly assigned to combinations of shade and watering conditions. A researcher proposes that shade improves germination mainly by reducing drying. If so, shade should provide a larger benefit under dry conditions than under wet conditions.

Germination by shade and water

Seeds germinated (%). Categories: Dry, Wet. No shade: 20, 70. Shade: 60, 80.No shadeShade020406080100Seeds germinated (%)Dry2060Wet7080

Equal-size batches; all other growing conditions held constant.

Which finding most directly supports the proposed explanation?

A. Shade increased germination by 40 percentage points under dry conditions but by only 10 under wet conditions.

B. Under shade, germination was 80% with wet conditions and 60% with dry conditions.

C. Without shade, germination was higher under wet conditions than under dry conditions.

D. The highest germination rate, 80%, occurred with both shade and wet conditions.

Practice this lesson: 4.15, 4.18.

Show worked solutionHide worked solution for W4.6

Answer: A

  1. The mechanism predicts that shade's benefit depends on drying risk.
  2. Shade adds 60 - 20 = 40 points when dry, but 80 - 70 = 10 when wet.
  3. A compares the effects needed to evaluate that prediction.

Why B fails: This compares water conditions within shade, not shade's benefit under each water condition.

Why C fails: This establishes a water-related difference but does not test how shade changes either result.

Why D fails: The highest endpoint alone does not reveal why shade helps.

Key takeaway: Use the passage to decide which comparison tests the proposed explanation.

4.7 Judge what the evidence can establish

Numerical results can support a pattern, a prediction, or a causal explanation with different degrees of strength. Check how the evidence was obtained. Random assignment and comparable conditions can strengthen a causal interpretation. An observational comparison can support an association and sometimes contribute to a causal argument, but self-selection or another difference may also explain it.

Do not reject useful evidence merely because it is not conclusive. Instead, match the conclusion's strength to the design. Also preserve the measured outcome: higher completion does not automatically mean greater enjoyment, and a group average does not describe every individual. Unless uncertainty information is supplied, do not label a small difference statistically significant.

W4.7: Association and self-selection

At a hypothetical cycling event, riders chose whether to use a new route-planning app. Organizers recorded route completion but did not measure prior fitness or navigation experience. They want to summarize what the results establish.

Route completion

Rider group

Riders

Completed route

Used app

100

75

Did not use app

100

50

Riders chose whether to use the app; prior fitness was not measured.

Which conclusion is best supported?

A. Using the app caused an increase of 25 percentage points in completion.

B. App users had greater prior fitness than riders who did not use the app.

C. Every app user was more likely to finish than every nonuser, regardless of prior fitness or navigation experience.

D. App use was associated with higher completion, although unmeasured differences between riders could help explain the pattern.

Practice this lesson: 4.04.

Show worked solutionHide worked solution for W4.7

Answer: D

  1. Completion is 75/100 = 75% among app users and 50/100 = 50% among other riders.
  2. Riders selected their own group, and potentially relevant differences were not measured.
  3. D reports the observed association while acknowledging the design's limit.

Why A fails: Self-selection prevents attributing the entire difference to the app from these data alone.

Why B fails: Fitness was not measured, so it is a possible explanation, not an established finding.

Why C fails: Group completion rates do not establish an individual-by-individual ranking of finishing probabilities.

Key takeaway: Distinguish an observed difference from an explanation of what caused it.

4.8 Check every condition in a complex claim

Hard quantitative questions may combine requirements: a gap shrinks while both groups improve, or a gain exceeds another gain without reducing a baseline outcome. Split the claim into separate checks. Every required condition must hold.

Distinguish false data from true but insufficient evidence. A fact about one group may not establish a two-group comparison. Also avoid inventing requirements: if the claim asks for a larger increase, the answer need not identify the highest final score.

  • Name the baseline or comparison group.
  • Test each required condition using the same units.
  • Reject an option that establishes only part of the requested relationship.
  • For weaken questions, identify the predicted pattern and look for a relevant contradiction or alternative explanation.

W4.8: Two requirements, one qualifying plan

A hypothetical museum compared two quiet-hour plans with its usual schedule. It considers a plan suitable if the increase in tour completion is larger for first-time visitors than for returning visitors, and returning visitors' completion does not decline.

Tour completion by schedule

Visitors completing tour (%). Categories: Usual, Plan P, Plan Q. First-time: 20, 35, 40. Returning: 40, 50, 65.First-timeReturning01020304050607080Visitors completing tour (%)Usual2040Plan P3550Plan Q4065

Rates are calculated separately within each visitor group.

Which statement identifies the plan that satisfies both requirements?

A. Q qualifies because its first-time visitors reached 40%, exceeding P's 35%.

B. P qualifies: first-time completion rose 15 points and returning completion rose 10; Q's increases were 20 and 25 points.

C. Both qualify because both visitor groups improved under each plan.

D. Neither qualifies because returning visitors had higher tour completion than first-time visitors under the usual schedule and under both plans.

Practice this lesson: 4.10, 4.20.

Show worked solutionHide worked solution for W4.8

Answer: B

  1. Compare each plan with the usual schedule, separately for each visitor group.
  2. P gives gains of 15 and 10 points; Q gives gains of 20 and 25.
  3. Both avoid declines, but only P gives first-time visitors the larger increase.

Why A fails: Q has a higher first-time endpoint, but its gain is larger for returning visitors.

Why C fails: Both satisfy the no-decline requirement, but Q fails the larger-first-time-increase requirement.

Why D fails: The criteria concern changes, not which group has the higher final rate.

Key takeaway: For a two-part claim, evidence must establish both parts using the correct baseline.

Independent practice

Keep these checks in mind

  • What quantity and population does each value represent?
  • Does the claim require levels, changes, rates, or changes in a gap?
  • Does the evidence establish every part of the requested relationship?
  • Does the conclusion stay within what was measured and how the data were collected?

Begin with four brief reasoning drills, then complete sixteen multiple-choice questions. Later items generally demand closer distinctions. Identify the evidence for each answer and the specific flaw in its nearest competitor.

These four written drills are ungraded. Write your reasoning before opening each model answer.

Drill 4.01

Two hypothetical museums surveyed visitors about whether they had visited before.

Surveyed museum visitors

Museum

Surveyed

Returning

X

100

20

Y

200

30

Returning visitors are a subset of all surveyed visitors.

Which museum had more returning visitors in its survey? Which had the higher percentage of returning visitors? Give both percentages.

Write a brief answer and identify the words that justify it.

Show model answerHide model answer for drill 4.01

Identify the denominator

Answer: Y had more returning visitors: 30 versus 20. X had the higher percentage: 20%, compared with Y's 15%.

For X, 20/100 = 20%. For Y, 30/200 = 15%. Count and proportion give different rankings because the sample sizes differ.

If you missed it: Write the denominator beside each numerator before comparing percentages.

Review lesson: 4.1 Read the display before interpreting it

Drill 4.02

A hypothetical library tested two versions of an event notice. Equal-size reader groups saw different versions.

Registration after reading a notice

Readers registering (%). Categories: Original, Revised. Readers registering (%): 40, 60.Readers registering (%)020406080100Readers registering (%)Original40Revised60

Each version was shown to a different group of 100 readers.

A student says, "The revised notice produced a 20% relative increase in registration." Correct the statement using both percentage points and relative change.

Write a brief answer and identify the words that justify it.

Show model answerHide model answer for drill 4.02

Percentage points versus relative change

Answer: The registration rate was 20 percentage points higher for the revised notice: 60% rather than 40%. Relative to the original rate, the revised rate was 50% higher.

The point difference is 60 - 40 = 20. Relative change uses the original 40%: 20/40 = 50%.

If you missed it: For values already written as percentages, label subtraction as percentage points before calculating relative change.

Review lesson: 4.5 Separate absolute change, relative change, and trend

Drill 4.03

A researcher predicts that planted buffer strips reduce sediment runoff more on steep slopes than on nearly flat slopes.

Without inventing numbers, name the comparisons needed to evaluate this prediction.

Write a brief answer and identify the words that justify it.

Show model answerHide model answer for drill 4.03

Translate a mechanism into a comparison

Answer: Compare buffered and unbuffered runoff on steep slopes; do the same on flat slopes; then compare the two reductions.

The prediction concerns a difference between effects. Comparing only the two buffered plots could confuse an initial slope-related difference with the buffer's benefit.

If you missed it: Underline "more" and identify the two changes that must be compared.

Review lesson: 4.2 Translate the claim into the comparison it needs

Drill 4.04

Households that chose a water-saving app used less water than households that did not. Their water use before installing the app was not recorded.

Give one possible alternative explanation and one additional measurement that would help evaluate it.

Write a brief answer and identify the words that justify it.

Show model answerHide model answer for drill 4.04

Separate a possible explanation from a finding

Answer: App users may already have used less water. Measuring both groups' water use before installation would help evaluate that possibility.

This is an untested explanation, not an established fact. Baseline data would clarify change, although other group differences could still matter.

If you missed it: Distinguish "could explain" from "the study found."

Review lesson: 4.7 Judge what the evidence can establish

Practice: 16 multiple-choice questions