One-variable data: distributions, center, and spread

Start here

Describe the data without inventing information.

The central idea: a display reveals some things and hides others

A one-variable data set records one characteristic for each observational unit. Numerical measurements support arithmetic such as means; category labels do not automatically do so. Before computing, identify the variable, its units, the number of observations, and whether the display shows exact values or grouped intervals.

Match each display to the information it preserves

Display

Read it this way

Frequency table

Each value is repeated its stated number of times. Add frequencies for the data count.

Dot plot

Each dot represents the stated number of observations at that horizontal value. Count stacks carefully.

Histogram

Heights give bin frequencies in this guide; a bin groups a range of values, not one exact value.

Box plot

Read the median and quartiles from the box; read the stated whisker convention before treating endpoints as extrema.

On a histogram with unequal bin widths, do not assume raw height alone equals frequency unless the axis says so. The examples here use equal-width frequency bins. A cumulative count can identify a median’s bin without determining the exact median. Midpoint-based calculations from bins are estimates unless the problem explicitly supplies additional information.

Mean, median, and mode answer different questions

The arithmetic mean is total divided by count: x¯ = ∑ x/n. To find a median, sort the observations. For odd n, use position (n + 1)/2. For even n, average the values at positions n/2 and n/2 + 1. A median can lie between two observed values. The mode is a most frequent value; it need not be unique.

For a frequency table, multiply each value by its frequency before summing:

x¯ = ∑(value × frequency)∑ frequency.

The denominator is the number of observations, not the number of distinct values. To combine groups, first recover each group’s total using nx¯, then divide the sum of those totals by the combined count.

Spread describes variation among the observations

Range and interquartile range

The range is maximum minus minimum. The interquartile range is Q3 − Q1, the width of the middle half of the data as summarized by the quartiles. Neither measure is the median. A narrow box can coexist with long whiskers: concentrated central values do not guarantee a small overall range.

Different conventions can produce slightly different quartiles for small raw data sets. In this guide, box-plot questions provide the quartiles directly. Also, some box plots show outliers separately and extend whiskers only to the most extreme non-outliers. Read the stated convention instead of assuming every whisker ends at an absolute minimum or maximum.

Standard deviation measures distances from the mean

A standard deviation is a nonnegative measure of spread in the same units as the observations. It is zero precisely when all observations are equal. Qualitatively, a distribution with more observations farther from its mean has more spread, provided the comparison holds the relevant data counts and structure fixed.

For explanation, the population standard deviation is

σ = ∑(x − x¯ )2n.

The sample convention uses n−1 instead of n. The SAT scope emphasizes interpreting and comparing standard deviations; you do not need to memorize a lengthy hand-calculation procedure here. In comparisons with equal data counts, comparing sums of squared deviations works under either consistent convention. Equal means and ranges do not force equal standard deviations.

Outliers and resistant summaries

An extreme observation can pull the mean substantially toward it. The median depends on order and is often less sensitive, but is not guaranteed to remain unchanged when an observation is removed or added. Range is controlled entirely by the extremes; standard deviation is sensitive to large deviations. A box plot alone generally does not determine an exact mean or standard deviation.

Do not ask a graph for information it does not contain

A histogram may identify an interval, not an exact median. A box plot may compare IQRs, not exact means. A table of group means needs group counts before you can find the combined mean. “Cannot be determined” is justified by missing information, not by difficult arithmetic.

Changing data without recomputing everything

Translate all observations

Adding a constant c to every observation adds c to the mean, median, and each quartile. All distances between observations stay unchanged, so range, IQR, and standard deviation stay unchanged. Subtracting a constant follows the same principle. A shifted location is not a changed spread.

Scale and shift all observations

For a transformation y = ax + b,

y¯ = ax¯ + b, SD(y) = |a| SD(x).

The absolute value is essential: multiplying by a negative number reverses order but cannot make spread negative. The median follows the same affine rule; range and IQR multiply by |a|. These results apply under either consistent standard-deviation convention. For a = 0, every transformed observation equals b and spread is zero.

Add, remove, or replace individual observations

Use totals for mean problems. If n values average m, their total is nm. After adding z, the mean is (nm+z)/(n+1). Replacing u with v changes the total by v − u while keeping the count fixed. Removing a value changes both the total and the count.

An observation equal to the old mean leaves the mean unchanged when added. For a nonconstant data set it reduces the population standard deviation, since it contributes zero squared deviation while increasing the denominator. For median questions, re-check positions; a mean shortcut does not establish what happens to the median.

Unknown groups and constraints

A combined mean can determine an unknown group size. Set total = mean × count for each group and solve the resulting equation. To maximize or minimize a mean under constraints, maximize or minimize the total while respecting every stated condition, including ordered positions, permitted values, and whether repetitions are allowed.

Data spread is not sampling uncertainty

Standard deviation describes variation among individual values. Margin of error describes uncertainty in an estimate of a population quantity. A large sample can estimate a mean precisely even when the individual measurements vary widely. Sections 6 and 7 explain why design and uncertainty belong to a different layer of reasoning.

What to practice next

Examples 3.01–3.06 focus on displays and summaries; 3.07–3.11 on totals, outliers, and transformations; and 3.12–3.15 on deeper comparisons and constraints.

15 worked examples

Try the question before reading the solution. The examples progress from Foundation to Challenge.

3.01. Read a mean as an equal‐share total

The numbers of books read by five students are 4, 6, 6, 8, and 11. What is the mean number of books read?

Show worked solutionHide worked solution for example 3.01

Recognize the structure. The mean is the total divided by the number of observations; repeated values still count separately.

Work it through. The total is 4 + 6 + 6 + 8 + 11 = 35 books. Since there are five students,

x̄ = 355 = 7.

The mean describes an equal-share value, even though no student in this data set read exactly 7 books.

Answer: 7 books.

Check. Five students at a mean of 7 account for 5(7) = 35 books, matching the total.

Avoid the trap. There are five observations, not four distinct values. Counting the repeated 6 only once changes the data set.

3.02. Sort before locating an even‐sized median

A data set consists of 9, 3, 12, 4, 7, and 7. What is its median?

Show worked solutionHide worked solution for example 3.02

Recognize the structure. The median is based on ordered positions, not on the middle entries in the original display.

Work it through. Sort the observations: 3, 4, 7, 7, 9, 12. There are six observations, so the two middle positions are 3 and 4. Their average is

7 + 72 = 7.

An even-sized data set requires the average of its two central observations.

Answer: 7.

Check. At least half the observations are at or below 7, and at least half are at or above 7.

Avoid the trap. Averaging 12 and 4 because they happen to occupy the middle of the unsorted list is not a median calculation.

3.03. Use frequencies as weights

The table summarizes the number of pets owned by 10 households. What is the mean number of pets per household?

Number of pets

0

1

2

3

Number of households

2

4

3

1

Show worked solutionHide worked solution for example 3.03

Recognize the structure. The top row gives values; the bottom row tells how many times each value occurs.

Work it through. Multiply each value by its frequency to recover the total number of pets:

0(2) + 1(4) + 2(3) + 3(1) = 13.

The number of households is 2 + 4 + 3 + 1 = 10, so the mean is 13/10 = 1.3 pets per household.

Answer: 1.3 pets per household.

Check. The expanded list is 0, 0, 1, 1, 1, 1, 2, 2, 2, 3. Its ten values sum to 13.

Avoid the trap. The mean of the four category labels, (0+1+2+3)/4 = 1.5, ignores their unequal frequencies. A noninteger mean is valid for count data.

3.04. Read the distribution in a dot plot

The dot plot shows the number of goals scored by a team in 10 games. What is the median number of goals?

A dot plot of goals per game: 2 has one dot, 3 has two, 4 has four, 5 has two, and 6 has one. Each dot represents one game.23456GoalspergameEachdotrepresentsonegame.
Show worked solutionHide worked solution for example 3.04

Recognize the structure. Each dot is one game. Count observations cumulatively from the smallest value.

Work it through. There is one game with 2 goals, two with 3, four with 4, two with 5, and one with 6. Positions 1–3 are below 4; positions 4–7 equal 4. Therefore both middle observations, positions 5 and 6, are 4, so the median is 4.

Answer: 4 goals.

Check. The plot is symmetric around 4, so its mean is also 4. Symmetry supports the result, although locating the middle positions is the decisive median calculation.

Avoid the trap. The tallest stack indicates the mode, which need not equal the median in another distribution. Do not assume the two measures are always the same.

3.05. Know what a histogram can and cannot reveal

A histogram summarizes 20 delivery times. Each interval includes its left endpoint but excludes its right endpoint. Which interval contains the median delivery time?

A histogram of time in minutes: 0–10 has frequency 3, 10–20 has 5, 20–30 has 8, and 30–40 has 4.010203040Deliverytime(minutes)02468Numberofdeliveries3584

A. 0 ≤ t < 10 B. 10 ≤ t < 20

C. 20 ≤ t < 30 D. 30 ≤ t < 40

Show worked solutionHide worked solution for example 3.05

Recognize the structure. A histogram gives counts within intervals, not the exact individual times.

Work it through. The bin counts are 3, 5, 8, and 4. The first two bins contain 3 + 5 = 8 times. The third bin therefore contains positions 9 through 16. With 20 times, the median averages positions 10 and 11; both lie in 20 ≤ t < 30. Their average lies there as well.

Answer: C: 20 ≤ t < 30.

Check. Eight observations are below 20, while sixteen are below 30, so the central pair must be in the third bin.

Avoid the trap. The exact median is not necessarily the bin midpoint 25. That would require information about the locations within the bin that the histogram does not supply.

Exact bounds from grouped data

A histogram fixes how many observations lie in each interval. It does not fix their individual values. For integer endpoints L and U, an integer in [L, U) may equal L but cannot equal U; its largest possible value is U − 1. The notation [L, U) means L ≤ x < U. If the endpoints are not integers, first identify the smallest and largest feasible integers; do not subtract 1 mechanically from a noninteger endpoint.

To bound a sum, multiply each bin’s frequency by its feasible lower or upper endpoint and add. Then divide by the total number of observations. These bounds are attained when every observation takes its chosen endpoint, provided there are no extra constraints such as distinctness or a specified median. Repeated values are allowed unless the question says otherwise.

Smallest sum = sum of (frequency × smallest feasible value).

Greatest sum = sum of (frequency × largest feasible value).

For independent data sets, to minimize mean A minus mean B, minimize A and maximize B. Divide each sum by its own count before subtracting. If the question asks for the absolute difference, first check whether the possible means overlap; minimizing a signed difference alone may not answer that question. Similarly shaped histograms do not prove that corresponding observations moved by one fixed amount. Bin midpoints yield an estimate under an extra assumption, not exact extrema.

Keep the domain visible. For unrestricted real values in [L, U), values can approach U arbitrarily closely. There is no largest value, so the integer rule U − 1 does not apply. Also, cumulative frequencies usually locate a median’s bin without determining its exact value.

Must versus could: If a new value is below every old value, it is below the old mean. The new mean is a weighted average of that smaller value and the old mean, so it must decrease. The median need not strictly decrease: adding 0 to 3, 7, 7, 7 leaves the median at 7. Adding 0 to 3, 5, 7, 9 changes the median from 6 to 5. A counterexample disproves “must”; one example of a decrease proves only “could.”

Worked example 1: Bound a mean from feasible endpoints

A data set contains six integers, summarized below. Find its smallest and greatest possible means.

Six integer observations[0, 4) has frequency 2; [4, 8) has frequency 3; [8, 12) has frequency 1. Each interval includes its lower endpoint and excludes its upper endpoint.Six integer observationsFrequency23104812Value
Six integer observations: [0, 4) has frequency 2; [4, 8) has frequency 3; [8, 12) has frequency 1. Each interval includes its lower endpoint and excludes its upper endpoint.
Show worked solutionHide worked solution: Bound a mean from feasible endpoints

The smallest values in the three bins are 0, 4, and 8. The largest integer values are 3, 7, and 11, not 4, 8, and 12.

Minimum sum = 2(0) + 3(4) + 1(8) = 20.

Maximum sum = 2(3) + 3(7) + 1(11) = 38.

Minimum mean = 206 = 103.

Maximum mean = 386 = 193.

Attainability is part of the answer: 0, 0, 4, 4, 4, 8 achieves the minimum; 3, 3, 7, 7, 7, 11 achieves the maximum. Both lists have exactly the stated frequencies. Midpoints 2, 6, and 10 would produce neither extreme.

Practice exact bounds from grouped data

Worked example 2: Compare independent sets with unequal counts

The two histograms summarize independent integer data sets. What is the smallest possible value of mean A minus mean B?

Data set A[8, 12) has frequency 2; [12, 16) has frequency 3. Each interval includes its lower endpoint and excludes its upper endpoint.Data set AFrequency2381216Value
Data set A: [8, 12) has frequency 2; [12, 16) has frequency 3. Each interval includes its lower endpoint and excludes its upper endpoint.
Data set B[0, 4) has frequency 1; [4, 8) has frequency 3. Each interval includes its lower endpoint and excludes its upper endpoint.Data set BFrequency13048Value
Data set B: [0, 4) has frequency 1; [4, 8) has frequency 3. Each interval includes its lower endpoint and excludes its upper endpoint.
Show worked solutionHide worked solution: Compare independent sets with unequal counts

To make the requested difference small, place A at its lower endpoints and B at its largest feasible integers. The counts differ, so calculate each mean separately.

Minimum mean A = 2(8) + 3(12)5 = 10.4.

Maximum mean B = 1(3) + 3(7)4 = 6.

Minimum difference = 10.4 − 6 = 4.4.

The lists A: 8, 8, 12, 12, 12 and B: 3, 7, 7, 7 achieve these means simultaneously because the data sets are independent. The gap stays positive even at these extremes, so the smallest absolute difference is also 4.4 here. Nothing establishes a fixed observation-by-observation shift.

Practice exact bounds from grouped data

Practice: 6 questions

3.06. Compare spread without confusing the middle and the extremes

The box plots summarize two data sets. The whiskers mark each minimum and maximum, and the box edges mark the first and third quartiles. Which statement is supported?

Boxplots on the same scale. A: minimum 2, first quartile 5, median 8, third quartile 11, maximum 16. B: minimum 1, first quartile 7, median 8, third quartile 9, maximum 18.02468101214161820ValueBA2581116178918Whiskersshowtheminimumandmaximum.

A. A has a larger median than B.

B. A and B have equal ranges.

C. A and B have equal medians, but A has the larger interquartile range.

D. B has the larger interquartile range.

Show worked solutionHide worked solution for example 3.06

Recognize the structure. Read the median line and the two box edges separately from the whiskers.

Work it through. Both median lines are at 8. For A, Q1 = 5 and Q3 = 11, so IQRA = 11 − 5 = 6. For B, Q1 = 7 and Q3 = 9, so IQRB = 2. The full ranges are 16 − 2 = 14 and 18 − 1 = 17, respectively.

Answer: C.

Check. A’s box is wider even though B’s whiskers extend farther overall. Middle-half spread and full-range spread can rank the data sets differently.

Avoid the trap. Box plots do not generally determine the exact mean or standard deviation. Equal medians do not imply equal distributions.

3.07. Combine means by recovering totals

Twelve students have a mean score of 78, and 18 other students have a mean score of 88. What is the mean score of all 30 students?

Show worked solutionHide worked solution for example 3.07

Recognize the structure. Multiply each group mean by its group size before combining the groups.

Work it through. The score totals are 12(78) = 936 and 18(88) = 1,584. Thus

x̄ = 936 + 1,58412 + 18 = 2,52030 = 84.

This is a weighted mean, with weights 12 and 18.

Answer: 84.

Check. The answer lies between 78 and 88 and is closer to 88, the mean of the larger group.

Avoid the trap. (78 + 88)/2 = 83 gives equal weight to unequal groups. A mean contains information about a total only when its sample size is also known.

Another route. Relative to 78, the larger group contributes an extra 10 points per student. The overall increase is (18/30)(10) = 6, giving 78 + 6 = 84.

3.08. Recover a missing observation from a mean

A set of eight numbers has mean 15. Seven of the numbers have sum 99. What is the eighth number?

Show worked solutionHide worked solution for example 3.08

Recognize the structure. A known mean and count determine the full sum.

Work it through. The sum of all eight numbers is 8(15) = 120. If the missing number is x, then

99 + x = 120 x = 21.

There is no need to know the seven individual values.

Answer: 21.

Check. (99 + 21)/8 = 120/8 = 15, the stated mean.

Avoid the trap. Subtracting 99 from 15 mixes a total with an average. Convert both pieces to totals before subtracting.

3.09. Separate an outlier’s effect on mean and median

The data set is 10, 11, 11, 12, 56. If 56 is removed, what happens to the mean and median?

A. Both decrease.

B. The mean decreases and the median stays the same.

C. The mean stays the same and the median decreases.

D. Both stay the same.

Show worked solutionHide worked solution for example 3.09

Recognize the structure. The mean uses every value; the median depends on the central ordered positions. Recompute both.

Work it through. Originally the sum is 100, so the mean is 20 and the median is 11. After 56 is removed, the remaining sum is 44 across four observations, giving mean 11. The median is the average of the two middle values, 11 and 11, and remains 11.

Answer: B.

Check. Removing a value far above the mean should lower the mean. The unchanged two central 11s explain the median result.

Avoid the trap. The median is resistant, not immune, to data changes. Removing or adding an extreme observation can change the median in another data set by shifting its central positions.

3.10. Shift every value without changing its spread

Data set A has mean 12 and standard deviation 3. Data set B is formed by adding 8 to every value in A. What are the mean and standard deviation of B?

Show worked solutionHide worked solution for example 3.10

Recognize the structure. Adding a constant moves the entire distribution without changing distances from its mean.

Work it through. The new mean is 12 + 8 = 20. For each old value x, the new deviation from the new mean is (x + 8) − (12 + 8) = x − 12.

All deviations are unchanged, so the standard deviation remains 3.

Answer: Mean 20; standard deviation 3.

Check. The range and interquartile range also stay unchanged under a common translation. Location changes, but spread does not.

Avoid the trap. Adding 8 to the standard deviation treats spread as location. Standard deviation measures distances from the mean, not distances from zero.

3.11. Transform the mean and scale the standard deviation

A data set has mean 5 and standard deviation 2. Every value x is replaced by y = 3x − 7. What are the mean and standard deviation of the new data set?

Show worked solutionHide worked solution for example 3.11

Recognize the structure. A linear transformation changes the mean by the full rule, but changes standard deviation only by the absolute scale factor.

Work it through. The new mean is 3(5) − 7 = 8. Deviations from that mean satisfy

y − 8 = (3x − 7) − 8 = 3(x − 5).

Every deviation is multiplied by 3, so standard deviation is 3(2) = 6.

Answer: Mean 8; standard deviation 6.

Check. The subtraction of 7 shifts all observations equally and cannot alter their spread.

Avoid the trap. Do not calculate 3(2) − 7 for standard deviation. A spread cannot become negative. In general the multiplier for spread is |a| in y = ax + b.

3.12. Compare standard deviations when ranges match

Data set A is 0, 5, 5, 5, 10, and data set B is 0, 0, 5, 10, 10. Which statement is true?

A. A has a larger mean.

B. B has a larger mean.

C. Their standard deviations are equal because their ranges are equal.

D. B has the larger standard deviation.

Show worked solutionHide worked solution for example 3.12

Recognize the structure. Compare how much data lie far from the common mean, not just the endpoints.

Work it through. Both means are 5, and both ranges are 10. The squared deviations from 5 sum to

A ∶ 25 + 0 + 0 + 0 + 25 = 50, B ∶ 25 + 25 + 0 + 25 + 25 = 100.

The sets have the same size, so either consistent standard-deviation convention divides these sums by the same denominator. B therefore has the larger standard deviation.

Answer: D.

Check. B places four observations 5 units from the mean; A places only two there. B is more dispersed despite its identical endpoints.

Avoid the trap. Equal ranges do not imply equal standard deviations. You do not need to take either square root to make this comparison.

3.13. Add the mean and compare spread

Data set A is 2, 4, 6, 8, 10. Data set B contains all values in A plus one additional value of 6. How do the mean, median, and population standard deviation change?

Show worked solutionHide worked solution for example 3.13

Recognize the structure. The added value equals the original mean and median. It adds no deviation from the center.

Work it through. The original sum is 30 across five values, so the mean is 6. The new sum is 36 across six values, again giving mean 6; the median is also 6 in both sets. The squared deviations from 6 sum to 40 in both. Population variance falls from 40/5 = 8 to 40/6, so population standard deviation decreases.

Answer: Mean unchanged; median unchanged; standard deviation decreases.

Check. One more observation is now located exactly at the center, making the distribution more concentrated around that center.

Avoid the trap. A larger number of observations does not automatically make a data set more spread out. The effect depends on the added values. Here the population convention is specified explicitly.

3.14. Recover an unknown group size from a combined mean

Group A has mean 18. Group B has 12 observations and mean 30. When the two groups are combined, the mean is 26. How many observations are in group A?

Show worked solutionHide worked solution for example 3.14

Recognize the structure. Write each total as count times mean, including the unknown count.

Work it through. Let group A contain n observations. Then

18n + 12(30)n + 12 = 26.

Multiplying by n + 12 gives 18n + 360 = 26n + 312, so 48 = 8n and n = 6.

Answer: 6 observations.

Check. The combined total is 6(18) + 12(30) = 468 across 18 observations; 468/18 = 26.

Avoid the trap. The combined mean lies closer to 30, so group B must have more observations than A. An answer exceeding 12 would contradict that weighting.

Another route. Relative to 26, each A observation is 8 below and each B observation is 4 above. Balance the deviations: 8n = 12(4), giving n = 6.

3.15. Optimize a mean subject to a median constraint

Five integers each lie between 2 and 14, inclusive. Their median is 8. What is the greatest possible mean of the five integers?

Show worked solutionHide worked solution for example 3.15

Recognize the structure. Sort the values. The median fixes the third value and bounds the first two.

Work it through. Write a ≤ b ≤ c ≤ d ≤ e. Since the median is 8, c = 8, and a, b ≤ 8. To maximize the sum, take a = b = 8 and d = e = 14. These choices satisfy every condition. The greatest mean is

8 + 8 + 8 + 14 + 145 = 525 = 10.4.

Answer: 10.4.

Check. The constructed set has five allowed integers and its middle value is 8. Any increase in either of the first two values would push it past the required middle bound.

Avoid the trap. Setting four values equal to 14 would force the median to be 14. Optimizing a sum requires preserving the position constraint, not just the individual upper bound.

Section 3: exit questions

Try the five section-exit questions before checking the explanations. Use scratch paper for your reasoning and enter the requested numerical value for a question without choices.

Topic practice