Sample-based inference and margin of error
Reading position saved on this browser; this is not a completion record.
Start here
An estimate is useful because its uncertainty is understood.
The central idea: samples estimate populations
A population is the full group of interest. A sample is the subset actually observed. A parameter describes the population, while a statistic is calculated from a sample. Different random samples can produce different statistics even when the underlying population is unchanged.
Quantity | Sample statistic | Population target |
|---|---|---|
Mean numerical value | Sample mean | Population mean μ |
Fraction in a category | Sample proportion | Population proportion p |
A mean of 6.4 calculated from 200 sampled households is a statistic even when the survey’s goal is to learn about all 5,000 households. The population’s size and the sample’s size serve different roles.
Estimate a count or total from a representative sample
For a population of size N, estimate a category count using N. Estimate a total numerical quantity using N. If 30% of a representative sample prefers an option, 30% of the population is the natural point estimate, not a guaranteed exact count.
A random sampling design supports extending sample information to the population from which the sample was selected. Coverage and response problems can still matter. A very large convenience sample is not made representative just by having many observations. Section 7 explains these design limits.[13]
A point estimate needs an uncertainty statement
A reported margin of error E around an estimate produces endpoints
− E and + E.
The full width is 2E, not E. The margin uses the same units as the estimate. For a percentage estimate, a margin stated in percentage points is added and subtracted directly from that percentage.
Identify the target before interpreting an interval
An interval for a population mean concerns the mean, not every individual’s value. An interval for a population proportion concerns a fraction, not the probability that a particular person has a numerical measurement inside those endpoints.
What a confidence interval does—and does not—say
Read the stated confidence level correctly
A 95% confidence procedure is designed so that, under its assumptions, about 95% of intervals constructed from repeated samples would contain the fixed population parameter. Once one interval is calculated, it either contains that parameter or it does not. The 95% describes the procedure, not the proportion of individuals inside an interval for a mean.[8]
For SAT-style interpretation, think “a reasonable interval estimate for the stated population quantity at the reported confidence level.” Avoid replacing that careful statement with a guarantee. A candidate value inside the interval is compatible with the interval; it is not proved to be the true value. A value outside is not logically impossible merely because one interval excludes it.
Convert the entire interval when the target changes
If the estimated population proportion is p ± E and the known population size is N, the corresponding count endpoints are N(p − E) and N(p + E). Use decimal proportions in the multiplication. The estimated count’s margin is NE, provided N is treated as known exactly.
If a claim requires more than half the population, check whether the whole interval lies above 50%. An estimate above 50% with an interval reaching below it does not establish that majority at the stated confidence level.
Larger samples generally improve precision, not certainty
Holding the sampling method, confidence level, and relevant variability comparable, a larger random sample generally produces a smaller margin of error.[3] It does not guarantee that a particular realized estimate is closer to the truth, and it does not repair bias in how participants were selected or measured.
A useful supplied model for sample size
Some practice problems give E
Halving the margin requires four times the sample size; reducing it to one-third requires nine times the size. Treat this as a supplied-model application, not a universal formula for every survey design. No memorization of z-scores, t-tables, or formal interval formulas is required in this guide.
Avoid overclaiming precision
Population size is not the same as sample size
When a simple random sample is a tiny fraction of a very large population, its precision depends far more on sample size than on the number of millions in the population. Sampling 1,000 rather than 100 generally matters much more than whether the target population contains one million or ten million people, all else comparable. When a sample is a substantial fraction of a finite population, that simplification is less appropriate.[13]
Use population weights after disproportionate sampling
Sampling equally many people from two unequal-sized groups deliberately gives the smaller group more sample representation. To estimate a full-population rate, combine subgroup estimates using population shares, not automatically the sample shares. If one group is three-quarters of the population, it should receive three-quarters of the weight in the combined estimate, assuming the subgroup estimates are appropriate.
Margin of error is not a blanket error allowance
A reported sampling margin does not normally account for every possible measurement mistake, misleading question, missing group, or nonresponse bias. A narrow interval can be centered on a biased estimate. Separate the precision of a sampling method from the quality of the full data-collection process.
Comparing two estimates requires restraint
A larger point estimate ranks the reported numbers; it does not by itself establish a difference between the underlying populations. Overlap between two separately reported intervals does not prove equality, and is not a universal test for whether the difference is statistically convincing. A formal comparison needs uncertainty for the difference, including the appropriate design and dependence information.[9]
This guide’s interval-comparison questions ask what the supplied information justifies. Do not invent a significance test or a confidence level for a difference when none is supplied. Similarly, “not convincingly different” is not the same as “identical.”
The inference checklist
Name the target population and parameter. Identify the sampling method. Compute the point estimate and any interval in matching units. State the result as an estimate, not a census fact. Check whether the proposed conclusion is stronger than the uncertainty and design allow.
What to practice next
Examples 6.01–6.09 build estimates and interpret intervals. Examples 6.10–6.12 examine sample size. Examples 6.13–6.15 distinguish precision, weighting, and justified comparisons.
15 worked examples
Try the question before reading the solution. The examples progress from Foundation to Challenge.
6.01. Distinguish a statistic from its target parameter
A researcher selects a simple random sample of 200 students from a school’s 5,000 students. All selected students respond. Their mean weekly library use is 6.4 hours. Which quantity is a sample statistic?
A. The unknown mean for all 5,000 students.
B. The observed mean of 6.4 hours for the 200 sampled students.
C. The number of students in every school in the district.
D. The exact library-use time of every student in the school.
Show worked solutionHide worked solution for example 6.01
Recognize the structure. A statistic is computed from a sample; a parameter describes the population.
Work it through. The value 6.4 is calculated using the 200 sampled students, so it is the sample mean, a statistic. It can be used to estimate the population mean for the 5,000 students, but that unknown population mean is a different quantity.
Answer: B.
Check. A different random sample could yield a different sample mean even though the population being studied remains the same.
Avoid the trap. Random selection does not turn an estimate into an exact census measurement. Keep the observed summary distinct from the quantity it estimates.
6.02. Project a sample proportion to a population count
A simple random sample of 120 customers is selected from 6,000 customers. All respond, and 42 prefer electronic invoices. Based on this sample, what is the best estimate of the number of all 6,000 customers who prefer electronic invoices?
Show worked solutionHide worked solution for example 6.02
Recognize the structure. Estimate the population share using the sample share, then multiply by the population size.
Work it through. The sample proportion is 42/120
0.35(6,000)
This is a projection from the sample, not a known exact count.
Answer: About 2,100 customers.
Check. The full population is 50 times the sample size, and 42(50)
Avoid the trap. Do not multiply 42 by 6,000 without dividing by 120. The estimate applies to the population sampled, not to customers of unrelated businesses.
6.03. Use a sample mean to estimate a total
A simple random sample of 50 recycling bins from a collection of 400 bins has mean contents mass 2.8 kilograms. What is the best estimate of the total contents mass, in kilograms, for all 400 bins?
Show worked solutionHide worked solution for example 6.03
Recognize the structure. A mean is an estimated amount per bin. Multiply by the number of population bins.
Work it through. Use 2.8 kg as the estimate of the population mean. Then
= 400 bins × 2.8 = 1,120 kg.
The sample size was used to obtain the sample mean; it is not the multiplier for the population total.
Answer: About 1,120 kilograms.
Check. The sample total is 50(2.8)
Avoid the trap. The estimate 2.8 is a mean, not a total. Also, the bins are not all required to contain exactly 2.8 kg for this estimation method to make sense.
6.04. Build an interval from a percentage margin of error
A survey estimates that 58% of a population supports adding weekend library hours, with a margin of error of 4 percentage points at the stated confidence level. What interval is reported for the population percentage?
Show worked solutionHide worked solution for example 6.04
Recognize the structure. A percentage-point margin is added to and subtracted from the estimate.
Work it through. The lower endpoint is 58 − 4
Its total width is 8 percentage points; the margin of error is the half-width.
Answer: 54% to 62%.
Check. The midpoint is (54 + 62)/2
Avoid the trap. 58%(1 ± 0.04) would treat the margin as 4% of the estimate, not 4 percentage points. Also, the interval is not a guarantee that the unknown parameter lies inside it.
6.05. Identify what a mean interval describes
From a random sample, a researcher estimates a population’s mean commute time as 72 minutes with a margin of error of 3 minutes at a stated confidence level. Which interpretation is appropriate?
A. Every individual commute is between 69 and 75 minutes.
B. The interval from 69 to 75 minutes estimates the population mean commute time.
C. Exactly 95% of individual commutes last 72 minutes.
D. The population mean is known to be exactly 72 minutes.
Show worked solutionHide worked solution for example 6.05
Recognize the structure. An interval for a mean concerns the population average, not the spread of individual observations.
Work it through. The interval is 72 ± 3, or 69 to 75 minutes. Its target is the mean commute time of the population sampled. Individual commutes may be far below 69 or above 75. The confidence level concerns the interval procedure, not a stated percentage of individual travel times.
Answer: B.
Check. A population containing both short and long commutes can have a mean near 72. Nothing in the reported interval forces the individual times to cluster within six minutes.
Avoid the trap. Margin of error is not the range or standard deviation of the observations. Never replace “population mean” with a statement about every individual.
6.06. Convert interval endpoints into estimated counts
A survey estimates that 42% of a town’s 18,000 households use a service, with a margin of error of 3 percentage points. What is the upper endpoint of the corresponding interval for the number of households using the service?
Show worked solutionHide worked solution for example 6.06
Recognize the structure. Convert the requested percentage endpoint to a count using the known population size.
Work it through. The upper percentage endpoint is 42% + 3%
For comparison, the full count interval runs from 0.39(18,000)
Answer: 8,100 households.
Check. The point estimate is 0.42(18,000)
Avoid the trap. This is an interval endpoint, not an absolute maximum possible count or a guaranteed upper bound. The question asks for the reported estimate’s upper endpoint.
6.07. Recover a point estimate and margin from an interval
A reported symmetric interval for a population proportion is [0.31, 0.39]. What are the point estimate and margin of error?
Show worked solutionHide worked solution for example 6.07
Recognize the structure. For a symmetric estimate-plus-or-minus-margin interval, the center is the estimate and half the width is the margin.
Work it through. The midpoint is
= = 0.35.
The margin is (0.39 − 0.31)/2
Answer: Point estimate 0.35; margin of error 0.04.
Check. 0.35 − 0.04
Avoid the trap. The full interval width is 0.08, not the margin of error. The stated symmetry is important: not every statistical interval is centered symmetrically on its estimate.
6.08. Check whether a proposed value is within the reported interval
A study estimates a population mean as 81 units with a margin of error of 2.5 units. Which proposed population mean lies within the reported interval?
A. 77 B. 78
C. 82 D. 85
Show worked solutionHide worked solution for example 6.08
Recognize the structure. Construct the interval and compare each candidate with its endpoints.
Work it through. The interval is [81 − 2.5, 81 + 2.5]
Answer: C: 82.
Check. The distance from 81 to 82 is 1, no more than the margin 2.5. The distances for 77, 78, and 85 are 4, 3, and 4.
Avoid the trap. A value outside the interval is not logically impossible. The question asks which value is included in the reported uncertainty interval, not which value is known with certainty.
6.09. Evaluate a majority claim with uncertainty
A random survey estimates that 53% of a population favors a proposal, with a margin of error of 4 percentage points. Which conclusion is best supported by the reported interval?
A. A population majority is guaranteed.
B. Exactly 53% of the population favors the proposal.
C. The interval includes values below and above 50%, so it does not establish a population majority at the stated confidence level.
D. A population majority is impossible.
Show worked solutionHide worked solution for example 6.09
Recognize the structure. Compare the entire interval, not just the point estimate, with the 50% threshold.
Work it through. The interval is [49%, 57%]. It includes population shares both below and above 50%. The sample estimate favors a majority, but the reported interval does not place the population share entirely above the majority threshold.
Answer: C.
Check. A share of 49.5% and a share of 55% both lie inside the interval, yet only one is a majority.
Avoid the trap. A point estimate above 50% is not the same as convincing interval evidence of a population majority. Conversely, uncertainty about a majority is not evidence that a majority is impossible.
6.10. Compare expected precision under comparable designs
Two simple random samples are drawn from the same very large population using the same survey question and confidence level. Sample A has 400 respondents and sample B has 1,600. Their estimated proportions are similar. Which sample would generally have the smaller margin of error?
Show worked solutionHide worked solution for example 6.10
Recognize the structure. Hold the method, confidence level, and underlying variability comparable before comparing sample sizes.
Work it through. Sample B has more observations and therefore generally provides a more precise estimate, with a smaller margin of error. This is an expectation about sampling uncertainty, not a guarantee that B’s point estimate is closer to the truth in every particular pair of samples.
Answer: Sample B.
Check. The question controls the other major conditions that could complicate the comparison. A larger well-designed sample contains more information about the same target.
Avoid the trap. More data do not automatically eliminate selection bias, and they do not necessarily reduce the standard deviation of individual observations. Margin of error describes uncertainty in an estimate.
6.11. Apply a supplied sample‐size model
For a particular survey design, the margin of error E follows the supplied model E = k/, with k fixed and n the sample size. When n = 225, E is 8 percentage points. According to this model, what is E when n = 900?
Show worked solutionHide worked solution for example 6.11
Recognize the structure. Use the stated inverse-square-root model. This formula is supplied, not assumed to be a universal margin-of-error rule.
Work it through. Since 900
Enew = 8 = 8 () = 4.
Answer: 4 percentage points.
Check. From the original case, k = 8 = 120. Then 120/ = 120/30 = 4.
Avoid the trap. Quadrupling the sample size does not divide this margin by four. The denominator is , not n. Exact sample-size arithmetic here depends on the explicitly given model.
6.12. Do not substitute population size for sample size
Survey A samples 1,000 people at random from a population of 1 million. Survey B samples 1,000 people at random from a population of 10 million. Both use the same method and confidence level, obtain similar estimated proportions, and sample only a negligible fraction of their populations. Which comparison is most reasonable?
A. B’s margin of error must be ten times larger.
B. A’s margin of error must be ten times larger.
C. Their margins of error should be approximately the same.
D. Both margins of error are zero.
Show worked solutionHide worked solution for example 6.12
Recognize the structure. Under the stated negligible-sampling-fraction condition, comparable sample sizes provide comparable precision.
Work it through. Both surveys observe 1,000 people under comparable conditions. The much larger population in B does not by itself require a tenfold margin. Their margins should be approximately the same. This comparison relies on the small sampling fractions; sampling a substantial fraction of a finite population can change the calculation.
Answer: C.
Check. Neither survey is a census, so zero uncertainty is not justified. The problem deliberately holds the sample sizes and other precision factors alike.
Avoid the trap. A sample’s percentage of a huge population is not the main measure of its precision. Distinguish n, the sample size, from N, the population size.
6.13. Separate sampling uncertainty from selection bias
A website reports the opinions of 50,000 visitors who voluntarily answer a poll. A second study randomly selects 800 people from the target population and obtains responses from all of them. Which study has the stronger design for estimating the target population’s opinion?
Show worked solutionHide worked solution for example 6.13
Recognize the structure. Representativeness depends on selection, not just on the number of responses.
Work it through. The second study uses a random sample of the target population with complete response, so its design supports population estimation. The voluntary poll may overrepresent people who visit that website and choose to respond. A large response count can reduce random variability while leaving systematic selection bias.
Answer: The second study.
Check. Even an enormous poll of a group with unrepresentative opinions need not describe the target population well. The source of respondents remains crucial.
Avoid the trap. Do not claim that the smaller random sample is guaranteed to be closer in this particular instance. It has the stronger inferential design, not guaranteed superior realized accuracy.
6.14. Weight subgroup estimates back to the target population
A district has 1,000 students at School A and 3,000 at School B. Separate simple random samples of 100 students from each school all respond. In A’s sample, 60% prefer a proposed schedule; in B’s sample, 40% prefer it. What is the best estimate of the percentage of all 4,000 district students who prefer the schedule?
Show worked solutionHide worked solution for example 6.14
Recognize the structure. The samples are equal-sized, but the population groups are not. Weight by population sizes.
Work it through. Estimated supporters are 0.60(1,000)
= = 0.45.
The district estimate is 45%.
Answer: 45%.
Check. The population weights are 1/4 and 3/4, so (1/4)(60%) + (3/4)(40%)
Avoid the trap. Pooling the 200 sample responses without weights gives 50%, overrepresenting the smaller school. A deliberately stratified design can be valid, but its estimates must respect the sampling design.
6.15. Avoid overclaiming from two separate intervals
District A reports a population-share estimate of 48% ± 4 percentage points. District B reports 52% ± 4 percentage points at the same confidence level. Which statement is justified from these summaries alone?
A. The population shares are exactly equal because the intervals overlap.
B. B’s true population share is certainly larger.
C. The point estimates differ by 4 percentage points, and the reported intervals overlap.
D. The overlap proves that a formal comparison could never find a difference.
Show worked solutionHide worked solution for example 6.15
Recognize the structure. Distinguish observable features of reported intervals from claims about unknown population shares.
Work it through. A’s interval is [44%, 52%] and B’s is [48%, 56%]. They overlap from 48% to 52%, and the point estimates differ by 52−48
Answer: C.
Check. Several different orderings of plausible population shares fit these separate intervals, so certainty about the true ranking is not warranted.
Avoid the trap. “The intervals overlap” is not equivalent to “the populations are equal” or “a difference is impossible.” Keep the claim no stronger than the reported evidence.
Section 6: exit questions
Try the five section-exit questions before checking the explanations. Use scratch paper for your reasoning and enter the requested numerical value for a question without choices.