Inference from sample statistics and margin of error
SAT-only content. What a margin of error does and, more importantly, does not mean.
~30 min · prequestion, worked examples, retrieval practice
A margin of error is not a measure of how wrong you might be — it is a measure of how much your number would wobble if you drew a different random sample the same way, and it is completely blind to every other way a study can be wrong. Almost every point lost on this topic is lost to a sentence that is arithmetically correct and says more than the sample can support.
Before you read on
Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.
Before any teaching. A random sample of 400 adults from a city reports a mean daily commute of 34.2 minutes, with a margin of error of 1.6 minutes at the 95% confidence level. Which statement does that support?
The same commute study is repeated with 1,600 randomly selected adults instead of 400 — same city, same method, same 95% confidence level. What happens to the margin of error, which was 1.6 minutes?
A magazine emails a survey to its own subscribers, receives 5,000 replies, and reports that 71% support a policy, with a margin of error of 1.3 percentage points. A statistician objects that the result says nothing about the country's adults. Does the very small margin of error answer that objection?
Scope note — SAT only. College Board's assessment framework marks this skill point, along with its sibling Evaluating Statistical Claims, as assessed on the SAT and not on the PSAT/NMSQT or PSAT 10. If you are working toward a PSAT first, none of this is on that test. If you are working toward the SAT, all of it is, and no amount of PSAT practice will have drilled it — which is exactly why it shows up as an unexplained gap on first full-length SAT attempts.
What the question is actually testing
This skill point is about using a number measured on a sample to say something defensible about a population that was never measured. The Digital SAT tests the reasoning, not the statistics: you will be given a sample statistic and a , and asked what may legitimately be concluded from them. No released item has required computing a margin of error from raw data — the value is always supplied, the standard-error formula is not on the reference sheet, and it is not needed.
Four shapes cover nearly every item. BUILD: given an estimate and a margin, identify the plausible interval, or a value inside it. INTERPRET: four sentences about the same study, of which three overreach, understate, or describe the wrong quantity. COMPARE: two studies, and a question about which has the smaller margin of error or what a wider interval implies. SCOPE: which population a result extends to, and what the study's design does or does not license.
College Board publishes the domain weight, not the skill weight: Problem-Solving and Data Analysis is about 15% of the scored Math questions — five to seven of them — split across seven skill points. Counts you will see for this particular skill, usually one item per form and sometimes none, are prep-industry inferences from released material rather than official figures. Treat them as a rough prior; the reason to learn this properly is not frequency but that the items are nearly all recoverable, because the arithmetic is trivial and the difficulty is entirely in the sentence.
That is the structural fact to carry into every item below. On most Math topics a wrong answer means you did the algebra wrong. Here, the arithmetic is one addition and one subtraction, and three of the four choices will contain the correct numbers. The question is which sentence they are wrapped in.
Foundations — from zero (skip if this is already automatic)
If "a sample mean of 12.5 with a margin of error of 0.8" already reads to you as "plausibly 11.7 to 13.3, for the population, not for any one member," skip to the next block. If it doesn't, this block builds the whole idea from nothing and it is the most valuable thing on this page.
A POPULATION is every member of the group you actually care about — all 1,200 employees at a plant, all adults in a city, all bottles produced in a day. A SAMPLE is the subset you actually measure. You take a sample for one reason: measuring everyone is impossible or unaffordable. A number computed from the population (its true mean) is called a parameter, and you almost never know it. A number computed from the sample (the sample mean) is called a statistic, and you always know it exactly. The whole topic is the bridge between the two.
Here is the problem the bridge has to solve. Draw 400 adults at random and you get a mean of 34.2 minutes. Send someone out to draw a different 400 the same way and they get 33.8. A third draw gives 34.9. Nobody made a mistake — different samples simply contain different people. That spread across hypothetical repeat samples is called sampling variability, and it is the only thing a margin of error measures.
A is reported so you can convert a single estimate into a range. The rule is one line: plausible interval = estimate − margin, up to estimate + margin. Written with the shorthand, estimate ± margin. For 34.2 ± 1.6, the interval runs 32.6 to 35.8. Write both endpoints out in full before you look at any answer choice — the most common careless loss on this topic is a subtraction slip on a decimal, and it is fully preventable.
The margin is the HALF-width of the interval, not the width. If a question hands you the interval instead of the margin — "plausibly between 61.4 and 68.2" — recover the two numbers by taking the midpoint and the half-gap: estimate = (61.4 + 68.2)/2 = 64.8, margin = (68.2 − 61.4)/2 = 3.4. Check it closes: 64.8 − 3.4 = 61.4 and 64.8 + 3.4 = 68.2. That check takes four seconds and catches the halving error every time.
"At the 95% confidence level" is a statement about the procedure, not about this one interval. It means that if you repeated the whole exercise many times — draw a fresh random sample, compute the estimate, build the interval — about 95% of the intervals you built would contain the population's true value. That is why every credited SAT answer says it is PLAUSIBLE that the parameter lies in the range, and never that it definitely does. You will hear people say "there is a 95% chance the true mean is in this interval"; that phrasing is loose, statisticians dispute it, and it is not how a credited answer is ever worded.
One notation trap before anything else. When the estimate is a percentage, the margin is given in PERCENTAGE POINTS, not as a percent of the estimate. "38%, margin of error 3 percentage points" means 35% to 41%. It does not mean 38 ± 3% of 38, which would be 36.86% to 39.14%. Percentage points are additive; percents of a number are multiplicative, and the SAT writes distractors from that confusion in more than one topic.
Finally, the design words. means members of the population were chosen at random to be measured; it is what licenses generalising the result back to that population, and nothing more. means subjects already in hand were assigned at random to a treatment or a control group; it is what licenses a causal claim. They are different operations, they license different conclusions, and confusing them is the most heavily tested idea in this corner of the syllabus.
The method, in four steps
Step 1 — name the population that was actually sampled, out loud, before touching a number. Not the population the writer of the study wants to talk about; the one the sampling procedure could physically reach. "150 students selected at random from those in the library on Tuesday afternoon" is a random sample of Tuesday-afternoon library visitors. Underline that clause. On scope items it is the answer.
Step 2 — name the parameter being estimated, with its unit. A mean number of hours. A percentage of riders. A mean fill volume in fluid ounces. Say whether it is a mean or a proportion, and say what it is a mean or proportion of. Half of the interpret-the-sentence items are decided here, because the tempting wrong choice is about a real quantity that is not this one.
Step 3 — build the interval and write both endpoints. Estimate − margin, estimate + margin. Do it before reading the choices, for the same reason you predict a word before reading the choices on Words in Context: once both numbers are on your page, three of the four sentences stop looking plausible.
Step 4 — test each choice against three filters, in this order. RIGHT POPULATION: is the sentence about the group the sample was drawn from? RIGHT QUANTITY: is it about the mean or proportion being estimated, rather than about individuals, or about the sample itself, or about a different variable? RIGHT STRENGTH: does it say plausible rather than certain, and does it avoid claiming a cause? A choice must pass all three. Most wrong choices fail exactly one, which is what makes them tempting.
The weak path on these items is to read the four sentences and pick whichever feels the least objectionable. That is recognition, and the choices are written to defeat it — they are usually all true-sounding, and often three of them are literally true statements about something. The strong path is to arrive at the choices already holding a population, a quantity, an interval, and a verb.
What a margin of error does and does not mean
It DOES quantify sampling variability: roughly how far this estimate could sit from the population's true value purely because of which members of the sampling frame happened to be selected. It DOES let you convert a point estimate into a plausible range for a population parameter. It DOES shrink as the sample grows and widen as the population's underlying grows.
It does NOT bound the maximum possible error — a value slightly outside the interval is not impossible, merely less plausible. It does NOT guarantee that the true value lies inside; at 95% confidence, about one interval in twenty built this way misses. It does NOT describe individuals: an interval for a mean says nothing about the range that any given person, bottle, or fish falls in. And it does NOT detect bias of any kind — not a sampling frame that excludes half the population, not non-response, not a leading question, not a survey answered only by people who felt strongly. Every one of those shifts the centre of the interval and leaves its width untouched.
That last point is the one the test rewards most. Ask which errors a margin of error could possibly have measured, and the answer is: only the ones that come from the draw. Everything upstream of the draw — who could be drawn at all, what they were asked, who bothered to answer — is outside its jurisdiction, so a small margin of error is never evidence that a study is sound.
What a WIDER interval implies, and what it does not. Three things widen it: a smaller sample, a more variable population, and a higher confidence level. Widening means less precision — a larger set of population values remains plausible — and, in practice, a weaker ability to distinguish two groups from each other. What widening does NOT mean is that the study is worse, that the estimate is biased, or that the estimate is wrong. A careful researcher who reports a 99% interval will publish a wider range than a colleague who reports a 95% interval from identical data, and the wider one is the more cautious claim, not the sloppier one.
The trade is worth stating explicitly because the SAT tests it directly: confidence and precision are bought from each other. Demand more confidence that your net contains the fish and you must use a bigger net. The only way to have both — high confidence and a narrow interval — is to collect more data, which is why sample size is the lever every comparison item turns on.
Mechanism
Why more data narrows the interval, and why it cannot rescue a biased one
The margin of error has the form (a number set by the confidence level) × (the spread of the data) ÷ √n. You are never asked to use that expression on the SAT, but its shape explains every comparison item on the topic. The √n in the denominator is there because the sample mean is an average of n independent draws, and averaging is a cancellation process: values above the true mean and values below it increasingly offset one another as n grows, so the average settles down even though the individual values are as spread out as they ever were. The cancellation accumulates with the square root of the count rather than the count itself, which is why quadrupling a sample halves its margin instead of quartering it, and why the fourth thousand respondents buy far less precision than the first hundred did. Now notice what that derivation assumed: that the draws are from the population you want to describe. Bias violates the assumption before the arithmetic starts. If your frame is a magazine's subscribers and your target is a country, every additional draw is another sample of subscribers, so the estimate converges — tightly, obediently, with a beautifully small margin of error — on the subscribers' opinion. The formula cannot notice, because n counts how many people you asked and contains no term for whether they were the right people. That is the precise sense in which precision and accuracy are independent quantities: √n governs one of them and the sampling frame governs the other, and only the first appears in the number the study reports.
Worked examples
Fully worked — build the interval, then say only what it supports
- 01Problem: a biologist nets 180 fish at random from a lake and records a mean length of 24.6 cm, with a margin of error of 1.4 cm at the 95% confidence level. State the strongest conclusion the data support.
- 02Step 1 — name the sampled population: all fish in that lake. Not fish in nearby lakes, not fish of that species generally. Any conclusion must be a sentence about this lake.
- 03Step 2 — name the parameter and its unit: the MEAN length, in centimetres, of all fish in the lake. Not the length of an individual fish, and not the mean of the 180 netted (which is known exactly and is 24.6).
- 04Step 3 — build the interval and write both endpoints: 24.6 − 1.4 = 23.2 and 24.6 + 1.4 = 26.0. Plausible range: 23.2 cm to 26.0 cm.
- 05Step 4 — state it with the right verb: it is plausible that the mean length of all fish in the lake is between 23.2 cm and 26.0 cm.
- 06Then run the three filters against the sentences you would reject. "Every fish is between 23.2 and 26.0 cm" fails RIGHT QUANTITY — that is a claim about individuals. "The mean length of all fish in the lake is 24.6 cm" fails RIGHT STRENGTH — 24.6 is the sample's mean, and the population's is unknown. "The mean length is certainly between 23.2 and 26.0" fails RIGHT STRENGTH — about one interval in twenty built this way misses.
- 07Answer: it is plausible that the mean length of all fish in the lake is between 23.2 cm and 26.0 cm.
One step hidden — comparing two sample sizes
- 01Problem: two firms estimate the mean household size in the same city. Firm A samples 250 households at random and reports a margin of error of 0.30 people at 95% confidence. Firm B samples 1,000 households at random from the same city, same method, same confidence level. Roughly what margin should Firm B report?
- 02Identify what changed and what did not. Sample size went from 250 to 1,000, a factor of 4. Population, method and confidence level are all unchanged, so every term in the margin except √n is the same for both firms.
- 03Apply the DIRECTION first, because it is what the SAT usually asks and it eliminates choices instantly: a larger sample means less sampling variability, so B's margin must be smaller than 0.30.
- 04Apply the SIZE: the margin scales with 1/√n, so multiplying n by 4 divides the margin by √4 = 2. B should report about 0.15 people.
Two steps hidden — do two intervals establish a difference?
- 01Problem: a random sample of first-year students slept a mean of 6.8 hours per night, margin of error 0.4 hours. A random sample of fourth-year students slept a mean of 7.1 hours, margin of error 0.5 hours. May the university conclude that fourth-year students sleep more on average?
- 02Build both intervals before comparing anything at all. First-years: 6.8 ± 0.4 → 6.4 to 7.2 hours. Fourth-years: 7.1 ± 0.5 → 6.6 to 7.6 hours.
- 03Compare the INTERVALS, not the two point estimates. The ranges overlap on 6.6 to 7.2, so a single value — 7.0, say — is a plausible population mean for both groups at once.
Solve alone
- 01Problem: a transit agency surveys 400 randomly selected riders and reports that 62% are satisfied with service, margin of error 4.8 percentage points at 95% confidence. A council member says: "Since 62 minus 4.8 is still well above half, we can be certain that most riders — and most city residents — are satisfied." Name every separate error in that sentence, then state what the data do support.
In your own words
In one sentence: why does step 4 of the second worked example divide the margin of error by 2 rather than by 4 when the sample size is multiplied by 4 — and what does that tell you about the price of precision?
Named traps
- Certainty upgrade
- Taking the correct interval and asserting it: "the mean IS between 7.07 and 7.63" instead of "it is plausible that the mean is between 7.07 and 7.63." The numbers match the credited answer exactly, which is what makes this the most-missed distractor on the topic. An interval estimate is a claim about plausibility; if the sample could establish the parameter, no interval would be needed.
- Interval applied to individuals
- Reading a range built for a MEAN as a range that every member falls inside — "every bottle contains between 7.07 and 7.63 ounces," "95% of bottles are in that range." Individual values are always far more spread out than the estimate of their average. Ask what the interval is an interval FOR, and this collapses immediately.
- Precision mistaken for accuracy
- Treating a small margin of error as evidence that a study is trustworthy. The margin measures only how much the draw wobbles; it is blind to a wrong sampling frame, non-response, and self-selection. Collecting more data the same way narrows the interval around the same off-target number — a bigger biased sample buys a more precise wrong answer.
- Half-width slip
- Confusing the margin with the width of the interval. The margin is half the width: an interval of 61.4 to 68.2 has a width of 6.8 and a margin of 3.4. The error runs in both directions — doubling the margin when building an interval, or reporting the full width when asked for the margin — and both produce a number that is exactly right for the wrong quantity.
- Overlap read as a verdict
- Two errors that mirror each other. Comparing point estimates while ignoring the intervals entirely ("15.0 is bigger than 14.2, so it went up"), or concluding from overlapping intervals that the two population values are equal. Overlap means the difference has not been established; it never means the values are the same.
- Scope drift
- Extending a conclusion past the population the sample was drawn from — from one plant's employees to a whole company, from riders to residents. Its under-corrected twin costs just as much: refusing to generalise at all, and insisting the result applies only to the people surveyed. A random sample exists precisely to license the step out to the frame it was drawn from, and no further.
The 800-level margin
At 1500 the concepts are not the problem; the sentence-level discrimination and the execution are. Start with the one that costs the most: two answer choices carrying identical numbers, differing only in the verb. "The mean is between 7.07 and 7.63" versus "it is plausible that the mean is between 7.07 and 7.63." A student who has done the arithmetic correctly and stops reading at the numbers picks the first one, and the item was written by someone who knew that. Read every choice to the end of its verb phrase before selecting, and treat any unhedged claim about a population parameter as disqualified on sight.
The second is percentage points versus percent, which turns up in this topic more than anywhere else because so many estimates are proportions. A margin of 3 percentage points on an estimate of 38% gives 35% to 41%. Read as "3 percent of the estimate" it gives 36.86% to 39.14% — a plausible-looking interval, centred correctly, that will match a distractor. Whenever the estimate is a percentage, say the words "percentage points" to yourself as you build the endpoints.
Third, answering the right question about the wrong quantity, which at the top of the scale costs more than any misconception does. The item asks for the margin of error and you report the interval's width. It asks for the largest plausible value and you report the estimate. It asks how much LARGER one study's margin is and you report the second study's margin. It asks about the population and every number you wrote down describes the sample. Before selecting, re-read the final clause of the prompt and say what your answer is a measurement of, and in what unit.
Fourth, the sample-size relationship in both directions. To halve a margin you must quadruple the sample; to cut it to a third, ninefold. Read backwards, a study with a quarter of the data has twice the margin. The trap built on top of this is subtler: an item that asks how the margin changes when the sampling method and confidence level are held fixed can be answered from the ratio alone, because the confidence multiplier and the population spread cancel. A strong student who knows the full formula sometimes selects "cannot be determined without the standard deviation" — a sophisticated answer to a question that was never asked.
Fifth, the over-correction. Students who have properly learned to distrust surveys start selecting the most sceptical sentence available, and the SAT punishes that as reliably as it punishes overreach. A genuine random sample of a plant's employees DOES support a conclusion about that plant's employees. "No conclusion can be drawn," "the sample is too small," and "only the surveyed people can be described" are wrong answers, written specifically to catch a student who has learned half of this topic. Scepticism is a filter, not a default; apply it to the frame, not to inference itself.
Sixth, the design vocabulary, because one clause in the stem decides which conclusions are on the table. Random SAMPLING from a population licenses generalising to that population, and nothing else — in particular, it never licenses a causal claim. Random ASSIGNMENT to groups licenses a causal claim about the people in the study, and generalising that cause to a wider population additionally requires the sample to have been randomly drawn. A study with neither supports description only. Locate both words in the stem before reading any choice; on hard items, one of the distractors is a perfectly correct statement about the other design.
Finally, the slips that are not knowledge failures at all and therefore survive to test day. Decimal subtraction under time pressure: 24.6 − 1.4 is 23.2, not 23.4. A margin stated in minutes against choices in hours. A study reporting thousands of gallons against a choice reporting gallons. An interval given as endpoints when the question wanted the margin. None of these are caught by understanding statistics better; they are caught only by writing both endpoints down and re-reading the ask, which is the cheapest habit on this page and the one that separates a 1500 from a 1600.
Retrieval — with feedback on every choice
A quality-control team selects 120 bottles at random from a day's production run and measures each one. The sample has a mean fill volume of 7.35 fluid ounces, with an associated margin of error of 0.28 fluid ounces at the 95% confidence level. Which conclusion about the day's entire production run is best supported?
INFERENCE FROM SAMPLE STATISTICS & MARGIN OF ERROR — reference card SAT only. Not assessed on the PSAT/NMSQT or PSAT 10. Plausible interval = estimate - margin, up to estimate + margin. Write BOTH endpoints. The margin is the HALF-width. Given an interval: estimate = midpoint, margin = (width)/2. Given as a percentage? The margin is in PERCENTAGE POINTS. 38% +/- 3 pts = 35% to 41%. Credited verb is always "it is plausible that" — never "is", never "certainly", never "proves". The interval estimates a MEAN or a PROPORTION for a population. Never an individual. Margin shrinks with 1/sqrt(n): quadruple the sample to halve the margin; x9 to cut it to a third. Wider interval = smaller sample, OR more variable population, OR higher confidence level. Higher confidence always means a wider interval from the same data. Never more precise. Margin of error measures sampling variability ONLY — blind to bias, non-response, bad frames. More data collected the same biased way = a more precise wrong answer. Random SAMPLING licenses generalising to the sampled population. Nothing causal. Random ASSIGNMENT licenses a causal claim. Different word, different conclusion. Overlapping intervals: difference NOT established. That is not the same as "equal". Over-correction is also wrong: a random sample DOES support a conclusion about its population. Before selecting: name the population, name the quantity, check the verb.
Every item on this page is Meridian-original, written to match the Digital SAT's format and difficulty — it is not a real SAT question. The only source that matches the live test exactly is College Board's own Bluebook and Question Bank.
A quality-control team selects 120 bottles at random from a day's production run and measures each one. The sample has a mean fill volume of 7.35 fluid ounces, with an associated margin of error of 0.28 fluid ounces at the 95% confidence level. Which conclusion about the day's entire production run is best supported?
- AThe mean fill volume of all bottles in the run is between 7.07 and 7.63 fluid ounces.
The arithmetic here is correct — this is exactly the right interval — and the verb is not. Stating that the population mean IS in the range claims something the sample cannot establish; a run whose true mean is 7.05 would produce a sample like this one often enough that it cannot be ruled out. When two choices carry identical numbers, the hedge is the question.
- It is plausible that the mean fill volume of all bottles in the run is between 7.07 and 7.63 fluid ounces.
Correct. Build the interval: 7.35 − 0.28 = 7.07 and 7.35 + 0.28 = 7.63; check it closes by confirming the midpoint of 7.07 and 7.63 is (7.07 + 7.63)/2 = 7.35. The population is the day's whole run (the bottles were drawn at random from it), the parameter is a mean fill volume in fluid ounces, and the verb is the hedged one an interval estimate requires.
- CIt is plausible that the mean fill volume of all bottles in the run is between 7.21 and 7.49 fluid ounces.
This comes from splitting the margin in half and applying 0.14 to each side: 7.35 − 0.14 = 7.21 and 7.35 + 0.14 = 7.49. The reasoning treats 0.28 as the total width of the interval, but the margin of error is the half-width — the amount added on ONE side. The check that catches it: this interval is 0.28 wide in total, so its margin would be 0.14, not the 0.28 the item gave.
- DIt is plausible that at least 95% of the bottles in the run contain between 7.07 and 7.63 fluid ounces.
Right interval, wrong quantity. The 95% attaches to the procedure that produced the interval, not to a share of bottles, and the interval estimates a single number — the run's mean — rather than describing how individual bottles are spread. Individual fill volumes vary far more than an estimate of their average does.
Traps tested: Certainty upgrade · Half width slip · Interval applied to individuals
A report states only this: "Based on a random sample, it is plausible that the mean weekly grocery spend of households in the region is between $61.40 and $68.20." What were the sample mean and the margin of error?
- Sample mean $64.80; margin of error $3.40
Correct. The estimate sits at the midpoint: (61.40 + 68.20)/2 = 129.60/2 = 64.80. The margin is half the width: (68.20 − 61.40)/2 = 6.80/2 = 3.40. Check that it closes in both directions — 64.80 − 3.40 = 61.40 and 64.80 + 3.40 = 68.20 — which is the four-second confirmation worth running every time.
- BSample mean $64.80; margin of error $6.80
The midpoint is right and the margin is the full width of the interval rather than half of it. Applying a margin of 6.80 to an estimate of 64.80 would produce an interval from 58.00 to 71.60 — twice as wide as the one the report actually gave, which is the check that rules this out.
- CSample mean $61.40; margin of error $6.80
This reads the lower endpoint as the estimate and the width as the margin. An interval is symmetric about the estimate by construction — the estimate is always the centre, never an endpoint — so a reported range immediately gives up the sample mean as its midpoint.
- DSample mean $64.80; margin of error $1.70
The midpoint is right, and the margin has been halved once too often: 6.80 is the width, 3.40 is the margin, 1.70 is half the margin. Substituting it back gives 63.10 to 66.50, which is not the interval the report published.
Traps tested: Half width slip · Endpoint read as estimate
To estimate the mean number of hours per week that students at a university spend on paid work, a researcher randomly selected 150 students from among those who visited the campus library on a Tuesday afternoon and surveyed all of them. The sample mean was 9.4 hours, with a margin of error of 1.2 hours. Which of the following best describes the study's principal limitation?
- AA margin of error of 1.2 hours is too wide for the estimate to support any conclusion.
There is no threshold at which a margin of error voids a study — 1.2 hours simply describes how precise the estimate is, and 8.2 to 10.6 hours is a perfectly usable range. This choice invents an accuracy standard the item never set, and it aims at precision when the actual problem is elsewhere entirely.
- The students were selected at random only from library visitors on one afternoon, so the result may not extend to the university's students as a whole.
Correct. The randomisation happened inside a frame — Tuesday-afternoon library visitors — that is narrower than the population the study wants to describe, and library visitors plausibly differ from other students in exactly the variable being measured. Random selection licenses generalising to the frame it was drawn from, and no further; the margin of error cannot detect this, because it measures only variation within the draw.
- CA sample of 150 is too small to estimate anything about a university's students.
Sample size determines precision, not validity, and its effect is already reported: 150 students produced a margin of 1.2 hours. Doubling the sample would narrow the interval and leave the real defect exactly where it is, since more Tuesday-afternoon library visitors are still Tuesday-afternoon library visitors.
- DThe study cannot establish that visiting the library causes students to work fewer paid hours.
A true statement about a claim nobody made. This is an estimation study reporting a mean, not a causal study, so noting the absence of random assignment identifies a limitation the study never needed to overcome. On hard items one distractor is regularly a correct sentence about the other design.
Traps tested: Arbitrary precision standard · Sample size fixation · Causal objection to non causal claim
Two independent random samples were drawn from the students of the same university. Sample P (200 students) gave a mean of 14.2 study hours per week, margin of error 1.1 hours. Sample Q (800 students), drawn a year later, gave a mean of 15.0 hours, margin of error 0.6 hours. Both intervals are at the 95% confidence level. Which conclusion is best supported?
- AStudents studied more in the later year, since 15.0 is greater than 14.2.
This compares the two point estimates and ignores the reason each was reported with an interval at all. Both means are single draws from populations that could each plausibly sit anywhere inside a range; a gap of 0.8 hours between estimates whose margins are 1.1 and 0.6 is well within the wobble either sample could produce on its own.
- The plausible ranges are 13.1 to 15.3 hours and 14.4 to 15.6 hours; because they overlap, the data do not establish that mean study time differs between the two years.
Correct, and the endpoints check out: 14.2 ± 1.1 gives 13.1 to 15.3, and 15.0 ± 0.6 gives 14.4 to 15.6. The ranges share 14.4 to 15.3, so a single value — 15.0, or 14.5 — is a plausible population mean for both years at once. When two plausible ranges overlap, the samples have not told the populations apart.
- CBecause Sample Q has the smaller margin of error, its mean of 15.0 must be closer to the true value than Sample P's 14.2.
A smaller margin means Q's estimate is more PRECISE — repeated samples of 800 cluster more tightly than repeated samples of 200. It does not mean this particular estimate landed closer to the truth, which is unknowable from the data, and it does not rule out the larger sample having drawn an unusual set of students. Precision is a property of the procedure; accuracy is a property of a result nobody can see.
- DBecause the plausible ranges overlap, mean study time was the same in both years.
This is the mirror of choice A and just as strong a claim. Overlap means the difference has not been established; it does not establish sameness. The honest conclusion is that this evidence does not settle the comparison in either direction — a genuine difference of half an hour would be entirely consistent with these two intervals.
Traps tested: Point estimates compared without intervals · Precision mistaken for accuracy · Overlap read as equality
A polling firm surveys 900 randomly selected residents and reports a margin of error of 2.4 percentage points at the 95% confidence level. It now wants a margin of error of about 1.2 percentage points, keeping the same sampling method, the same population, and the same confidence level. Approximately how many residents must it survey?
- A1,800
This halves the margin by doubling the sample — precision bought one-for-one with data. Because the margin scales with 1/√n, doubling the sample from 900 to 1,800 divides it by √2, giving roughly 2.4/1.41 ≈ 1.7 percentage points: better than 2.4, but well short of 1.2.
- 3,600
Correct. With the method and confidence level fixed, the margin scales with 1/√n, so the ratio of margins is √(900/n). Setting 1.2/2.4 = 1/2 gives √(900/n) = 1/2, so 900/n = 1/4 and n = 3,600. Check forward: √3,600 = 60 against √900 = 30, so the margin is divided by exactly 2, and 2.4 becomes 1.2.
- C225
This divides the sample by 4 because the margin is to be divided by 2, inverting the relationship. Running it forward exposes the direction error: √225 = 15 against √900 = 30, so the margin would DOUBLE to about 4.8 percentage points. A smaller sample can never produce a narrower interval.
- DIt cannot be determined without the population's standard deviation.
The standard deviation and the confidence multiplier are both held fixed by the stem (same population, same method, same confidence level), so they appear identically in both margins and cancel in the ratio — leaving only √n. This is the sophisticated wrong answer on the item: correct about the full formula, and applied to a question that only ever asked for a ratio.
Traps tested: Linear sample size scaling · Inverted sample size relationship · Unnecessary parameter demand
From one random sample of 500 households, an analyst computes two intervals for the mean monthly water use: 3,420 to 3,580 gallons, and 3,395 to 3,605 gallons. One is a 95% interval and the other is a 99% interval. Which statement is correct?
- 3,395 to 3,605 is the 99% interval, because a higher confidence level requires a wider interval from the same data.
Correct. Both intervals are centred on the same estimate — (3,420 + 3,580)/2 = 3,500 and (3,395 + 3,605)/2 = 3,500 — so they differ only in margin: 80 gallons against 105 gallons. Confidence is bought with width: to be right more often, the net has to be bigger, so the wider interval is the more confident one.
- B3,395 to 3,605 is the 95% interval, because a lower confidence level allows more room for error.
This reads "more room for error" as belonging to the weaker claim, which reverses the trade. A 95% interval accepts being wrong one time in twenty and can therefore afford to be narrow; a 99% interval accepts being wrong one time in a hundred and must be wide enough to earn that. Less confidence buys a tighter interval, not a looser one.
- C3,420 to 3,580 is the 99% interval, because a more confident claim is a more precise one.
Confidence and precision are the two things being traded against each other, not two names for the same property. From a fixed sample you can have a narrow interval you are less sure of, or a wide interval you are more sure of; having both at once requires more data, which this analyst does not have.
- DThe two intervals must come from different samples, because one sample produces exactly one interval.
An interval is built from the data AND a chosen confidence level, so one sample produces a different interval at every level you ask for. The shared centre of 3,500 is positive evidence that the same sample generated both: two different samples would almost certainly have produced two different estimates.
Traps tested: Confidence width inverted · Confidence conflated with precision · One sample one interval
Up next
Evaluating statistical claims
SAT-only. Random sampling licenses generalising; random assignment licenses causation. Never both.
28 min