Evaluating statistical claims

SAT-only. Random sampling licenses generalising; random assignment licenses causation. Never both.

~28 min · prequestion, worked examples, retrieval practice

Random selection buys you a population; random assignment buys you a cause. Neither one buys you the other, and every wrong answer on this skill is one of them being spent on something it does not pay for. This is the only skill point on the Math section with nothing to compute — no formula, no Desmos, no arithmetic to slip on. Problem-Solving and Data Analysis is about 15% of the Math section, and an item on this skill is worth full marks to a student who reads two phrases carefully and nothing at all to one who reads the study as a story and picks the conclusion that sounds most sensible.

Before you read on

Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.

Question 1
Medium

Before any teaching. Researchers randomly selected 900 adults from a city's residency records and recorded, for each, whether the adult belongs to a gym and his or her resting blood pressure. Gym members had lower resting blood pressure on average. Which conclusion does this study design support?

Question 2
Medium

Before any teaching. A researcher put up flyers at a community centre, and 180 people volunteered for a study on sleep. Each volunteer was randomly assigned to one of two evening routines and followed it for three weeks. The group using routine A fell asleep 14 minutes faster on average. Which conclusion does this design support?

Question 3
Hard

Before any teaching. A survey is mailed to every one of the 20,000 households in a town. Three thousand households return it, and 68% of those returning it say they support a new library tax. Which statement about the 68% figure is best supported?

What the question is actually testing

An item on this skill hands you three or four sentences describing a study and asks which of four conclusions is appropriate, best supported, or most reasonable. There is nothing to calculate. Every number in the description is decoration — the 68%, the 900 adults, the 14 minutes — and changing any of them would not change the answer. What decides the answer is two structural facts about how the study was built, and both are stated in the passage in plain words.

College Board publishes the domain weight: Problem-Solving and Data Analysis is about 15% of the Math section, roughly 5 to 7 of the 44 questions, and it shares that 15% with Geometry and Trigonometry as the two smaller Math domains. What College Board does not publish is how those 5 to 7 questions split among the domain's seven skill points. The commonly repeated prep-industry figure of about one statistical-claims item per test is reverse-engineered from released material — plausible, consistent with what the practice tests look like, and not an official number. Plan with it; do not cite it as fact.

The reason this skill is worth more than its frequency suggests is that it is close to free. There is no arithmetic to slip on, no sign to lose, no unit to convert. An item you can read correctly you will get right every time, at a cost of about forty seconds — which is the cheapest mark on the Math section and a source of time for the Advanced Math items that actually need it.

Weak path: read the study as a story, form an opinion about whether the finding sounds true, and pick the choice that best matches the opinion. This fails reliably, because three of the four choices are written to sound true. Strong path: ignore what the study found entirely, check two things about how it was run, write down the one sentence those two things license, and then find the choice that matches that sentence. The finding is irrelevant to the question being asked. What is being tested is the design.

Scope note. This is SAT-only content in a strict sense: nothing on this page is computed, and the two-switch rule below is a deliberately clean version of a question that real researchers argue about at length. Actual causal inference has a large literature on natural experiments, instrumental variables, matching, and what can and cannot be recovered from observational data when randomisation is impossible — none of it is on this exam, and importing it will cost you marks, because the exam's credited answer is always the one the clean rule produces. Learn the clean rule for the test, and know that it is the simple version.

Foundations — the vocabulary, from zero (skip this block if the two-randomisation rule is already automatic)

If you can already say, without pausing, that random selection licenses generalisation and random assignment licenses causation, and that the two are independent, skip ahead to the next block. If any part of that sentence is unfamiliar, this is the highest-value block on the page — and the vocabulary in it is the entire content of the skill.

A POPULATION is the whole group a claim is about: every adult in a county, all 6,400 fourth-graders in a district, every bolt produced by a factory this month. A SAMPLE is the smaller group actually studied. A STATISTIC is a number computed from the sample — 62% of the 500 people contacted. A PARAMETER is the corresponding number for the population, which is usually unknown and is what the study is trying to learn about. The whole point of sampling is to say something about a parameter you cannot measure using a statistic you can.

The SAMPLING FRAME is the list the sample was actually drawn from, and it is the quiet centre of this whole skill. A frame is not the same thing as the population. "Randomly selected from the county's registered voters" has a frame of registered voters, even if the study is written up as being about county adults. Adults who are not registered had zero chance of selection, so nothing in the result speaks about them. When you read "randomly selected from", the noun that follows is the only group the study can generalise to.

An OBSERVATIONAL STUDY is one where the researcher only watches and records. Nobody is told what to do; people are already doing whatever they do, and the study measures it. A survey is an observational study. So is a study that compares people who already exercise with people who already do not. The defining test is simple: did the researcher DO anything to the subjects, or only record what was already happening?

An EXPERIMENT is one where the researcher imposes a TREATMENT — gives one group the drug, the reading block, the study schedule — and imposes a different treatment, or none, on a COMPARISON GROUP (also called a control group). Two things have to be present for it to count. The researcher must decide who gets what, and there must be at least two conditions to compare. One group given a treatment with nothing to compare against is not an experiment in the sense that licenses anything.

A CONFOUNDING VARIABLE is a third thing tied to both of the two you are looking at, which makes it impossible to say which one is responsible. Ice-cream sales and drownings rise together; the confounder is summer. Students who take music lessons score higher on tests; family income is tied to both. The problem with an observational study is never that the pattern is fake — the pattern is usually real. The problem is that at least one confounder is always available, and the study contains nothing that rules it out.

ASSOCIATION means two quantities move together in the data. CAUSATION means changing one changes the other. Association is what a pattern shows; causation is a claim about what would happen if you intervened. Every credited answer on this skill is one of those two words, attached to one specific group of people, and choosing the wrong word or the wrong group is the whole failure mode.

One last word, because the exam uses it in the technical sense and not the everyday one. BIAS here does not mean prejudice or unfairness — it means a systematic tendency for a procedure to miss in one direction. A survey that reaches only people at home during the day is biased whether or not anyone intended anything. Bias is a property of the method, which is why running the same method on more people does not reduce it.

The two switches

Every study on this exam has two independent switches, set at two different moments, doing two different jobs. Find both, and the licensed conclusion writes itself.

SWITCH 1 — , which happens BEFORE the study, and decides WHO THE CONCLUSION APPLIES TO. It is on when subjects were chosen at random from a stated list. When it is on, the result generalises to that list and to nothing wider. When it is off — volunteers, respondents to an advertisement, one convenient school, whoever was in the shopping centre that afternoon — the result applies to the people studied and to nobody else.

SWITCH 2 — , which happens AFTER you already have the subjects, and decides WHETHER YOU MAY SAY "CAUSED". It is on when the researcher randomly assigned the subjects to two or more conditions. When it is on, a causal conclusion is licensed. When it is off — because the researcher only observed, or because the subjects sorted themselves into the groups by their own behaviour — the conclusion stops at an association.

The two switches are independent, and the exam's whole item bank lives in that independence. All four combinations are writable, and three of the four occur regularly. Random selection AND random assignment: a causal conclusion that generalises to the frame — the strongest result available, rare in real research, and usually the credited answer to a "which design would be appropriate" item. Random selection ONLY: an association, generalisable to the frame. Random assignment ONLY: a causal conclusion, restricted to the subjects who took part. NEITHER: an association among the people studied, and nothing else.

Say that grid out loud once and notice what it refuses to do. Random assignment says nothing about who was in the room, so an experiment on 80 volunteers stays about those 80 volunteers no matter how clean the randomisation was. Random selection says nothing about what was done to the subjects, so a beautifully drawn national sample still cannot tell you what causes what. Neither switch ever compensates for the other, and no amount of either one substitutes for the other.

The procedure

Step 1 — read the description looking for two phrases and ignoring everything else: how the subjects GOT IN, and how they got into their GROUPS. Underline them. The numbers, the finding, and the topic are all irrelevant to the question being asked.

Step 2 — set switch 1. Does the passage say the subjects were randomly selected from a stated group? If yes, write the exact noun that follows "selected from" — that is the ceiling on the population. If no, the ceiling is "the people in this study".

Step 3 — set switch 2. Did the researcher randomly assign the subjects to different conditions, with at least two conditions to compare? If yes, the verb is "caused". If no, the verb is "is associated with".

Step 4 — write the licensed sentence before reading a single answer choice: "[caused / is associated with] ... among [the ceiling population]." This is a prediction step, exactly as on Words in Context, and it exists for the same reason: all four choices are engineered to be readable, and you cannot be talked out of a sentence you already wrote down.

Step 5 — eliminate on two axes, not one. First cut every choice whose claim is stronger than switch 2 allows: causes, leads to, results in, improves, and any recommendation about what someone should do. Then cut every choice whose population is wider than the ceiling. Then cut anything weaker than what you earned — a choice that refuses to generalise when switch 1 was on, or refuses causation when switch 2 was on, is wrong too. Usually exactly one choice survives, and if two do, the difference between them is the hedge.

Mechanism

Why each randomisation does exactly one job, and cannot do the other

Random assignment works by BALANCE. When a coin flip decides which condition each subject gets, every other characteristic those subjects carry — age, motivation, prior sleep quality, income, diet, genetics, things nobody thought of and nothing measured — lands in the two groups by chance rather than by any rule connected to the outcome, so the groups end up alike in everything except what the researcher controlled. When the outcome then differs by more than chance would produce, the treatment is the only systematic difference left standing. That is the whole argument, and it is why randomisation beats the alternative of listing confounders and adjusting for them: it neutralises the confounders you never thought of, which are exactly the ones that ruin observational studies. It is also why no observational study can be repaired by a bigger sample or a cleverer analysis — with nothing assigned, the subjects sorted themselves, and that sorting is itself a systematic difference carrying along everything that produced it. Random selection works by REPRESENTATIVENESS, which is a different property with a different reach. Drawing at random from a list gives every member of that list an equal chance of appearing, so the sample's composition tracks the list's composition on average — the same mix of ages, incomes and opinions, up to sampling variability — which is what makes a statistic computed from the sample an estimate of the parameter for the list. Note the word "list": the mechanism is a fact about the frame the draw was made from, and it says nothing whatsoever about anyone who was not on it. That is also the precise reason a large sample cannot rescue a biased one. Size governs random error, the scatter around the truth, and it shrinks as the sample grows; bias is systematic error, a fixed offset built into the procedure, and it does not shrink at all. Averaging more readings from a mis-calibrated instrument buys a more confident wrong number, and a voluntary-response survey of 9,400 is not a better version of one with 94 — it is the same error, measured more precisely. So the two mechanisms answer two different questions. Balance answers "is the treatment responsible?" Representativeness answers "responsible for whom?" Neither contains any part of the other, which is exactly why the exam can build — and does build — items where one is present and the other is conspicuously missing.

Worked examples

Fully worked — random selection, no assignment

  1. 01Problem: "A researcher randomly selected 350 of the 8,000 students enrolled at a university and recorded, for each, the number of hours per week the student worked at a paid job and the student's grade point average. Students who worked more hours tended to have lower grade point averages. Which conclusion is most appropriate?"
  2. 02Switch 1 — how did the subjects get in? "Randomly selected 350 of the 8,000 students enrolled at a university." Random selection is present, and the frame is students enrolled at this university. The ceiling on the population is therefore students at this university — not university students generally, and not students anywhere else.
  3. 03Switch 2 — how did the subjects get into their groups? They did not. Nobody assigned anybody to work 5 hours or 25; the researcher recorded what students were already doing. This is an observational study, so the verb is "is associated with".
  4. 04Write the licensed sentence before looking at a single choice: "There is an association between hours worked and grade point average among students at this university." That sentence is the ceiling in both directions — anything stronger is unsupported and anything weaker gives away a conclusion the random selection already paid for.
  5. 05Eliminate on two axes. Anything containing causes, lowers, hurts, or a recommendation ("students should work fewer hours") exceeds switch 2 — cut. Anything naming a population wider than this university exceeds switch 1 — cut. Anything saying no conclusion can be drawn falls below what switch 1 licensed — cut.
  6. 06Convince yourself the causal choice is a trap rather than a bold-but-defensible answer by naming one confounder out loud: a student working 25 hours a week probably also has less study time, a longer commute, more family obligations and less access to paid tutoring, and this design cannot separate the hours from any of them. Naming one takes four seconds and settles the item.

One step hidden — random assignment, no selection

  1. 01Problem: "Eighty adults responded to an online advertisement for a study on memory. Each was randomly assigned either to a group that reviewed a word list once a day for four days or to a group that reviewed the same list four times in a single day. Tested a week later, the first group recalled 22% more words on average. Which conclusion is most appropriate?"
  2. 02Switch 1 — the subjects responded to an advertisement, which means they selected themselves. There is no random selection from any stated population, so the ceiling is the adults in this study and nobody else.
  3. 03Switch 2 — "randomly assigned" to one of two review schedules, and there are two conditions to compare. Random assignment is present, so the verb is "caused".
  4. 04Licensed sentence: "Reviewing once a day over four days caused better one-week recall than reviewing four times in one day, for the adults in this study." Both halves are set by the design: a causal verb, attached to a population that is the study itself.

Two steps hidden — the "which design" variant

  1. 01Problem: "A district wants to know whether a new 20-minute daily reading block raises reading scores for its 6,400 fourth-graders. Which study design would allow a cause-and-effect conclusion that applies to all 6,400?"
  2. 02Read what the question demands rather than what the designs offer. It asks for a cause AND a named population, which means both switches must be on. Any design missing either one is eliminated regardless of how careful the rest of it sounds.
  3. 03Switch 1 requirement: the participants must be drawn at random from the district's 6,400 fourth-graders — not recruited from teachers who volunteered, not taken from one school, not limited to classes that already use a reading block.

Solve alone

  1. 01Problem: "A company randomly selected 1,500 of the 62,000 patients listed in a national registry of adults with a particular chronic condition and invited them to take part in a trial. Of those invited, 900 agreed, and each of the 900 was randomly assigned to receive either a new drug or an existing drug. After six months, patients receiving the new drug reported fewer symptoms on average. State the strongest conclusion this design supports, and name the feature that stops you from stating a stronger one."

In your own words

In one sentence: why does random assignment neutralise confounding variables that nobody in the study ever measured or even thought of — and why does that same argument say nothing at all about whether the result applies to anyone outside the room?

Named traps

Causation from observation
Reading a pattern in recorded data as a cause. The tell is a verb: causes, lowers, improves, leads to, results in. It also arrives disguised as advice — "students who want higher grades should work fewer hours" is a causal claim wearing an imperative, and it is wrong for exactly the same reason. The defence is mechanical: if nobody was assigned anything, name one confounder out loud and the choice dies.
Comparison mistaken for assignment
Two groups appear in the description, so the study reads like an experiment. "Researchers compared students who take music lessons with students who do not." Nobody assigned anybody — the students sorted themselves, and everything that made them sort that way came along with them. Assignment means the RESEARCHER decided who got what. Comparing groups that already existed is observation with two columns.
Frame overreach
Generalising past the list the sample was actually drawn from. Registered voters are not adults, gym members are not the public, one university's students are not university students, and people reachable by landline at 2pm are a narrow slice of anyone. Read the noun that follows "randomly selected from" and treat it as a hard ceiling; every choice naming a wider group is cut without further thought.
Reflexive correlation-is-not-causation
The trap built specifically for strong students. Having learned that observational data cannot support a cause, the reader applies the rule to a genuine randomised experiment and refuses a conclusion the design has already earned — usually reasoning that the subjects were volunteers, which is switch 1 doing switch 2's job. When "randomly assigned" appears and there are two conditions, the causal verb is available. The only question left is who it applies to.
Big-n laundering
Treating sample size as a remedy for a broken selection procedure. Size shrinks random error and leaves systematic error exactly where it was, so a voluntary-response survey of 9,400 is not a better version of one with 94 — it is the same bias, estimated more precisely. Any answer choice whose reasoning turns on how large or how small the sample was is almost always naming the wrong defect.
One-axis elimination
The characteristic execution error at the top of the score range. Each of the four choices varies on TWO axes — strength of claim and breadth of population — and a reader who checks only one will find two survivors and pick the wrong one. The choice that gets the population right and the verb wrong is written to catch exactly this. Check both axes on every choice, every time; it costs about eight seconds.

The 800-level margin

At this point the two switches are automatic and the standard item is free. Six things are still moving points here, and five of them are ways a switch that appears to be on is actually off.

THE FRAME IS NOT THE POPULATION. Random selection generalises to the list, and the exam builds items where the list quietly excludes people who matter: registered voters when the claim is about adults, current customers when the claim is about consumers, students in an honours programme when the claim is about a school. This is UNDERCOVERAGE, and it is invisible unless you read the noun after "selected from" and compare it with the noun in the answer choice. Two choices will often be identical except for that noun.

NONRESPONSE UN-RANDOMISES A RANDOM SAMPLE. Selecting 2,000 people at random and hearing back from 240 does not give you a random sample of 240; it gives you a self-selected sample of 240 drawn from a random invitation list, which is a different object. The same applies to ATTRITION in an experiment — if a third of one group drops out before the end, the groups that finish are no longer the groups that were randomised. Whenever the passage reports both an invited number and a participating number, the gap between them is the item.

NO COMPARISON GROUP MEANS NO CAUSE, EVEN WITH PERFECT RANDOMISATION. "Two hundred randomly selected patients took the supplement for eight weeks, and their cholesterol fell." Both randomisation words are absent for a reason: nobody was assigned anything, because there was only one condition. Time is running alongside the treatment, and so are seasonal change, regression toward the mean, and the fact that people who know they are in a study behave differently. A before-and-after design on a single group licenses an association with time, and nothing more.

HEDGE CALIBRATION. Credited answers on this skill hedge, because studies support conclusions with hedges. Words that mark a wrong choice far more often than a right one: proves, guarantees, must, will, always, every, any, all, only. Words that mark a credited one: likely, evidence, suggests, tends to, association, among, about. This is a tie-breaker rather than a rule — do not select on wording alone when the design analysis has already settled the item — but on a choice set where three options overclaim, it is the fastest first cut available.

THE CONCLUSION IS ABOUT WHAT WAS MEASURED, AT THE LEVEL ADMINISTERED, OVER THE PERIOD RUN. Self-reported sleep is not sleep. Reaction time on one computer task is not alertness. A trial of one dose says nothing about double the dose, and six months says nothing about five years. A comparison against an existing drug ranks the new drug against that drug only — not against no treatment, and not against a third option nobody tested. And a study that finds no significant difference has not shown there is no effect; it has failed to detect one, which is a different sentence.

AVERAGES ARE NOT INDIVIDUALS. A group scoring six points higher on average does not mean any particular member gained six points, or gained anything at all. Answer choices that convert a mean difference into a promise about a person — "a student who takes the weekly quizzes will score six points higher" — are wrong on this ground alone, independently of whether the design licenses a cause.

EXECUTION ERRORS, which is where the last of the points actually live. Reading the study carefully and the question carelessly, so that a "which design would be appropriate" item gets answered as though it were a "which conclusion is supported" item. Missing the single word "volunteered" in a sentence that otherwise reads like a clean experiment. Seeing "random" and stopping, when the word was attached to the order the questions were asked in rather than to selection or assignment. And answering about the sample when the question named the population, or the population when it named the sample — the same wrong-quantity error that costs points across every other Math skill, in the one place on the exam where it involves no arithmetic at all.

Retrieval — with feedback on every choice

Question 1
Medium

A researcher selected 400 of the 6,200 employees at a large manufacturing firm at random and recorded, for each, the number of hours of sleep the employee reported getting on a typical night and the employee's score on a standard reaction-time test. Employees who reported more sleep tended to have faster reaction times.

Which of the following is the most appropriate conclusion based on this study?

Reference — not a study method, a lookup
EVALUATING STATISTICAL CLAIMS — reference card
Nothing is computed here. Two switches decide the answer; the finding is irrelevant.
SWITCH 1 — random SELECTION, before the study -> decides WHO the conclusion covers.
SWITCH 2 — random ASSIGNMENT, after you have the subjects -> decides whether you may say CAUSED.
Selection ON -> generalise to the list sampled FROM, and no further. OFF -> the people studied only.
Assignment ON (with a comparison group) -> "caused". OFF -> "is associated with".
Both ON: cause, generalisable. Selection only: association, generalisable.
Assignment only: cause, these subjects only. Neither: association, these subjects only.
The two are independent. Neither ever substitutes for the other.
Observational = the researcher only recorded. Experiment = the researcher imposed a treatment.
Two pre-existing groups compared is still observational. Comparison is not assignment.
Write the licensed sentence BEFORE reading the choices: [verb] ... among [ceiling population].
Eliminate on TWO axes: strength of claim, then breadth of population. Then cut anything under-claimed.
A recommendation ("they should ...") is a causal claim. Cut it from any observational study.
Bias is systematic; size only shrinks random error. A bigger biased sample is precisely wrong.
Read the noun after "randomly selected from" — registered voters are not adults.
Invited 2,000, heard from 240 -> a self-selected 240, not a random sample. Same for dropouts.
One group, before and after, no comparison -> no cause, however random the selection was.
Credited answers hedge: likely, evidence, tends to, association, about. Cut proves, guarantees, all, any, will.
A mean difference is not a promise about any individual.
The conclusion covers what was MEASURED, at the dose given, over the period run, against the treatment compared.
"No significant difference" is a failure to detect an effect, not proof there is none.

Every item on this page is Meridian-original, written to match the Digital SAT's format and difficulty — it is not a real SAT question. The only source that matches the live test exactly is College Board's own Bluebook and Question Bank.

Question 1Medium

A researcher selected 400 of the 6,200 employees at a large manufacturing firm at random and recorded, for each, the number of hours of sleep the employee reported getting on a typical night and the employee's score on a standard reaction-time test. Employees who reported more sleep tended to have faster reaction times.

Which of the following is the most appropriate conclusion based on this study?

  • AGetting more sleep causes employees at the firm to have faster reaction times.

    Nothing was assigned. The researcher recorded how much sleep employees were already getting, which means every characteristic that travels with sleep travels into the comparison too — shift pattern, age, caffeine, second jobs, children at home. Any one of those could produce this pattern with sleep doing no causal work at all.

  • BThere is an association between reported sleep and reaction time among all adults who work full time.

    The strength of the claim is right and the population is wrong. The 400 were drawn from the 6,200 employees at one manufacturing firm, so that is the frame and that is the ceiling. Full-time workers elsewhere had no chance of being selected, and a manufacturing workforce is not a random slice of them.

  • There is an association between reported sleep and reaction time among employees at this firm.

    Correct on both axes, which is how this item is checked. Switch 1: randomly selected from the 6,200 employees at this firm, so the conclusion generalises to employees at this firm and stops there. Switch 2: nothing assigned, so the verb is "is associated with". The licensed sentence is exactly this one, and note that it also correctly says REPORTED sleep — the study measured what employees said, not what they slept.

  • DEmployees at the firm who want faster reaction times should get more sleep.

    A recommendation is a causal claim in imperative clothing: advising the action only makes sense if the action produces the outcome, which is precisely what an observational study cannot establish. This choice is tempting because the advice is probably good in real life — but the question asks what THIS study supports, and the answer does not change because the conclusion happens to be true for other reasons.

Traps tested: Causal upgrade · Scope inflation · Recommendation as conclusion

Question 2Hard

A university researcher recruited 250 participants from among students who responded to a poster advertising a study on memory. Each participant was randomly assigned either to a group that studied a word list in 20-minute blocks spread over four days or to a group that studied the same list for 80 minutes in a single session. On a test administered one week later, the group that studied in spread-out blocks recalled significantly more words on average.

Which of the following is the most appropriate conclusion based on this study?

  • AStudying in blocks spread over four days causes better one-week recall for university students in general.

    The causal verb is earned; the population is not. The participants answered a poster, so they selected themselves, and no random selection from any stated group took place. Self-selected participants may well differ from students in general in motivation, in schedule, and in prior interest in how memory works — none of which the randomisation touches, because randomisation happened after they were already in the room.

  • BThere is an association between study schedule and one-week recall among the participants, but no cause-and-effect conclusion is possible because the participants were not randomly selected.

    This is the trap for a reader who knows the rules and has attached them to the wrong switch. Random selection governs WHO a conclusion covers; random assignment governs WHETHER it may be causal. Assignment is present here, with two conditions and a researcher deciding who got which, so the causal conclusion is available. Refusing it gives away a point the design already paid for.

  • CThe participants assigned to the spread-out schedule were more motivated than those assigned to the single session.

    Random assignment is the specific procedure that makes a systematic motivation gap unlikely — the coin flip does not know or care how motivated anyone is, so motivation is spread across both groups by chance. Nothing in the study measured motivation either, so this is an unmeasured explanation invented to replace one the design has already ruled out.

  • Studying in blocks spread over four days caused better one-week recall for the participants in this study.

    Correct, and the check is the two switches read separately. Switch 2 is on: participants were randomly assigned to one of two schedules, and the two groups differ systematically in nothing else, so the schedule is the only available explanation for the gap. Switch 1 is off: they answered a poster, so the sentence has to end at the participants in this study. Causal verb, narrow population — the combination that catches readers who assume the two always travel together.

Traps tested: Scope inflation · Selection required for causation · Confound invented despite assignment

Question 3Hardest on the test

To estimate the proportion of the 40,000 residents of a city who support a proposed transit levy, a newspaper posted the question on its website and invited readers to respond. Of the 9,400 people who responded, 71% expressed support. An analyst argued that because the sample was so large — nearly a quarter of the city's population — the result can be treated as a reliable estimate of city-wide support.

Which of the following best identifies the flaw, if any, in the analyst's reasoning?

  • AThe reasoning is sound: with 9,400 responses, the margin of error is small enough that the estimate is reliable.

    This is the sample-size fallacy stated plainly, and the is exactly the wrong tool to reach for. Margin of error describes sampling variability — how much a RANDOM sample of this size would bounce around the truth — and it is silent on whether the sampling procedure aims at the truth in the first place. A narrow interval around a biased estimate is a precise wrong answer.

  • BThe reasoning is flawed because 9,400 responses is too small a number to estimate a proportion in a population of 40,000.

    This objects to the size when the problem is the selection procedure, and it objects in the wrong direction as well — 9,400 is a very large sample by any standard, and a properly drawn random sample of 1,000 would serve this purpose comfortably. Naming the wrong defect leaves the real one untouched, and would suggest a remedy (collect more responses) that makes nothing better.

  • CThe reasoning is flawed because the newspaper did not randomly assign residents to support or to oppose the levy.

    Random assignment is impossible here and would not help if it were possible: you cannot assign someone an opinion, and the study is not trying to establish a cause. This choice is built for a reader who has memorised "random assignment" as the phrase that fixes studies and applies it without asking which switch the question is about. Assignment answers "is the treatment responsible?" — a question nobody asked.

  • The reasoning is flawed because respondents chose whether to reply, so the sample is not representative of city residents, and increasing the number of respondents does not remove that bias.

    Correct, and it names both halves of the problem. Voluntary response means the sample is defined by willingness to answer a newspaper's transit question, which is plausibly related to how strongly someone feels about transit — the sample selects on something connected to the very thing being measured. And bias is a property of the procedure, so it does not shrink with size: 9,400 self-selected respondents make the same systematic error as 94, just with a tighter interval around it.

Traps tested: Sample size fixation · Random assignment misapplied to survey

Question 4Hardest on the test

A county health office wants to estimate the proportion of adults living in the county who have received a particular vaccine. It selects 500 names at random from the county's list of registered voters and successfully contacts all 500. Of those contacted, 62% report having received the vaccine.

Which of the following is the most appropriate conclusion based on this study?

  • AIt is likely that about 62% of adults living in the county have received the vaccine.

    The selection was random, the response rate was perfect, and the conclusion is still one noun too wide. The draw was made from the list of registered voters, so adults in the county who are not registered had zero chance of appearing — and registration is associated with age, length of residence and engagement with public institutions, all of which are plausibly related to vaccination. Random selection generalises to the frame, and the frame here is voters.

  • It is likely that about 62% of registered voters in the county have received the vaccine.

    Correct. Switch 1 is on and its ceiling is the list actually sampled: registered voters in this county. Switch 2 is off — nothing was assigned and no cause is claimed — so a descriptive estimate is the right shape of conclusion. Note the hedges: "likely" and "about", which are what a sample supports about a population. How wide the "about" is belongs to the neighbouring skill point on inference from sample statistics; what this item tests is which population the estimate is about at all.

  • CRegistering to vote makes an adult more likely to receive the vaccine.

    A causal verb from a survey that assigned nothing to anybody. Even if registered voters really are vaccinated at higher rates, the explanation would be everything that travels with registration — age, stable housing, contact with public services — rather than the act of registering. This choice also compares registered voters against a group the study never measured.

  • DNo conclusion is possible, because 500 people is too small a sample for a county.

    Five hundred randomly selected people is an ordinary and perfectly serviceable sample for estimating a proportion — national surveys routinely run on about a thousand. The defect in this design is which list the 500 came from, not how many of them there were, and objecting to the size points at the one feature that is not the problem.

Traps tested: Frame mistaken for population · Causal upgrade · Sample size fixation

Question 5Hard

A nutritionist wants to determine whether a new meal-replacement drink lowers resting heart rate in adults aged 40 to 60, and wants the result to apply to all adults aged 40 to 60 in her region. Which of the following study designs would allow both a cause-and-effect conclusion and a conclusion about that entire age group in the region?

  • ASelect 300 adults aged 40 to 60 at random from the region and record, for each, whether the adult already drinks the product and his or her resting heart rate.

    This design has switch 1 and not switch 2. Random selection from the region licenses a conclusion about adults aged 40 to 60 in the region, which is half of what was asked — but nobody is assigned anything, so the adults who already drink the product chose to, and everything else that goes with that choice comes along. The result would be a generalisable association and no cause.

  • BRecruit 300 volunteers aged 40 to 60 from a local fitness centre and randomly assign half to drink the product daily and half to drink a similar-tasting drink without the active ingredients.

    This design has switch 2 and not switch 1, which is the mirror image of the previous choice. The random assignment and the comparison drink make the causal conclusion clean, but the participants are volunteers from one fitness centre — a group whose resting heart rates are unlikely to represent the region's 40-to-60-year-olds. Valid cause, wrong population, and the question asked for both.

  • Select 300 adults aged 40 to 60 at random from the region and randomly assign half to drink the product daily and half to drink a similar-tasting drink without the active ingredients.

    Correct, and it is the only choice containing both required phrases: "select ... at random from the region" turns switch 1 on with a frame that matches the target population, and "randomly assign half ... half" turns switch 2 on with a comparison group. That combination is the only one that licenses a causal conclusion generalisable to the region's adults aged 40 to 60 — which is exactly what the question demanded, and why reading the demand before the designs saves the item.

  • DSelect 300 adults aged 40 to 60 at random from the region, have all 300 drink the product daily for eight weeks, and compare their resting heart rates before and after.

    Everything here sounds controlled, and the word "randomly" appears — attached to selection only. With one group and no comparison, there is nothing to assign anyone to, so switch 2 is off. Whatever change appears over eight weeks is confounded with the eight weeks: seasonal change, the effect of knowing one is being measured, and any other habit shift that accompanies joining a study. Random assignment requires at least two conditions.

Traps tested: Causal upgrade · Assignment without selection · No comparison group

Question 6Hardest on the test

A researcher randomly selected 1,200 of the 45,000 undergraduates enrolled at a state university and randomly assigned each of them to one of two versions of the same online course: one version included a short quiz every week, and the other did not. All 1,200 completed the course and sat the same final examination, on which students in the weekly-quiz version scored 6 points higher on average.

Which of the following is the most appropriate conclusion based on this study?

  • Including weekly quizzes caused an increase in average final-examination scores among undergraduates at this university.

    Correct, and it is the rare item where both switches are on. Switch 2: students were randomly assigned to the two versions, so the weekly quiz is the only systematic difference between the groups and the causal verb is licensed. Switch 1: the 1,200 were randomly selected from the 45,000 undergraduates at this university, so the conclusion generalises to that frame — and stops there. Note the word "average", which keeps the claim at the level the evidence supports.

  • BIncluding weekly quizzes caused an increase in average final-examination scores among undergraduates at every university.

    The design is strong and this claim is still one step past its frame. The draw was made from one state university's 45,000 undergraduates; students elsewhere had no chance of selection, and this course, this examination and this student body are all particular. A frame is a hard ceiling even when both randomisations are present — being able to say "caused" never widens who it was caused for.

  • CIncluding weekly quizzes guarantees that any undergraduate at this university will score 6 points higher on the final examination.

    Two separate overstatements in one sentence. A difference in group MEANS says nothing about any individual — some students in the quiz version certainly scored lower than some in the other. And "guarantees" plus "any" converts a statistical result into a promise, which no study of this kind supports. This choice gets both switches right and fails on the arithmetic of averages.

  • DThere is an association between weekly quizzes and final-examination scores among undergraduates at this university, but the design does not support a cause-and-effect conclusion.

    This is the correlation-is-not-causation reflex firing on a genuine randomised experiment. The rule it invokes applies to observational studies, where subjects sort themselves into groups; here the researcher assigned them, which is the specific condition under which "caused" is permitted. The population half of this choice is right, which is what makes it survive a one-axis check — and the verb half gives away a conclusion the design has already earned.

Traps tested: Scope inflation · Average read as individual · Reflexive noncausation

Meridian · progress saved in this browser

Up next

Problem-Solving & Data Analysis — module quiz

Interleaved across all seven skills.

1 min