Probability and conditional probability

Two-way tables, where the whole skill is choosing the right denominator.

~30 min · prequestion, worked examples, retrieval practice

Every probability question on this test is a single fraction, and the numerator is almost never the hard part — the entire skill is deciding which population the question is drawing from, because that is the denominator. A two-way table hands you three legitimate denominators for the very same cell: the grand total, its row total, and its column total. The words "given that," "among," and "of the students who" pick one of the three, and the test writes a distractor for each of the other two. Get the denominator wrong and every step afterwards is flawless and the answer is still wrong.

Before you read on

Two or three questions on exactly what this lesson teaches. Being wrong here is fine — it's the fastest way to find out what to pay attention to next.

Question 1
Medium

TABLE. "Commuting method by age, 200 employees at one firm." Employees who drive: 45 are under 40, 105 are 40 or older, row total 150. Employees who take public transit: 30 are under 40, 20 are 40 or older, row total 50. Column totals: 75 employees are under 40, 125 are 40 or older. Grand total: 200.

Before any teaching. One of the 200 employees is selected at random from those who take public transit. What is the probability that the selected employee is under 40?

Question 2
Medium

Which wording in a probability question tells you the denominator is one row or one column of the table rather than the grand total?

Question 3
Hard

At a school, 90% of the students who play a varsity sport also take a music course. Does it follow that 90% of the students who take a music course also play a varsity sport?

What the question is actually testing

Problem-Solving and Data Analysis is about 15% of the Math section — that is College Board's published domain weight, and it works out to roughly six or seven of the 44 Math questions. Probability is one of the domain's seven skill points. How the questions divide among those seven is not published, so the "one or two probability questions per test" figure that circulates in prep materials is an estimate drawn from released item pools, not an official statement. What is not in dispute is the format: the overwhelming majority of probability items on the digital SAT are built on a two-way frequency table.

A two-way frequency table classifies one group of people or things by two traits at once. The rows are the categories of one trait, the columns are the categories of the other, each interior cell holds the count of members with that row's trait and that column's trait, and the margins hold the totals. Every count you need is already printed. Nothing has to be multiplied, no formula from a statistics course is required, and the arithmetic is one division.

So the item cannot be testing whether you can count. It is testing one decision: which total goes underneath. A single interior cell can legitimately sit over three different denominators — its row total, its column total, or the grand total — and each of the three answers a genuinely different question about the same people. The test knows this, which is why a well-built probability item puts all three numbers in front of you as choices.

It is worth saying plainly how narrow the scope is, because probability is a large subject and almost none of it is here. Conditional probability on the digital SAT means reading a denominator off a table. There is no multiplication rule to memorise, no independence formula to apply, no permutations or combinations, and no probability distributions. You will never be asked to apply Bayes' theorem by name — although the reversed-conditional questions in this lesson are exactly what it computes, done by counting instead.

Foundations — probability from zero (skip this block if "of the 50 who took the bus, 30 were late" already reads as a fraction)

If the phrase "of the 50 employees who took the bus, 30 were late" already reads to you as 30/50, skip ahead; the rest of the page is one idea applied repeatedly and you have it. If probability still feels like a topic with rules in it, this is the highest-value block here, because there are no rules — there is one fraction.

A probability is a fraction: how many outcomes count, over how many outcomes there are to draw from. That is the whole definition. It runs from 0 (never happens) to 1 (always happens), so a value outside that range is an arithmetic error rather than a strange result. A number above 1 nearly always means two overlapping groups were added; a negative one means a subtraction went the wrong way.

"Selected at random" is doing one specific job: it means every member of the pool is equally likely, which is exactly what allows the answer to be one count divided by another rather than a weighted average. It is not a hint about which pool.

Fraction, decimal, and percent are the same number in three costumes: 30/50, 0.6, and 60%. Divide to get the decimal (30 ÷ 50 = 0.6), multiply by 100 to get the percent. The test specifies which it wants — in the choices, or in a phrase like "to the nearest hundredth" — and the right number in the wrong form still loses the point.

The complement: P(not A) = 1 − P(A). If 0.6 of a group were late, 0.4 were not. This is the fast route whenever the trait you want spans more cells than the trait you do not want — count the small side and subtract.

Now the table itself, using the one from the prequestion. Two hundred employees, each classified twice: by how they commute (drives, or takes transit) and by age (under 40, or 40 or older). Four interior cells hold four counts — 45 drive and are under 40, 105 drive and are 40 or older, 30 take transit and are under 40, 20 take transit and are 40 or older. Add across a row and you get how many drive (150) and how many take transit (50). Add down a column and you get how many are under 40 (75) and how many are 40 or older (125). Both directions must land on the same grand total of 200, and checking that they do is the cheapest error-detector in the skill: if a row does not sum to its printed total, a cell has been misread.

Three kinds of question live in those four cells, and they differ only in the denominator. A joint probability — "the probability that a randomly selected employee takes transit and is under 40" — is one cell over the grand total: 30/200 = 0.15. A marginal probability — "the probability that a randomly selected employee takes transit" — is one row total over the grand total: 50/200 = 0.25. A conditional probability — "given that the employee takes transit, the probability that the employee is under 40" — is one cell over that row's total: 30/50 = 0.6. Same table, three answers, and two of them are built from the very same 30 people.

The word "conditional" carries no extra machinery. It does not name a different kind of probability with different rules; it names a probability taken after part of the table has been discarded. That is why is a denominator skill and not a formula skill.

The method — circle the denominator before you look for the numerator

One protocol, run in this order, every time. It costs about eight seconds and it is the entire skill.

Step 1 — find the conditioning phrase and underline it. The vocabulary is small and it repeats: "given that," "among," "of the," "from those who," "if the selected student is." If one is present, the group it names becomes the new population. If none is present, the population is everyone in the table.

Step 2 — write the denominator down before you go looking for the numerator. This inverts the order almost everyone uses, and inverting it is the point. The numerator is easy and it is a magnet: once a cell is in your hand you will divide it by whichever total your eye lands on, and on a printed table the grand total is the one that catches the eye. Commit to the pool first and the numerator has nowhere wrong to go.

Step 3 — the numerator is the count of members inside that denominator who also have the trait the question asks about. Inside is the operative word. If the denominator is a row total, the numerator can only be assembled from cells in that row.

Step 4 — check two things before you write anything down. The numerator must be less than or equal to the denominator, and the value must lie between 0 and 1. A numerator drawn from outside its own pool fails the first check almost every time it happens.

Step 5 — reread the last line of the question. It may want a percent when you produced a decimal, a count of people when you produced a fraction, or the other conditional. This step recovers more points near the top of the score range than any amount of additional technique.

The phrase-to-denominator map, stated once so it never has to be re-derived under time pressure. "A randomly selected member is A and B" → the A-and-B cell, over the grand total. "A randomly selected member is A" → A's row (or column) total, over the grand total. "Given that the member is B, the probability that the member is A" → the A-and-B cell, over B's total. "Of the members who are B, what fraction are A" → identical to the previous one. The SAT alternates those last two wordings freely and they mean exactly the same thing.

One more pattern, because it uses a different rule: "A or B" always means A, or B, or both — the inclusive reading, on every item this test writes. Count it as (A's total) + (B's total) − (the cell where they cross), over the grand total. Subtracting the shared cell is not a refinement to apply when you have time; skipping it counts those members twice and can push the answer above 1.

Mechanism

Why the numerator survives the condition and the denominator does not

A probability is not a property of an event. It is a ratio between an event and a population, and "given that the employee takes transit" does not add information to the old ratio — it replaces the population. Before the condition there were 200 people to draw from; after it there are 50, and the other 150 have not become less likely, they have stopped existing for the purposes of this question. Meanwhile the people who are both transit riders and under 40 are the same 30 people either way. That single asymmetry explains everything else on this page. It is why P(A given B) and P(B given A) share a numerator and differ only in denominator: the overlap is one group of people, and the two questions divide it by two different populations, so they agree only when those two populations are the same size. It is why a conditional can sit far away from the unconditional in either direction — narrowing to transit riders raised the probability of being under 40 from 75/200 = 0.375 to 30/50 = 0.6, because transit riders skew young in this table. And it is why the complement has to be taken inside the conditioned pool: 1 − P(A) computed over 200 people is not the complement of a fraction whose denominator is 50, and subtracting it from 1 produces a number that is not a probability of anything. The case where the conditional does not move is worth naming, because the test asks about it in words rather than symbols: when the conditional equals the unconditional, the condition carried no information and the two traits are independent. In a table that shows up as every row having the same internal split as the grand total. When the splits differ, the traits are associated, and the size of the gap between one row's split and another's is the whole of the evidence for it.

Worked examples

Fully worked — one cell, three denominators

  1. 01TABLE. "Flowering by fertilizer, 300 seedlings." Fertilizer A: 84 flowered, 56 did not, row total 140. Fertilizer B: 126 flowered, 34 did not, row total 160. Column totals: 210 flowered, 90 did not. Grand total 300. QUESTION: one of the 300 seedlings is selected at random. If the selected seedling received Fertilizer A, what is the probability that it flowered?
  2. 02Step 1 — find the conditioning phrase. "If the selected seedling received Fertilizer A" is a condition even though the word "given" never appears. The Fertilizer B row is now irrelevant: 160 seedlings have left the problem.
  3. 03Step 2 — denominator first, before looking at anything else: the Fertilizer A row total, 140.
  4. 04Step 3 — numerator: inside those 140, the ones that flowered. That is the cell where the Fertilizer A row meets the Flowered column: 84.
  5. 05Step 4 — 84/140. Divide top and bottom by 28: 3/5 = 0.6. Checks: 84 is less than 140, and 0.6 lies between 0 and 1.
  6. 06Step 5 — reread the last line. It asked for a probability, and 3/5 is one. Answer: 3/5.
  7. 07Worth seeing once, because it is the whole lesson in one line: the same 84 over the other two denominators. 84/210 = 2/5 answers "given that it flowered, what is the probability it received Fertilizer A." 84/300 = 0.28 answers "what is the probability it both received Fertilizer A and flowered." Three questions, one cell, three answers — 0.6, 0.4, 0.28 — and a well-built item offers all three.

One step hidden — a complement taken inside the pool

  1. 01TABLE. "E-book borrowing by membership type, 480 library patrons in one month." Student members: 126 borrowed, 54 did not, row total 180. Adult members: 150 borrowed, 150 did not, row total 300. Column totals: 276 borrowed, 204 did not. Grand total 480. QUESTION: a patron is selected at random from those with a student membership. What is the probability that the patron did not borrow an e-book that month?
  2. 02Step 1 — "from those with a student membership" is the condition, so the denominator is the student row total, 180. Not 480, and not the 204 in the "did not borrow" column.
  3. 03Step 2 — the trait wanted is a negative one, which gives two equivalent routes to the numerator. Count it directly from the student row: 54. Or take the complement inside that same row: 180 − 126 = 54, which is the same statement as 1 − 126/180.
  4. 04Step 3 — note why the complement route is legal here. It was taken inside the conditioned pool. Computing 1 − 276/480 instead would be the complement over the whole table, and it answers nothing that was asked.

Two steps hidden — finish the table, then pool two columns

  1. 01TABLE, with two entries missing. "Response to a proposed schedule change, 300 students." Grade 11: Support is missing, Oppose 45, No opinion 15, row total 120. Grade 12: Support 96, Oppose is missing, No opinion 24, row total 180. Column totals: Support 156, Oppose 105, No opinion 39. Grand total 300. QUESTION: one student is selected at random from those who did not support the change. What is the probability that the student is in Grade 11?
  2. 02Fill the missing cells before touching the question, using the rule that every row sums to its printed total. Grade 11 Support = 120 − 45 − 15 = 60. Grade 12 Oppose = 180 − 96 − 24 = 60. Then verify each one against the other direction rather than trusting the subtraction: 60 + 96 = 156 ✓ and 45 + 60 = 105 ✓.
  3. 03"Those who did not support the change" is not a column — it is two columns pooled, Oppose and No opinion. Denominator = 105 + 39 = 144, and the cross-check confirms it: 300 − 156 = 144.

Solve alone

  1. 01PROBLEM. TABLE. "Reported side effect by age group, 640 patients." Under 18: 36 reported a side effect, 84 did not, row total 120. Ages 18 to 64: 84 reported, 276 did not, row total 360. Ages 65 or over: 60 reported, 100 did not, row total 160. Column totals: 180 reported a side effect, 460 did not. Grand total 640. QUESTION: one patient is selected at random from those who reported a side effect. What is the probability that the patient is not in the 18-to-64 age group?

In your own words

In one sentence: when a question adds "given that the patient reported a side effect," why does the numerator stay the same while the denominator changes — and what does that fact tell you about why P(A given B) and P(B given A) are usually different numbers?

Named traps

Grand-total denominator
The condition was read and then ignored: the right cell is found and divided by the table's grand total instead of the conditioned group's total. It is the highest-frequency wrong answer on this skill, the test writes a distractor for it on essentially every conditional item, and it is always smaller than the correct answer — which makes "that came out surprisingly low" a usable alarm.
Reversed conditional
Computing P(A given B) when the question asked for P(B given A). Same cell on top, the other margin underneath. The tell is grammatical, not mathematical: whatever noun follows "given that," "among," or "of the" is the denominator, and whatever noun the question then asks about supplies the numerator. Read those two nouns in order, out loud, before touching the numbers.
Joint mistaken for conditional
"The probability that a randomly selected student is a senior and opposes the change" is one cell over 300. "The probability that a senior opposes the change" is the same cell over the senior total. Nothing but the wording separates them, both wordings appear on the same test, and the word doing the work is a single "and" versus a single "a."
Complement taken in the wrong pool
Using 1 − P(A) computed over the grand total when the question needed 1 − P(A given B) computed inside the row. A complement is always taken inside whatever population the question has already narrowed to; taken outside it, the two fractions do not even share a denominator, so the subtraction is meaningless before it is wrong.
Pooled-category miss
"Did not support," "not in the 18-to-64 group," and "either biology or chemistry" each span more than one row or column, and the fraction needs every cell in the span — on the top and on the bottom, wherever the span applies. Taking a single cell where a category was pooled is the standard way a hard item is lost by a student who understood it perfectly.
Overlap double-count on "or"
Adding a row total to a column total to answer "A or B." The cell where they cross belongs to both groups and gets counted twice, so the answer comes out too large and can exceed 1. Subtract the shared cell exactly once: (row total) + (column total) − (shared cell), over the grand total.

The 800-level margin

Once the denominator rule is automatic, the points left in this skill sit in six places: tables you have to finish building, conditionals stated as percentages of a subgroup, counts asked for instead of probabilities, relative-frequency tables, rate-versus-count comparisons, and a short list of execution errors that survive knowing every rule above.

Incomplete tables. The hard version of this item prints half the entries and hides a constraint in the surrounding sentence, and the reconstruction is the question. Two rules recover everything: every row and every column sums to its printed total, and the row totals and the column totals must both sum to the same grand total. Fill from whichever line has exactly one blank left, then verify each filled cell against the other direction — subtracting once and trusting it is how a single slip propagates into a wrong answer that looks entirely reasonable. When the constraint arrives as a percentage ("40% of the students who chose the morning session were seniors"), convert it to a count immediately and write the count into the table. Percentages left as percentages get applied to the wrong base a few lines later.

Conditionals stated as percentages of a subgroup. "Seventy percent of the 400 volunteers completed the training; 30% of those who completed it and 55% of those who did not took the follow-up survey" is a two-way table written in prose. Build it: 280 completed and 120 did not; 0.30 × 280 = 84 and 0.55 × 120 = 66 took the survey; 150 took the survey in total. Every percentage applies to its own subgroup and never to the whole. Once the counts are on paper the item is an ordinary denominator question again — "given that a volunteer took the survey, what is the probability the volunteer completed the training" is 84/150 = 0.56. The characteristic error here is not arithmetic. It is answering 0.30, which is the reversed conditional, and which is sitting in the choices precisely because building the table correctly and then reading the wrong direction off it is what a strong student does under time pressure.

Counts, not probabilities. A large share of the hardest items in this skill end with "how many" rather than "what is the probability," and every step before the last line is identical. Two failure modes, both of which produce a number that looks like an answer: reporting the fraction when a count was asked for, and stopping at a partial count — the subgroup you happened to compute rather than the total across subgroups. The direction is worth stating explicitly: a probability multiplied by the size of the population it was measured on gives the count. Multiply by any other population and the answer is wrong by exactly the ratio between the two.

Relative-frequency tables. Sometimes the table holds proportions or percentages instead of counts, which means a denominator has already been applied — and the question is which one. If each row sums to 1 (or to 100%), the entries in a row are already conditional probabilities given that row, and the answer can be read straight off; the column conditionals are not recoverable without the underlying counts. If the whole table sums to 1, every entry is a joint probability and the conditionals still have to be built by dividing by a row or column sum. Check what the table sums to before reading a single value out of it, because the same printed number means different things in the two cases.

Rates versus counts. When two groups are different sizes, the group with more members having a trait can be the group with the lower rate of it. Two clinics recording 414 successes out of 460 and 132 out of 150 have the first ahead on both count and rate — but change 132 to 141 and the counts still favour the first clinic (414 to 141) while the rates reverse (0.90 to 0.94). Any comparison across groups of unequal size has to be made on the conditional probabilities, never on the cell counts, and items are built specifically to make the raw count the tempting answer to a question about likelihood.

Independence, when it is asked in words. If a conditional equals the corresponding unconditional — a row's internal split matches the table's overall split — the condition carried no information and the two traits are independent. The test rarely uses the word. It asks whether the data "suggest an association," and the answer is decided by comparing one row's conditional probability against another row's, not by comparing counts across rows of different sizes. Note also the ceiling on what any two-way table can support: an association between two survey columns is not evidence that one caused the other, and no arrangement of counts becomes a causal claim without random assignment.

Execution errors, which is where the last few points actually live. Answering the right question about the wrong quantity — the reversed conditional, the count when the probability was wanted, the complement when the trait itself was wanted. Rounding before the final step: keep 8/15 as a fraction and convert once at the end, because 0.53 entered where 0.533 was required is marked wrong. Reading a cell from the wrong row when a table has three or more rows and the question lists them in a different order than the table does. And the format rules, which are College Board's own published directions and not prep-industry folklore: a fraction is accepted, an unreduced fraction is accepted, symbols such as the percent sign are not, and a mixed number is not — five halves is entered as 5/2 or as 2.5, never as 2 1/2, which the scoring reads as twenty-one halves. If the question asks for a percent, the grid takes the number of percent (60), not the decimal (0.6), so the last line decides the format as well as the quantity.

Retrieval — with feedback on every choice

Question 1
Medium

TABLE. "Fitness class attendance in one week, 400 community centre members." Members under 30: 72 attended, 48 did not, row total 120. Members 30 or over: 108 attended, 172 did not, row total 280. Column totals: 180 attended, 220 did not. Grand total 400.

Given that a randomly selected member is under 30, what is the probability that the member attended a fitness class that week?

Reference — not a study method, a lookup
PROBABILITY AND CONDITIONAL PROBABILITY — reference card
Every answer is (how many count) / (how many you are drawing from). Find the second one FIRST.
Denominator words: given that / among / of the / from those who / if the selected X is Y.
No condition stated -> the denominator is the grand total.
Condition names a row -> that row's total. Names a column -> that column's total.
The numerator is counted INSIDE the denominator's pool. Nowhere else.
One cell, three answers: cell/row, cell/column, cell/grand total. All three will be offered.
P(A given B) and P(B given A) share a numerator; equal only if the two totals are equal.
Complement: P(not A) = 1 - P(A), taken inside the conditioned pool, never over the whole table.
"A or B" = A's total + B's total - the shared cell, over the grand total. Subtract the overlap once.
Pooled categories ("did not support", "not 18-64") span several cells — add every one, top and bottom.
A percent of a subgroup applies to that subgroup only. Convert it to a count and write it in the table.
Missing cells: every row and column sums to its total, and both directions reach the same grand total.
Relative-frequency table: check what it sums to (each row = 1, or the whole table = 1) before reading.
Compare groups of unequal size by rate, never by cell count.
Independent = the row's split matches the table's split. Association is never causation without random assignment.
Asked for a count, not a probability? Probability x the size of the population it was measured on.
Before answering: numerator <= denominator, value between 0 and 1, and reread the last line.

Every item on this page is Meridian-original, written to match the Digital SAT's format and difficulty — it is not a real SAT question. The only source that matches the live test exactly is College Board's own Bluebook and Question Bank.

Question 1Medium

TABLE. "Fitness class attendance in one week, 400 community centre members." Members under 30: 72 attended, 48 did not, row total 120. Members 30 or over: 108 attended, 172 did not, row total 280. Column totals: 180 attended, 220 did not. Grand total 400.

Given that a randomly selected member is under 30, what is the probability that the member attended a fitness class that week?

  • A0.18

    This is 72/400: the right cell over the grand total. It answers "what is the probability that a randomly selected member is under 30 and attended" — a joint probability, with no condition applied. "Given that a randomly selected member is under 30" has already removed the 280 members aged 30 or over from the pool, so they cannot remain in the denominator.

  • B0.30

    This is 120/400, the probability that a member is under 30 at all. It reports the probability of the condition instead of the probability of the thing being asked about, so attendance — the entire subject of the question — never enters the calculation.

  • C0.40

    This is 72/180: the same 72 members, divided by the "attended" column total. It answers the reversed question — "given that a member attended, what is the probability the member is under 30." Same numerator, wrong pool. The noun that follows "given that" is under 30, so the under-30 total is the denominator.

  • 0.60

    Correct. "Given that a randomly selected member is under 30" fixes the pool at the under-30 row, total 120. Inside that row, 72 attended. 72/120 = 0.6. Checks: 72 is less than 120, and 0.6 lies between 0 and 1.

Traps tested: Grand total denominator · Marginal for conditional · Reversed conditional

Question 2Hard

TABLE. "Housing status by year of study, 900 university students." First-year students: 252 live on campus, 36 live off campus, 72 live with family, row total 360. Upper-year students: 216 live on campus, 252 live off campus, 72 live with family, row total 540. Column totals: 468 on campus, 288 off campus, 144 with family. Grand total 900.

One student is selected at random from those who do not live on campus. What is the probability that the selected student is a first-year student?

  • 1/4

    Correct. "Do not live on campus" pools two columns: 288 + 144 = 432, confirmed by 900 − 468 = 432. Inside that pool, the first-year students are also spread across two cells: 36 + 72 = 108. 108/432 = 1/4 = 0.25. Checks: 108 is less than 432, and the value lies between 0 and 1.

  • B3/25

    This is 108/900: the right numerator over the grand total. The pooling was done correctly and the conditioning was not — the 468 students who do live on campus are still in the denominator, even though the question removed them before the draw.

  • C3/10

    This is 108/360, the same 108 students divided by the first-year row total. It answers the reversed question — "given that a student is a first-year, what is the probability the student does not live on campus." The question conditions on housing and asks about year, not the other way round.

  • D1/8

    This is 36/288, which treats "does not live on campus" as the off-campus column alone. Living with family is also not living on campus, so that category belongs in the pool — and it belongs on both the top and the bottom, which is why dropping it changes the numerator from 108 to 36 as well as the denominator from 432 to 288.

Traps tested: Grand total denominator · Reversed conditional · Pooled category miss

Question 3Hard

At a conference with 500 attendees, 60% are researchers and the rest are industry practitioners. Of the researchers, 40% attended the opening keynote; of the practitioners, 65% attended it. An attendee is selected at random from those who attended the keynote. What is the probability that the selected attendee is a researcher?

  • A0.24

    This is 120/500 — the researchers who attended, over all 500 attendees. It answers "what is the probability that a randomly selected attendee is a researcher and attended the keynote," a joint probability with no condition applied. The 250 attendees who skipped the keynote were removed by the question and cannot stay in the denominator.

  • B0.40

    This is the rate given in the passage: 40% of the researchers attended. That is P(attended given researcher) — the reversed conditional. The question conditions on attendance and asks about occupation, which is the other direction, and the two agree only if the number of researchers equals the number of attendees at the keynote (300 versus 250 here, so they do not).

  • 0.48

    Correct. Build the table first: 0.60 × 500 = 300 researchers and 200 practitioners. Keynote attendance: 0.40 × 300 = 120 researchers and 0.65 × 200 = 130 practitioners, so 250 attended in total. Conditioning on attendance fixes the denominator at 250, and the numerator is the researchers inside it: 120/250 = 0.48. Check: 120 is less than 250, 130 + 120 = 250 matches the pool, and 0.48 lies between 0 and 1.

  • D0.60

    This is 300/500, the probability that any attendee is a researcher, which is the unconditional figure the passage opens with. It ignores the condition entirely. It is also the tell for a real result: the conditional came out at 0.48 rather than 0.60 because practitioners attended at a higher rate, so knowing someone was at the keynote makes them less likely to be a researcher, not equally likely.

Traps tested: Grand total denominator · Reversed conditional · Marginal for conditional

Question 4Hardest on the test

TABLE, partially filled. "Streaming subscription and broadband access, 800 households." Households that subscribe to a streaming service: 440 have broadband, the number without broadband is not shown, and the row total is not shown. Households that do not subscribe: the number with broadband is not shown, 90 do not have broadband, and the row total is not shown. Column totals: 600 households have broadband, the no-broadband total is not shown. Grand total: 800.

One household is selected at random from those that do not have broadband. What is the probability that the selected household subscribes to a streaming service?

  • 11/20

    Correct. Rebuild the table: no-broadband total = 800 − 600 = 200; subscribers without broadband = 200 − 90 = 110; non-subscribers with broadband = 600 − 440 = 160; row totals 550 and 250, which sum to 800 ✓. Conditioning on "do not have broadband" fixes the denominator at 200, and the subscribers inside it number 110. 110/200 = 11/20 = 0.55.

  • B9/20

    This is 90/200: the correct denominator with the other category on top. It answers "what is the probability the household does not subscribe" — the complement of what was asked. The reconstruction was done correctly and the final trait was read off the wrong row, which is why this distractor is exactly 1 minus the credited answer.

  • C1/5

    This is 110/550: the same 110 households over the subscriber row total instead of the no-broadband column total. It answers the reversed question — "given that a household subscribes, what is the probability it has no broadband." The condition here is broadband, not subscription.

  • D11/80

    This is 110/800: the right numerator over the grand total. It answers the joint question, "what is the probability that a household both subscribes and lacks broadband." The 600 broadband households were removed by the condition and are still being counted here.

Traps tested: Complement answered · Reversed conditional · Grand total denominator

Question 5Hardest on the test

TABLE. "Regional science fair results, 250 projects." Biology: 28 won an award, 62 did not, row total 90. Chemistry: 22 won an award, 38 did not, row total 60. Physics: 30 won an award, 70 did not, row total 100. Column totals: 80 projects won an award, 170 did not. Grand total 250.

One project is selected at random. What is the probability that the selected project is a biology project or won an award?

  • A0.112

    This is 28/250, the biology projects that won an award. That is the "and" question — a project carrying both traits at once — and it is the intersection rather than the union. "Or" on this test always means at least one of the two, so the 62 biology projects without awards and the 52 award winners outside biology all belong in the numerator too.

  • B0.456

    This is 114/250, from 62 + 22 + 30 — the biology projects without awards plus the non-biology award winners. That is the exclusive reading of "or": exactly one of the two traits. The digital SAT uses the inclusive reading, so the 28 projects that are biology and won an award are counted, once.

  • 0.568

    Correct. Union: (biology total) + (award total) − (biology and award) = 90 + 80 − 28 = 142, over 250, which is 0.568. Verify by direct count instead of formula: 90 biology projects, plus the award winners outside biology, 22 + 30 = 52, gives 142 as well. The subtraction is exactly the 28 that would otherwise be counted in both groups.

  • D0.680

    This is 170/250, from 90 + 80 with no subtraction. The 28 projects that are biology and won an award sit in both totals and are therefore counted twice. Adding two margins without removing the cell where they cross is the standard failure on union items, and the resulting figure is always too large by exactly that shared cell over the grand total.

Traps tested: Joint for union · Exclusive or read · Overlap double count

Question 6Hardest on the test

A bookstore recorded 720 customer visits in one week. The probability that a randomly selected visit involved a loyalty card is 5/12. Given that a visit involved a loyalty card, the probability that a discounted title was bought is 3/5. Among the visits that did not involve a loyalty card, 84 included the purchase of a discounted title. How many of the 720 visits included the purchase of a discounted title?

  • A180

    This is the loyalty-card half of the answer and it stops there: (5/12)(720) = 300 loyalty-card visits, and (3/5)(300) = 180 of them included a discounted title. Correct as far as it goes, but the question asks for all 720 visits, and the 84 discounted purchases stated for the non-card visits have not been added.

  • 264

    Correct. Loyalty-card visits: (5/12)(720) = 300, leaving 420 without a card. The conditional 3/5 was measured on the card visits, so it multiplies 300, not 720: (3/5)(300) = 180. The non-card figure is given directly as a count, 84. Total = 180 + 84 = 264. Check the fit: 264 is less than 720, and 180 is less than the 300 it came from.

  • C384

    This is 300 + 84 — the loyalty-card row total added to the non-card discounted count, instead of the loyalty-card discounted count. It treats every one of the 300 card visits as a discounted purchase, which discards the 3/5 the problem supplied. A row total and a cell are different quantities and the conditional is what converts one into the other.

  • D432

    This is 180 + 252, where 252 = (3/5)(420): the card holders' discount rate applied to the non-card visits as well. The problem gives a separate figure for that group — 84 — precisely because the two groups behave differently. A conditional probability is only valid on the population it was measured on, and transplanting it to another group is the same error as multiplying by the wrong total.

Traps tested: Partial total · Row total for cell · Rate transplanted across groups

Meridian · progress saved in this browser

Up next

Inference from sample statistics and margin of error

SAT-only content. What a margin of error does and, more importantly, does not mean.

30 min