Math · Checkpoint 20 of 30 · Problem-Solving and Data Analysis · ~13 min
Statistics, data & probability
Mean vs. median, scatterplots, probability from tables, and what a study can actually claim.
The 30-second version
the mean chases outliers; the median doesn't — one billionaire in the room wrecks the average income but barely moves the middle. Probability from a table is (cell you want)/(row or column you're restricted to). And only a randomized experiment can support a cause-and-effect claim; observational studies support association only, and random sampling is what lets you generalize to the population.
Learn it
Center: mean vs. median
Mean = total ÷ count. Median = middle value when sorted. Skewed data or an outlier drags the mean toward the extreme while the median stays put — so "which is greater?" questions are really "which way does the data lean?" Also useful: if the mean of 5 tests is 84, the total is 420. Many SAT mean questions are secretly total questions ("what must she score on the sixth test…").
Scatterplots and best fit
The line of best fit is a linear model: its slope is a rate ("each additional hour of practice predicts 4 more points"), its intercept a baseline. A point above the line means the actual value beat the prediction; the vertical gap is the size of the miss. Questions asking "the y-value predicted for x = 30" want the line's height there, not a data point's.
Probability from two-way tables
"Given that a randomly chosen participant is a junior, what's the probability they chose robotics?" — the "given" restricts you to the junior row: answer = robotics-juniors ÷ all juniors. Unrestricted probability divides by the grand total. Read which universe the question locks you into before dividing.
Margin of error & study design
A poll's "38% ± 3%" means the population value is plausibly between 35% and 41% — larger random samples shrink the margin. And the claims ladder: random assignment to treatments → may conclude causation; random sampling → may generalize; neither → the study describes only the people in it. The SAT asks "what's the strongest supported conclusion?" — pick the rung the design actually earns.
Common traps
Dividing by the grand total on a "given that" question.
"The mean must be one of the data values." It usually isn't.
Causal language from observational data — "improves," "causes," "leads to" need an experiment.
Try it
Nine houses on a street sell for about $300,000 each; a tenth sells for $3 million. Which statement is true?
The median stays near $300,000 (the middle house is ordinary), but the $3M outlier inflates the mean to roughly $570,000. Outliers drag the mean toward themselves and barely touch the median — the single most-tested statistics fact.
Free structured lessons, video instruction, and practice exercises for both SAT sections, built with College Board as the official practice partner. After a Bluebook practice test, you can link your results to get personalized practice recommendations.
The official library of real SAT Suite questions. You can filter by section (Math or Reading and Writing), by content domain, by specific skill, and by difficulty, then export a custom practice set with answer explanations.
Official College BoardCompletely Free
Practice next
Question Bank filter: Math → Problem-Solving and Data Analysis → the statistics, probability, and data-inference skills. These questions are wordy; practice extracting the numbers before the story.