AP Statistics Guide: Two-Sample $t$-Test vs. Matched Pairs Inference & Verification
1. Introduction & AP Exam Weight
In quantitative analysis, determining whether two sets of numerical observations are independent or dependent is a core statistical threshold. On the AP Statistics Exam, Inference for Quantitative Data: Means (Units 6 & 7) accounts for 10–18% of the multiple-choice section and appears consistently on the Free-Response section (FRQs).
The distinction between a Two-Sample $t$-Test for the Difference of Two Means ($\mu_1 - \mu_2$) and a Matched Pairs $t$-Test for a Mean Difference ($\mu_d$) is one of the most frequent discriminators between a Score 4 and a Score 5. AP Readers frequently grade FRQs where students apply a two-sample procedure to dependent data or fail to state conditions in terms of sample differences.
Mastering this distinction requires a solid grasp of mathematical probability: recognizing how pairing controls for extraneous variation by exploiting covariance, and verifying conditions on the correct statistical distribution.
2. Deep Concept Breakdown
Mathematical Foundations: The Role of Covariance
Consider two random variables $X_1$ and $X_2$ representing two measurements. The variance of their difference is defined as:
$$\text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$
Case 1: Independent Samples (Two-Sample $t$-Test)
When observations in Group 1 are independent of observations in Group 2, $\text{Cov}(X_1, X_2) = 0$. The variance of the sample mean difference $\bar{x}_1 - \bar{x}_2$ simplifies to:
$$\text{Var}(\bar{x}_1 - \bar{x}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$
Replacing parameters with sample variances gives the Standard Error of $\bar{x}_1 - \bar{x}_2$:
$$SE_{\bar{x}_1 - \bar{x}_2} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$
The test statistic is calculated as:
$$t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)_0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$$
Degrees of freedom ($df$) are calculated using the Satterthwaite approximation:
$$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{\left(s_1^2/n_1\right)^2}{n_1 - 1} + \frac{\left(s_2^2/n_2\right)^2}{n_2 - 1}}$$
(Note: AP Statistics permits the conservative lower-bound estimate $df = \min(n_1 - 1, n_2 - 1)$, though technological Satterthwaite output is preferred).
Case 2: Dependent/Paired Samples (Matched Pairs $t$-Test)
When observations are paired (e.g., pre-test vs. post-test on the same subject, or twin studies), $X_1$ and $X_2$ are positively correlated ($\text{Cov}(X_1, X_2) > 0$). Subtracting $2\text{Cov}(X_1, X_2)$ reduces the variability of the difference, removing subject-to-subject baseline variation.
We reduce the bivariate data $(x_{1i}, x_{2i})$ to a single univariate sample of differences:
$$d_i = x_{1i} - x_{2i} \quad \text{for } i = 1, 2, \dots, n_d$$
The hypotheses are stated in terms of the single parameter $\mu_d$, the population mean difference:
$$H_0: \mu_d = \Delta_0 \quad (\text{typically } \Delta_0 = 0)$$
The Standard Error of $\bar{x}_d$ is:
$$SE_{\bar{x}_d} = \frac{s_d}{\sqrt{n_d}}$$
where $s_d = \sqrt{\frac{\sum (d_i - \bar{x}_d)^2}{n_d - 1}}$. The test statistic follows a standard $t$-distribution with $df = n_d - 1$:
$$t = \frac{\bar{x}d - \mu{d,0}}{\frac{s_d}{\sqrt{n_d}}}$$
Condition Verification Comparison
| Condition | Two-Sample $t$-Test ($\mu_1 - \mu_2$) | Matched Pairs $t$-Test ($\mu_d$) |
|---|---|---|
| Randomness | Two independent random samples OR random assignment of units to two treatment groups. | A single random sample of paired units OR random assignment of treatments within pairs. |
| 10% Rule | $n_1 \le 0.10 N_1$ and $n_2 \le 0.10 N_2$ (if sampling without replacement). | $n_d \le 0.10 N_{\text{pairs}}$ (if sampling without replacement). |
| Normality / Large Sample | Both populations normal OR $n_1 \ge 30$ and $n_2 \ge 30$. If $n_1, n_2 < 30$, inspect both sample distributions for strong skew/outliers. | Population of differences is normal OR $n_d \ge 30$. If $n_d < 30$, inspect the single distribution of sample differences $d_i$ for strong skew/outliers. |
Computational Simulation: Variance Reduction via Pairing
The following Python snippet demonstrates how blocking (pairing) isolates variation and increases statistical power compared to an un paired analysis:
import numpy as np
from scipy import stats
# Seed for reproducibility
np.random.seed(42)
# Simulate 20 subjects with high baseline variability (Subject effect ~ N(100, 15^2))
subject_baselines = np.random.normal(100, 15, 20)
# Treatment adds a subtle true mean improvement of mu_d = 3.5
# Individual response variance = 2.0^2
pre_treatment = subject_baselines + np.random.normal(0, 2.0, 20)
post_treatment = subject_baselines + 3.5 + np.random.normal(0, 2.0, 20)
# 1. Incorrect Approach: Independent Two-Sample t-test
t_two_sample, p_two_sample = stats.ttest_ind(post_treatment, pre_treatment)
# 2. Correct Approach: Matched Pairs t-test
differences = post_treatment - pre_treatment
t_paired, p_paired = stats.ttest_rel(post_treatment, pre_treatment)
print(f"--- INDEPENDENT TWO-SAMPLE t-TEST ---")
print(f"t-statistic: {t_two_sample:.4f} | p-value: {p_two_sample:.4f}")
print(f"Outcome: {'Fail to Reject H0' if p_two_sample > 0.05 else 'Reject H0'}\n")
print(f"--- MATCHED PAIRS t-TEST ---")
print(f"t-statistic: {t_paired:.4f} | p-value: {p_paired:.4f}")
print(f"Outcome: {'Fail to Reject H0' if p_paired > 0.05 else 'Reject H0'}")
Output Insights: The two-sample test fails to reject $H_0$ because the variance from subject baselines inflates the standard error denominator. The paired $t$-test eliminates baseline variance by operating solely on differences, lowering standard error and detecting the treatment effect.
3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances
Score 4 vs. Score 5 Student Response Architecture
AP INFERENCE FRQ EVALUATION
│
┌────────────────────────────┴────────────────────────────┐
▼ ▼
SCORE 4 RESPONSE SCORE 5 RESPONSE
────────────────────────────── ──────────────────────────────
• Selects Two-Sample t-Test • Identifies Matched Pairs
for paired data. design based on structure.
• Checks Normality on raw • Explicitly defines mu_d:
Group 1 and Group 2 data. "mean difference in [units]".
• States generic hypotheses: • Graphs and verifies Normal
H0: mu1 = mu2. condition ONLY on differences.
• Omits context in final • States decision relative
p-value linkage. to alpha with full context.
Free-Response Scoring Rubric Breakdown (Four Essential Components)
To earn an Essentially Correct (E) on an AP Statistics inference question, your solution must satisfy four distinct components:
1. Statement of Hypotheses & Parameter Definition
- Pitfall: Writing $H_0: \mu_1 - \mu_2 = 0$ when data are paired, or defining $\mu_d$ simply as "the mean difference" without direction or context.
- Score 5 Requirement: $$H_0: \mu_d = 0 \quad \text{vs.} \quad H_a: \mu_d > 0$$ Where $\mu_d$ represents the true population mean difference in [Variable Name] (Post $-$ Pre) for [Target Population].
2. Identification of Procedure & Verification of Conditions
- Pitfall: Checking normality on Group 1 and Group 2 separately when running a matched pairs test, or stating "Normality is met because $n \ge 30$" without showing data plots when $n_d < 30$.
- Score 5 Requirement: Name the procedure as a "One-Sample $t$-test for Sample Differences ($\mu_d$)" or "Matched Pairs $t$-test". Construct a dotplot, boxplot, or stem-and-leaf plot of the differences $d_i$, and explicitly write: "The plot of sample differences shows no strong skewness or outliers. Thus, the condition for population normality of differences is plausible."
3. Correct Mechanics & Statistics
- Pitfall: Plugging incorrect degrees of freedom into test statistic calculations.
- Score 5 Requirement: Report test statistic ($t$), degrees of freedom ($df = n_d - 1$), and exact $P$-value. $$t = \frac{\bar{x}_d - 0}{\frac{s_d}{\sqrt{n_d}}}, \quad df = n_d - 1, \quad P\text{-value} = P(T > t)$$
4. Contextual Conclusion Linked to $P$-Value
- Pitfall: Writing "We accept $H_0$" or leaving out explicit comparison to $\alpha$.
- Score 5 Requirement: Compare $P$-value directly to $\alpha$. State explicit decision (Reject/Fail to Reject $H_0$) and contextualize the conclusion: "Because $P\text{-value} = 0.012 < \alpha = 0.05$, we reject $H_0$. There is convincing statistical evidence that the true mean difference in [Variable Name] is greater than zero."
4. Harvard University Placement Pathway
Mastering statistical inference yields concrete academic advantages at elite institutions such as Harvard University.
HARVARD QUANTITATIVE PATHWAY
│
AP Statistics Exam (Score 5)
│
┌─────────────────────┴─────────────────────┐
▼ ▼
EXEMPTED FROM STAT 100 DIRECT ADVANCED PLACEMENT
(General Quantitative Reasoning) │
┌───────────────────┴───────────────────┐
▼ ▼
STAT 110 STAT 111
(Introduction to Probability) (Statistical Inference)
Course Exemption & Placement Analysis
- Exempted Requirement: A Score of 5 on AP Statistics satisfies Harvard's College General Education or Quantitative Reasoning distribution requirement, granting equivalent placement for Stat 100: Introduction to Quantitative Methods for the Social Sciences.
- Accelerated Placement Track: Passing out of introductory statistics allows immediate enrollment into higher-level foundational courses:
- Stat 110: Introduction to Probability (taught by Prof. Joe Blitzstein): A standard rigor milestone for concentrators in Applied Math, Computer Science, Data Science, Economics, and Statistics.
- Stat 111: Introduction to Statistical Inference: Theoretical treatment of point estimation, confidence regions, hypothesis testing, Neyman-Pearson Lemma, and Likelihood Ratio Tests.
Strategic Academic Advantage
Understanding the core distinction between independent sample variance and paired block design ($\text{Cov}(X_1, X_2)$ variance reduction) directly prepares students for experimental design and econometric modeling: * In Harvard’s Economics Department (e.g., Econ 1123 / Econ 1126), the matched pairs framework serves as the conceptual foundation for Difference-in-Differences (DiD) estimation and panel data fixed-effects estimators. * In Government/Data Science tracks (e.g., Gov 50 / Gov 2001), controlling for subject-level confounders via matched pair matching algorithms (e.g., propensity score matching) is fundamental to establishing modern causal inference under potential outcomes modeling.
5. High-Yield Practice Problem & Step-by-Step Solution Checklist
Problem Statement
A clinical research team at a university medical center investigates whether a specialized cognitive training software reduces the time (in seconds) required for older adults to complete a complex spatial task.
A random sample of 10 participants aged 65 or older is selected. Each participant's task completion time is measured prior to training ($x_{\text{pre}}$) and again after completing a 4-week training regimen ($x_{\text{post}}$). The table below displays the completion times in seconds:
| Subject | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 | 9 | 10 |
|---|---|---|---|---|---|---|---|---|---|---|
| Pre-Training ($x_{\text{pre}}$) | 45.2 | 38.0 | 52.1 | 41.5 | 49.8 | 44.0 | 51.3 | 39.7 | 46.5 | 43.0 |
| Post-Training ($x_{\text{post}}$) | 41.0 | 36.5 | 47.0 | 41.8 | 44.2 | 40.1 | 46.5 | 38.0 | 42.1 | 39.5 |
Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the training software reduces the mean completion time for this population?
Complete Step-by-Step Solution Checklist
Step 1: Identify Parameter and Hypotheses
Let $d_i = x_{\text{pre}, i} - x_{\text{post}, i}$ denote the pairwise difference in time (in seconds) for subject $i$.
Define $\mu_d$ as the true population mean difference in completion time (Pre $-$ Post) for adults aged 65 and older undergoing this training program.
- Hypotheses: $$H_0: \mu_d = 0$$ $$H_a: \mu_d > 0$$
(Note: Defining $d = \text{Pre} - \text{Post}$ means a reduction in time corresponds to $d > 0$.)
Step 2: Identify Procedure and Verify Conditions
Procedure Name: Matched Pairs $t$-test for a Mean Difference ($\mu_d$).
Calculate the 10 pairwise difference values $d_i = x_{\text{pre}, i} - x_{\text{post}, i}$:
$$\text{Differences } (d_i): {+4.2, +1.5, +5.1, -0.3, +5.6, +3.9, +4.8, +1.7, +4.4, +3.5}$$
Condition Verification: 1. Random Sample: The problem explicitly states that a random sample of 10 participants aged 65 or older was selected. 2. 10% Condition: $n_d = 10 \le 0.10 \times (\text{Total population of adults }\ge 65)$. This is reasonable to assume. 3. Normality of Differences: Since sample size $n_d = 10 < 30$, we must construct a plot of sample differences to verify the distribution shows no severe skewness or outliers.
DOTPLOT OF DIFFERENCES (d_i)
* * * * * *
───┼───┼───┼───┼───┼───┼───┼───┼───┼───┼───┼───
-1 0 1 2 3 4 5 6
d_i (seconds)
Verification Statement: The dotplot of the 10 sample differences shows no extreme outliers or strong skewness. Therefore, the condition that the population of differences is approximately normal is satisfied.
Step 3: Compute Test Mechanics
Calculate sample statistics from the difference vector $d_i$:
$$\bar{x}_d = \frac{\sum d_i}{n_d} = \frac{34.4}{10} = 3.44 \text{ seconds}$$
$$s_d = \sqrt{\frac{\sum (d_i - \bar{x}_d)^2}{n_d - 1}} = 1.833 \text{ seconds}$$
$$SE_{\bar{x}_d} = \frac{s_d}{\sqrt{n_d}} = \frac{1.833}{\sqrt{10}} \approx 0.5796$$
Compute Test Statistic $t$:
$$t = \frac{\bar{x}_d - 0}{\frac{s_d}{\sqrt{n_d}}} = \frac{3.44 - 0}{0.5796} \approx 5.935$$
Degrees of Freedom:
$$df = n_d - 1 = 10 - 1 = 9$$
Calculate $P$-value:
$$P\text{-value} = P(T_9 \ge 5.935) < 0.0001$$
(Using $t$-table or calculator: $\text{tcdf}(5.935, \infty, 9) \approx 0.000112$)
Step 4: Write Conclusion in Context
Because the $P\text{-value} \approx 0.00011$ is less than the significance level $\alpha = 0.05$, we reject $H_0$.
There is convincing statistical evidence that the cognitive training software reduces the true mean task completion time (in seconds) for older adults aged 65 and older.