AP Statistics Mastery Guide: Two-Sample $t$-Test vs. Matched Pairs Inference & Verification
Target Institution: Georgia Institute of Technology (Georgia Tech)
Score Target: 5 (Essential for STEM/Analytics Placement)
Course Credit Exempted: ISYE 3770 (Statistics and Applications — 3 Credit Hours)
Accelerated Track: MATH 3215 (Introduction to Probability & Statistics) / ISYE 2028 (Basic Statistical Methods)
1. Introduction & AP Exam Weight
Inference for two-sample data and paired data accounts for 10–15% of the AP Statistics Exam multiple-choice section and appears consistently on Free-Response Questions (FRQs)—frequently on FRQ 3, 4, or 5, and as a crucial component of the Investigative Task (FRQ 6).
The critical challenge evaluated by AP Chief Readers is a student's ability to distinguish between two independent samples and a single sample of paired observations (matched pairs). Confusing a Matched Pairs $t$-procedure with a Two-Sample $t$-procedure is an immediate automatic downgrade on Section II of the AP Exam (reducing a response from Essentially Correct $[E]$ to Incomplete $[I]$ or Incorrect $[I]$).
Why Georgia Tech Demands Perfection Here
Georgia Tech’s H. Milton Stewart School of Industrial and Systems Engineering (ISyE)—consistently ranked #1 nationally—and the College of Computing require absolute statistical precision.
By achieving a Score 5 on the AP Statistics Exam, you earn credit for ISYE 3770, waiving a 3-credit core requirement. This allows immediate acceleration into MATH 3215 or advanced operations research modules in your freshman year.
In industrial design, computer architecture benchmarking, and algorithmic optimization, mistaking dependent paired data for independent samples introduces massive covariance bias. Master this distinction to demonstrate true university-level mastery.
2. Deep Concept Breakdown
Mathematical Foundations & Variance Structure
The core difference between a Two-Sample $t$-test and a Matched Pairs $t$-test lies in the covariance structure of the random variables under observation.
1. Two Independent Samples (Two-Sample $t$-Test)
Let $X_1 \sim N(\mu_1, \sigma_1^2)$ and $X_2 \sim N(\mu_2, \sigma_2^2)$ be two independent random variables. The linear combination representing their difference is $D = X_1 - X_2$.
By the properties of expectation and variance for independent variables: $$\mathbb{E}[\bar{X}_1 - \bar{X}_2] = \mu_1 - \mu_2$$
$$\text{Var}(\bar{X}_1 - \bar{X}_2) = \text{Var}(\bar{X}_1) + \text{Var}(\bar{X}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$
Standard Error ($SE_{\bar{X}1 - \bar{X}_2}$): $$SE{\bar{X}_1 - \bar{X}_2} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$
Degrees of Freedom ($df$): Computed using the Welch–Satterthwaite equation: $$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{\left(\frac{s_1^2}{n_1}\right)^2}{n_1 - 1} + \frac{\left(\frac{s_2^2}{n_2}\right)^2}{n_2 - 1}}$$ (Note: On the AP Exam, technology calculates Welch-Satterthwaite $df$, or the conservative approximation $df = \min(n_1-1, n_2-1)$ is used).
2. Dependent / Paired Samples (Matched Pairs $t$-Test)
When observations are matched in pairs $(X_{1i}, X_{2i})$ for $i = 1, 2, \dots, n$, the variables are not independent. $$\text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$
If subjects are positively correlated ($\text{Cov}(X_1, X_2) > 0$), pairing reduces overall variance, drastically increasing the statistical power of the test to detect true differences!
We collapse the paired data into a single univariate random variable $d_i = X_{1i} - X_{2i}$: $$\bar{d} = \frac{1}{n}\sum_{i=1}^n d_i, \quad s_d = \sqrt{\frac{\sum (d_i - \bar{d})^2}{n - 1}}$$
Standard Error ($SE_{\bar{d}}$): $$SE_{\bar{d}} = \frac{s_d}{\sqrt{n}}$$
Degrees of Freedom: $$df = n - 1 \quad \text{(where } n \text{ is the number of pairs)}$$
Comparative Structural Matrix
| Architectural Feature | Two-Sample $t$-Test | Matched Pairs $t$-Test |
|---|---|---|
| Data Structure | Two separate, independent groups ($n_1, n_2$) | One sample of paired units, or repeated measures on same units ($n$ pairs) |
| Parameter of Interest | Difference of means: $\mu_1 - \mu_2$ | Mean difference: $\mu_d$ (where $d = x_1 - x_2$) |
| Null Hypothesis ($H_0$) | $H_0: \mu_1 - \mu_2 = 0$ | $H_0: \mu_d = 0$ |
| Test Statistic | $t = \frac{(\bar{x}_1 - \bar{x}_2) - 0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$ | $t = \frac{\bar{x}_d - 0}{\frac{s_d}{\sqrt{n}}}$ |
| Degrees of Freedom | Satterthwaite or $\min(n_1-1, n_2-1)$ | $n_{pairs} - 1$ |
| Normality Condition | Both sample distributions must be normal/large ($n_1, n_2 \ge 30$) | The single distribution of differences ($d_i$) must be normal/large ($n_{pairs} \ge 30$) |
Python Verification Script: Algorithmic Differentiation
This Python script demonstrates how treating paired data as independent two-sample data severely inflates standard error and distorts the $p$-value:
import numpy as np
from scipy import stats
# Simulated Data: Pre-treatment vs Post-treatment optimization benchmark (ms)
np.random.seed(42)
pre_treatment = np.array([120.4, 135.2, 128.1, 145.0, 118.9, 133.5, 140.1, 129.8])
# Strong positive correlation added to simulate paired observations
post_treatment = pre_treatment - np.random.normal(loc=5.0, scale=1.5, size=8)
print("=== CORRECT METHOD: Matched Pairs t-Test ===")
differences = pre_treatment - post_treatment
t_stat_paired, p_val_paired = stats.ttest_rel(pre_treatment, post_treatment)
print(f"Mean Difference (d_bar): {np.mean(differences):.4f}")
print(f"Std Dev of Differences (s_d): {np.std(differences, ddof=1):.4f}")
print(f"Paired t-statistic: {t_stat_paired:.4f}")
print(f"Paired p-value: {p_val_paired:.6f}\n")
print("=== INCORRECT METHOD: Two-Sample t-Test (Ignoring Pairing) ===")
t_stat_2samp, p_val_2samp = stats.ttest_ind(pre_treatment, post_treatment, equal_var=False)
print(f"Two-Sample t-statistic: {t_stat_2samp:.4f}")
print(f"Two-Sample p-value: {p_val_2samp:.6f}")
# Output Analysis:
# Notice how the paired test detects significance due to controlling baseline noise!
3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances
To secure a 5, your response must be complete, precise, and contextualized. AP Readers follow strict analytical rubrics.
Top AP Exam Pitfalls
- Incorrect Parameter Definition:
- Score 4 mistake: Defining $\mu_d$ as "the difference in mean scores." (Ambiguous: sounds like $\mu_1 - \mu_2$).
-
Score 5 execution: "Let $\mu_d$ be the true mean difference in benchmark execution time (Pre-optimization minus Post-optimization) for Georgia Tech servers."
-
Checking Conditions on Separate Groups for Paired Data:
- Score 4 mistake: Plotting $X_1$ and $X_2$ separately, showing histogram graphs of both groups for a matched-pairs test.
-
Score 5 execution: Calculating $d_i = X_{1i} - X_{2i}$, plotting a single boxplot or histogram of the differences, and stating: "The plot of differences shows no extreme skewness or outliers, satisfying the Normality condition for $n = 8 < 30$."
-
Confusing Random Assignment with Random Sampling:
- If data comes from a paired experiment (e.g., subjects assigned both treatments in random order), state: "Randomization condition is met via random assignment of treatment order within each pair/subject." Do not claim it is a "random sample of subjects" unless explicitly stated.
Rubric Comparison Matrix: Score 4 vs. Score 5 Solutions
| Rubric Component | Score 4 Response (Sub-Optimal) | Score 5 Response (AP Reader Ideal) |
|---|---|---|
| Hypotheses & Parameter | $H_0: \mu_1 = \mu_2$, $H_a: \mu_1 > \mu_2$ (Fails to recognize pairing structure) |
$H_0: \mu_d = 0$ vs $H_a: \mu_d > 0$ Where $\mu_d = \text{true mean difference in... } (X_1 - X_2)$ |
| Conditions Check | Writes "Normal because $N > 30$" without checking differences or providing graph sketches when $n < 30$. | Explicitly lists: 1. Paired data/random order; 2. $n_{pairs} \le 10\%$ of population; 3. Sketch of $d_i$ showing no strong skewness/outliers. |
| Mechanics | Reports $t = 3.42, p = 0.002$ without listing formula, substitution, or $df$. | Shows $t = \frac{\bar{d} - 0}{s_d / \sqrt{n}} = \frac{4.81}{1.38/\sqrt{8}} = 9.86$, $df = 7$, $p\text{-value} = 0.000012$. |
| Conclusion | "Reject $H_0$. There is a difference." | "Because $p$-value $\approx 0.000012 < \alpha = 0.05$, we reject $H_0$. There is convincing statistical evidence that the true mean difference in..." |
4. Georgia Tech Placement Pathway
Exemption Dynamics: ISYE 3770 Waiver
Achieving a Score 5 on AP Statistics opens immediate academic advantages at Georgia Tech:
[AP Statistics: Score 5]
│
▼
[Waives ISYE 3770: Statistics & Applications (3 Credits)]
│
├─────────────────────────────────────────┐
▼ ▼
[Immediate Entry: MATH 3215] [Accelerated Path: ISYE 2028]
Intro to Probability & Stats Basic Statistical Methods
│ │
▼ ▼
[Advanced Stochastic Modeling] [ISYE 3044: Simulation Analysis]
Strategic Value for GT Majors
- Industrial & Systems Engineering (ISyE):
- ISyE students who clear ISYE 3770 early unlock ISYE 2028 (Basic Statistical Methods) in their second semester.
- This pushes prerequisite completion for critical junior-level tracks like ISYE 3044 (Simulation Analysis) and ISYE 3231 (Deterministic Operations Research) ahead by a full academic year.
- Computer Science (CS - AI/ML Threads):
- Enables immediate enrollment in MATH 3215 (Probability & Statistics for Applications), bypassing introductory survey courses and moving directly into CS 4641 (Machine Learning) and CS 4731 (Game AI).
- Variance Reduction Expertise:
- Matched pairs inference is the foundational baseline for Design of Experiments (DoE) and A/B testing protocols taught in GT’s high-level engineering sequences.
5. High-Yield Practice Problem & Step-by-Step Solution Checklist
The Problem
A Georgia Tech computer science researcher is testing a new memory-compression algorithm designed to reduce thread execution times. Eight server cores are selected at random. Each core runs a complex benchmark under two conditions: Control (standard memory management) and Treatment (compressed memory management). The order in which each core runs the two conditions is randomized.
The benchmark completion times (in milliseconds) are recorded below:
| Core ID | 1 | 2 | 3 | 4 | 5 | 6 | 7 | 8 |
|---|---|---|---|---|---|---|---|---|
| Control ($X_1$) | 142.5 | 158.0 | 133.2 | 167.4 | 129.1 | 151.3 | 160.8 | 144.7 |
| Treatment ($X_2$) | 138.1 | 151.2 | 129.9 | 160.1 | 126.5 | 144.8 | 153.2 | 139.0 |
Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the new compression algorithm reduces mean benchmark completion time across server cores?
Step-by-Step Complete Solution Checklist
Step 1: Identify Procedure & Hypotheses
- Procedure: Matched Pairs $t$-Test for a Mean Difference.
- Parameter Definition: Let $\mu_d$ be the true mean difference in benchmark completion time ($\text{Control} - \text{Treatment}$, in ms) for server cores.
- Hypotheses: $$H_0: \mu_d = 0$$ $$H_a: \mu_d > 0$$
Step 2: Check Conditions
- Paired Data & Randomization: Data are paired by server core. The order of treatments (Control vs. Treatment) was randomized within each core.
- 10% Condition: $n = 8$ server cores is reasonable to assume as $\le 10\%$ of all available server cores in high-performance computing clusters.
- Normality Condition: Calculate the sample differences $d_i = X_{1i} - X_{2i}$:
$$\mathbf{d = [4.4, 6.8, 3.3, 7.3, 2.6, 6.5, 7.6, 5.7]}$$
- $n = 8 < 30$, so we must plot the sample differences.
- Sketch/Description of Graph: A dotplot or boxplot of the differences shows values clustered between $2.6$ and $7.6$. The distribution displays no strong skewness and contains no outliers. The Normality condition is satisfied.
Dotplot of Differences (d):
o o o o o o o o
──┼───┼───┼───┼───┼───┼───┼───┼───┼───
2.0 3.0 4.0 5.0 6.0 7.0 8.0 (ms)
Step 3: Compute Test Statistic & Mechanics
Calculate summary statistics for differences: * Sample size $n = 8$ * Mean of differences $\bar{d} = \frac{4.4 + 6.8 + 3.3 + 7.3 + 2.6 + 6.5 + 7.6 + 5.7}{8} = 5.525\text{ ms}$ * Standard deviation of differences $s_d = 1.838\text{ ms}$
Test Statistic Computation: $$t = \frac{\bar{d} - \mu_0}{\frac{s_d}{\sqrt{n}}} = \frac{5.525 - 0}{\frac{1.838}{\sqrt{8}}} = \frac{5.525}{0.6498} \approx 8.503$$
Degrees of Freedom: $$df = n - 1 = 8 - 1 = 7$$
$p$-Value Calculation: $$p\text{-value} = P(T > 8.503 \mid df = 7) \approx 0.0000302 \quad (3.02 \times 10^{-5})$$
Step 4: State Conclusion in Context
Because the $p$-value ($\approx 0.0000302$) is substantially less than the significance level $\alpha = 0.05$, we reject $H_0$.
There is convincing statistical evidence that the true mean benchmark completion time under the compressed memory algorithm (Treatment) is significantly lower than under standard memory management (Control) across server cores.
Key Rubric Self-Check for FRQ Credit
- [x] Did you explicitly name the Matched Pairs $t$-test?
- [x] Is $\mu_d$ defined with the correct order of subtraction ($\text{Control} - \text{Treatment}$)?
- [x] Did you construct a plot of the differences (not the raw groups)?
- [x] Did you link the final decision explicitly to $\alpha$ ($p < \alpha \implies \text{Reject } H_0$)?