Statistics • Score 5 Strategy

Two-Sample t-Test vs. Matched Pairs Inference & Verification Guide: AP Statistics Score 5 for Carnegie Mellon University

AP Statistics Mastery Guide: Two-Sample $t$-Test vs. Matched Pairs Inference & Verification


1. Introduction & AP Exam Weight

In statistical inference for quantitative data (AP Statistics Unit 7: Inference for Quantitative Data: Means), few topics test conceptual fluency and structural precision as rigorously as distinguishing between a Two-Sample $t$-Test for the Difference of Two Independent Means and a Matched Pairs $t$-Test for a Mean Difference.

Combined, quantitative inference topics constitute 10–18% of the multiple-choice section and appear systematically in Free-Response Questions (FRQs)—frequently in Question 1, Question 3, or Question 6 (the Investigative Task).

The critical challenge lies not in arithmetic computation, but in experimental and observational design verification. Misidentifying dependent (paired) data as independent two-sample data—or vice versa—results in structural error, cascading incorrect condition checks, improper degrees of freedom ($df$), flawed standard errors, and ultimately a score drop from an Essentially Correct (E) to an Incomplete (I) on the College Board AP grading scale.


2. Deep Concept Breakdown

Mathematical Foundations & Covariance Structure

To understand why design determines inference, we must examine the variance of the difference between two random variables $X_1$ and $X_2$:

$$\text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$

Case 1: Two Independent Samples

When two samples are selected independently (e.g., randomly assigning subjects to Treatment A or Treatment B, or sampling two independent populations), the covariance between observations across groups is zero:

$$\text{Cov}(X_1, X_2) = 0 \implies \text{Var}(\bar{X}_1 - \bar{X}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$

The sample standard error for the difference between sample means $\bar{x}_1 - \bar{x}_2$ is:

$$\text{SE}_{\bar{x}_1 - \bar{x}_2} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

Degrees of freedom are calculated using the Welch–Satterthwaite approximation:

$$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{1}{n_1-1}\left(\frac{s_1^2}{n_1}\right)^2 + \frac{1}{n_2-1}\left(\frac{s_2^2}{n_2}\right)^2}$$

(Note: AP Statistics permits the conservative linear approximation $df = \min(n_1 - 1, n_2 - 1)$ when software outputs are unavailable, though statistical software uses Welch's exact value).

Case 2: Matched Pairs (Dependent Samples)

When data are paired—such as measuring the same subject before and after a treatment, or pairing twins/matched subjects—a positive covariance is intentionally introduced ($\text{Cov}(X_1, X_2) > 0$). This pairing reduces noise by controlling for subject-specific baseline variability.

We collapse the paired observations $(x_{1i}, x_{2i})$ into a single univariate random variable of differences:

$$d_i = x_{1i} - x_{2i}$$

The statistical model simplifies to a One-Sample $t$-procedure on the differences $D$:

$$\bar{d} = \frac{1}{n}\sum_{i=1}^n d_i, \quad s_d = \sqrt{\frac{1}{n-1}\sum_{i=1}^n (d_i - \bar{d})^2}$$

$$\text{SE}_{\bar{d}} = \frac{s_d}{\sqrt{n}}$$

Degrees of freedom simplify directly to:

$$df = n - 1$$

where $n$ is the number of pairs, not total observations.


Computational Proof: Statistical Power Gain via Matching

The following Python script illustrates how treating paired data as two independent samples artificially inflates variance, leading to a loss of statistical power (a Type II error risk).

import numpy as np
from scipy import stats

# Seed for reproducibility
np.random.seed(42)

# Generate baseline traits (nuisance variability across individuals)
baseline_ability = np.random.normal(loc=100, scale=15, size=30)

# Apply treatment with a true mean effect size delta = +4.0
treatment_effect = np.random.normal(loc=4.0, scale=2.0, size=30)

# Construct Paired Data
pre_treatment = baseline_ability
post_treatment = baseline_ability + treatment_effect

# Compute explicit differences
differences = post_treatment - pre_treatment

# --- 1. Correct Analysis: Matched Pairs t-Test ---
t_stat_paired, p_val_paired = stats.ttest_rel(post_treatment, pre_treatment)
df_paired = len(differences) - 1

# --- 2. Incorrect Analysis: Independent Two-Sample t-Test ---
t_stat_indep, p_val_indep = stats.ttest_ind(post_treatment, pre_treatment, equal_var=False)

print(f"=== MATCHED PAIRS t-TEST ===")
print(f"Sample Mean Diff (d_bar) : {np.mean(differences):.4f}")
print(f"Standard Error SE(d_bar) : {np.std(differences, ddof=1)/np.sqrt(30):.4f}")
print(f"t-statistic              : {t_stat_paired:.4f}")
print(f"Degrees of Freedom (df)  : {df_paired}")
print(f"p-value                  : {p_val_paired:.6e}\n")

print(f"=== INCORRECT TWO-SAMPLE t-TEST ===")
print(f"Difference of Means      : {np.mean(post_treatment) - np.mean(pre_treatment):.4f}")
print(f"t-statistic              : {t_stat_indep:.4f}")
print(f"p-value                  : {p_val_indep:.6f}")

Output Insight: The paired $t$-test produces a significantly larger $t$-statistic and a lower $p$-value because it removes the inter-subject variability ($\text{Var}(\text{baseline}) = 15^2 = 225$), isolating the true effect variance.


3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

                       ┌──────────────────────────────────────┐
                       │ Are observations linked in pairs?    │
                       │ (e.g., Before/After, Matched Units) │
                       └──────────────────┬───────────────────┘
                                          │
                        ┌─────────────────┴─────────────────┐
                        ▼                                   ▼
                     [ YES ]                             [ NO ]
                        │                                   │
                        ▼                                   ▼
          ┌───────────────────────────┐       ┌───────────────────────────┐
          │     MATCHED PAIRS         │       │     TWO-SAMPLE t-TEST     │
          │   (Single Mean Diff)      │       │   (Diff of Two Means)     │
          └─────────────┬─────────────┘       └─────────────┬─────────────┘
                        │                                   │
                        ▼                                   ▼
          • Parameter: μ_d                    • Parameter: μ₁ - μ₂
          • Data: Differences (d_i = x₁ - x₂) • Data: Group 1 (x₁) & Group 2 (x₂)
          • Check: Normality of DIFFERENCES   • Check: Normality of BOTH groups
          • df = n_pairs - 1                  • df = Welch-Satterthwaite / min(n₁-1, n₂-1)

High-Frequency AP Exam Pitfalls

  1. Parameter Definition Mismatch:
  2. Incorrect: Writing $\mu_1 - \mu_2$ for a matched pairs test, or defining $\mu_d$ without establishing the direction of subtraction (e.g., $d = \text{Post} - \text{Pre}$).
  3. Correct: For Matched Pairs, state: "$\mu_d$ is the true mean difference in [variable] (Post - Pre) for [population]."\ For Two-Sample, state: "$\mu_1 - \mu_2$ is the true difference in mean [variable] between [Population 1] and [Population 2]."

  4. Checking Conditions on the Wrong Distributions:

  5. For Matched Pairs, do not construct separate histograms/boxplots for Sample 1 and Sample 2. You must construct a single graph of the differences $d_i$ and assess its symmetry/skewness.
  6. For Two-Sample, you must assess the normality/sample size conditions independently for both groups.

  7. Confusing Degrees of Freedom:

  8. Stating $df = n_1 + n_2 - 2$ on an AP exam without explicitly showing standard pooled calculations is penalized. AP Statistics assumes unequal variances by default.

Score 4 vs. Score 5 Performance Matrix

Evaluation Dimension Score 4 Response (Competent) Score 5 Response (Mastery)
Model Identification Correctly chooses Matched Pairs $t$-test. Explicitly states: "Matched Pairs $t$-test for a mean difference because data are paired by subject (pre/post measurement)."
Hypothesis Framing $H_0: \mu_d = 0$, $H_a: \mu_d > 0$. Defines $H_0: \mu_d = 0$ versus $H_a: \mu_d > 0$, where $\mu_d = \mu_{\text{after}} - \mu_{\text{before}}$ explicitly defining the context and subtraction direction.
Condition Verification States "Data is normal, $n > 30$." Checks: 1. Paired data from a random sample; 2. $10\%$ condition ($n = 25 \le 10\%$ of all subjects); 3. Plots differences, notes no strong skewness or outliers ($n < 30$).
Mechanics & Degrees of Freedom Writes formula, plugs in values, states $t$ and $p$. Explicitly states test statistic formula, $t = 3.42$, $df = n - 1 = 24$, exact $p\text{-value} = 0.0011$, avoiding "calculator dump" penalties.
Conclusion Context "Reject $H_0$. There is evidence that the treatment works." "Since $p = 0.0011 < \alpha = 0.05$, we reject $H_0$. There is convincing statistical evidence that the true mean difference in score ($\mu_d$) is significantly greater than zero."

4. Carnegie Mellon University Placement Pathway

Mastering statistical inference yields concrete credit and academic acceleration advantages at Carnegie Mellon University (CMU).

   AP Statistics Score: 5
             │
             ▼
   Exempts: 36-200 (Reasoning with Data - 9 Units)
             │
             ├────────────────────────────────────────┐
             ▼                                        ▼
   Accelerated Entry Path 1                  Accelerated Entry Path 2
   36-225: Intro to Probability Theory       36-401: Modern Regression
   (Foundation for ML / AI in SCS)           (Applied Statistical Modeling)

Credit Exemption Details

Strategic Acceleration Path

  1. Bypassing Introductory Data Analysis: By scoring a 5, students place directly out of 36-200 and can immediately enroll in 36-225: Introduction to Probability Theory or 36-226: Introduction to Statistical Inference.

  2. Foundational Bridge to Machine Learning (SCS & Dietrich): At CMU, statistical inference is foundational for advanced machine learning:

  3. Difference of Means $\to$ Two-Sample Hypothesis Testing: Bridges to non-parametric tests, permutation tests, and A/B testing frameworks in 36-401 Modern Regression and 10-315 Intro to Machine Learning.
  4. Matched Pairs $\to$ Dependent Data & Block Designs: Directly translates into linear mixed-effects models, covariance matrix operations ($\boldsymbol{\Sigma}$ matrices), and variance-reduction techniques (e.g., control variates, blocking in experimental design).

5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Problem Statement

An educational technology firm develops an algorithm to increase reading speed (words per minute, WPM). To evaluate its effectiveness, 12 high school students are randomly selected from a large school district.

Each student’s reading speed is evaluated before training ($X_{\text{Pre}}$) and after completing a 4-week training module ($X_{\text{Post}}$). The results are summarized below:

$$\begin{array}{|c|c|c|c|} \hline \mathbf{Student} & \mathbf{Pre\text{-}Training\ (WPM)} & \mathbf{Post\text{-}Training\ (WPM)} & \mathbf{Difference\ (Post - Pre)} \ \hline 1 & 210 & 225 & +15 \ \hline 2 & 195 & 202 & +7 \ \hline 3 & 240 & 238 & -2 \ \hline 4 & 205 & 218 & +13 \ \hline 5 & 220 & 235 & +15 \ \hline 6 & 180 & 192 & +12 \ \hline 7 & 215 & 214 & -1 \ \hline 8 & 225 & 238 & +13 \ \hline 9 & 200 & 215 & +15 \ \hline 10 & 190 & 198 & +8 \ \hline 11 & 230 & 245 & +15 \ \hline 12 & 210 & 222 & +12 \ \hline \end{array}$$

Summary Statistics:

Task: Perform an appropriate statistical test at the $\alpha = 0.01$ significance level to determine if the training module significantly increases mean reading speed.


Step-by-Step Solution Checklist (4-Part AP FRQ Format)

Part 1: State Hypotheses & Parameter Definitions


Part 2: Plan (Verify Conditions)

  1. Random Sampling Condition:
  2. Check: The problem states that 12 high school students were randomly selected from a large school district.

  3. 10% Independence Condition:

  4. Check: $n = 12$. It is reasonable to assume that 12 students is less than $10\%$ of all high school students in this large school district ($N \ge 120$).

  5. Normality / Large Sample Condition ($n < 30$):

  6. Check: Since $n = 12 < 30$, we must check the distribution of sample differences $d_i$ for strong skewness or outliers.
  7. Analysis of Differences: The differences are ${-2, -1, 7, 8, 12, 12, 13, 13, 15, 15, 15, 15}$.
  8. A dotplot or boxplot of differences reveals a mild left skew due to two minor negative values, but no extreme outliers or severe skewness. Therefore, $t$-procedures remain robust and appropriate.

Part 3: Do (Mechanics)

Standard Error Calculation:

$$\text{SE}_{\bar{d}} = \frac{s_d}{\sqrt{n}} = \frac{6.03}{\sqrt{12}} = \frac{6.03}{3.4641} \approx 1.7407$$

Test Statistic $t$ Calculation:

$$t = \frac{\bar{d} - \mu_0}{\text{SE}_{\bar{d}}} = \frac{10.17 - 0}{1.7407} = 5.8425$$

$p$-value Determination: For a one-tailed upper test with $t = 5.8425$ and $df = 11$:

$$p\text{-value} = P(T_{11} \ge 5.8425) \approx 0.0000572 \quad (5.72 \times 10^{-5})$$


Part 4: Conclude


AP Scoring Rubric Checklist for this Problem

Rubric Element Essential Requirements for "E" (Essentially Correct)
Section 1: State Must state $H_0$ and $H_a$ using correct symbol $\mu_d$ (or $\mu_{\text{diff}}$) AND explicitly define the direction of subtraction in context ($Post - Pre$).
Section 2: Plan Must identify the test by name OR formula. Must state and verify all 3 conditions (Randomness, 10% Rule, Normality of differences). Mentioning normality of original data instead of differences drops score to "P".
Section 3: Do Must display correct $t$-statistic ($t \approx 5.84$), correct $df = 11$, and $p\text{-value} \le 0.0001$.
Section 4: Conclude Must link $p$-value to $\alpha$, explicitly state decision (Reject $H_0$), and provide a contextual conclusion supporting $H_a$.

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Carnegie Mellon University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断