Statistics • Score 5 Strategy

Two-Sample t-Test vs. Matched Pairs Inference & Verification Guide: AP Statistics Score 5 for Harvard University

AP Statistics Guide: Two-Sample $t$-Test vs. Matched Pairs Inference & Verification


1. Introduction & AP Exam Weight

In quantitative analysis, determining whether two sets of numerical observations are independent or dependent is a core statistical threshold. On the AP Statistics Exam, Inference for Quantitative Data: Means (Units 6 & 7) accounts for 10–18% of the multiple-choice section and appears consistently on the Free-Response section (FRQs).

The distinction between a Two-Sample $t$-Test for the Difference of Two Means ($\mu_1 - \mu_2$) and a Matched Pairs $t$-Test for a Mean Difference ($\mu_d$) is one of the most frequent discriminators between a Score 4 and a Score 5. AP Readers frequently grade FRQs where students apply a two-sample procedure to dependent data or fail to state conditions in terms of sample differences.

Mastering this distinction requires a solid grasp of mathematical probability: recognizing how pairing controls for extraneous variation by exploiting covariance, and verifying conditions on the correct statistical distribution.


2. Deep Concept Breakdown

Mathematical Foundations: The Role of Covariance

Consider two random variables $X_1$ and $X_2$ representing two measurements. The variance of their difference is defined as:

$$\text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$

Case 1: Independent Samples (Two-Sample $t$-Test)

When observations in Group 1 are independent of observations in Group 2, $\text{Cov}(X_1, X_2) = 0$. The variance of the sample mean difference $\bar{x}_1 - \bar{x}_2$ simplifies to:

$$\text{Var}(\bar{x}_1 - \bar{x}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$

Replacing parameters with sample variances gives the Standard Error of $\bar{x}_1 - \bar{x}_2$:

$$SE_{\bar{x}_1 - \bar{x}_2} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

The test statistic is calculated as:

$$t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)_0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$$

Degrees of freedom ($df$) are calculated using the Satterthwaite approximation:

$$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{\left(s_1^2/n_1\right)^2}{n_1 - 1} + \frac{\left(s_2^2/n_2\right)^2}{n_2 - 1}}$$

(Note: AP Statistics permits the conservative lower-bound estimate $df = \min(n_1 - 1, n_2 - 1)$, though technological Satterthwaite output is preferred).

Case 2: Dependent/Paired Samples (Matched Pairs $t$-Test)

When observations are paired (e.g., pre-test vs. post-test on the same subject, or twin studies), $X_1$ and $X_2$ are positively correlated ($\text{Cov}(X_1, X_2) > 0$). Subtracting $2\text{Cov}(X_1, X_2)$ reduces the variability of the difference, removing subject-to-subject baseline variation.

We reduce the bivariate data $(x_{1i}, x_{2i})$ to a single univariate sample of differences:

$$d_i = x_{1i} - x_{2i} \quad \text{for } i = 1, 2, \dots, n_d$$

The hypotheses are stated in terms of the single parameter $\mu_d$, the population mean difference:

$$H_0: \mu_d = \Delta_0 \quad (\text{typically } \Delta_0 = 0)$$

The Standard Error of $\bar{x}_d$ is:

$$SE_{\bar{x}_d} = \frac{s_d}{\sqrt{n_d}}$$

where $s_d = \sqrt{\frac{\sum (d_i - \bar{x}_d)^2}{n_d - 1}}$. The test statistic follows a standard $t$-distribution with $df = n_d - 1$:

$$t = \frac{\bar{x}d - \mu{d,0}}{\frac{s_d}{\sqrt{n_d}}}$$


Condition Verification Comparison

Condition Two-Sample $t$-Test ($\mu_1 - \mu_2$) Matched Pairs $t$-Test ($\mu_d$)
Randomness Two independent random samples OR random assignment of units to two treatment groups. A single random sample of paired units OR random assignment of treatments within pairs.
10% Rule $n_1 \le 0.10 N_1$ and $n_2 \le 0.10 N_2$ (if sampling without replacement). $n_d \le 0.10 N_{\text{pairs}}$ (if sampling without replacement).
Normality / Large Sample Both populations normal OR $n_1 \ge 30$ and $n_2 \ge 30$. If $n_1, n_2 < 30$, inspect both sample distributions for strong skew/outliers. Population of differences is normal OR $n_d \ge 30$. If $n_d < 30$, inspect the single distribution of sample differences $d_i$ for strong skew/outliers.

Computational Simulation: Variance Reduction via Pairing

The following Python snippet demonstrates how blocking (pairing) isolates variation and increases statistical power compared to an un paired analysis:

import numpy as np
from scipy import stats

# Seed for reproducibility
np.random.seed(42)

# Simulate 20 subjects with high baseline variability (Subject effect ~ N(100, 15^2))
subject_baselines = np.random.normal(100, 15, 20)

# Treatment adds a subtle true mean improvement of mu_d = 3.5
# Individual response variance = 2.0^2
pre_treatment = subject_baselines + np.random.normal(0, 2.0, 20)
post_treatment = subject_baselines + 3.5 + np.random.normal(0, 2.0, 20)

# 1. Incorrect Approach: Independent Two-Sample t-test
t_two_sample, p_two_sample = stats.ttest_ind(post_treatment, pre_treatment)

# 2. Correct Approach: Matched Pairs t-test
differences = post_treatment - pre_treatment
t_paired, p_paired = stats.ttest_rel(post_treatment, pre_treatment)

print(f"--- INDEPENDENT TWO-SAMPLE t-TEST ---")
print(f"t-statistic: {t_two_sample:.4f} | p-value: {p_two_sample:.4f}")
print(f"Outcome: {'Fail to Reject H0' if p_two_sample > 0.05 else 'Reject H0'}\n")

print(f"--- MATCHED PAIRS t-TEST ---")
print(f"t-statistic: {t_paired:.4f} | p-value: {p_paired:.4f}")
print(f"Outcome: {'Fail to Reject H0' if p_paired > 0.05 else 'Reject H0'}")

Output Insights: The two-sample test fails to reject $H_0$ because the variance from subject baselines inflates the standard error denominator. The paired $t$-test eliminates baseline variance by operating solely on differences, lowering standard error and detecting the treatment effect.


3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

Score 4 vs. Score 5 Student Response Architecture

                               AP INFERENCE FRQ EVALUATION
                                            │
               ┌────────────────────────────┴────────────────────────────┐
               ▼                                                         ▼
       SCORE 4 RESPONSE                                          SCORE 5 RESPONSE
 ──────────────────────────────                            ──────────────────────────────
 • Selects Two-Sample t-Test                              • Identifies Matched Pairs
   for paired data.                                         design based on structure.
 • Checks Normality on raw                                 • Explicitly defines mu_d:
   Group 1 and Group 2 data.                                "mean difference in [units]".
 • States generic hypotheses:                              • Graphs and verifies Normal
   H0: mu1 = mu2.                                           condition ONLY on differences.
 • Omits context in final                                  • States decision relative
   p-value linkage.                                         to alpha with full context.

Free-Response Scoring Rubric Breakdown (Four Essential Components)

To earn an Essentially Correct (E) on an AP Statistics inference question, your solution must satisfy four distinct components:

1. Statement of Hypotheses & Parameter Definition

2. Identification of Procedure & Verification of Conditions

3. Correct Mechanics & Statistics

4. Contextual Conclusion Linked to $P$-Value


4. Harvard University Placement Pathway

Mastering statistical inference yields concrete academic advantages at elite institutions such as Harvard University.

                         HARVARD QUANTITATIVE PATHWAY
                                       │
                       AP Statistics Exam (Score 5)
                                       │
                 ┌─────────────────────┴─────────────────────┐
                 ▼                                           ▼
      EXEMPTED FROM STAT 100                    DIRECT ADVANCED PLACEMENT
 (General Quantitative Reasoning)                           │
                                         ┌───────────────────┴───────────────────┐
                                         ▼                                       ▼
                                     STAT 110                                STAT 111
                            (Introduction to Probability)           (Statistical Inference)

Course Exemption & Placement Analysis

Strategic Academic Advantage

Understanding the core distinction between independent sample variance and paired block design ($\text{Cov}(X_1, X_2)$ variance reduction) directly prepares students for experimental design and econometric modeling: * In Harvard’s Economics Department (e.g., Econ 1123 / Econ 1126), the matched pairs framework serves as the conceptual foundation for Difference-in-Differences (DiD) estimation and panel data fixed-effects estimators. * In Government/Data Science tracks (e.g., Gov 50 / Gov 2001), controlling for subject-level confounders via matched pair matching algorithms (e.g., propensity score matching) is fundamental to establishing modern causal inference under potential outcomes modeling.


5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Problem Statement

A clinical research team at a university medical center investigates whether a specialized cognitive training software reduces the time (in seconds) required for older adults to complete a complex spatial task.

A random sample of 10 participants aged 65 or older is selected. Each participant's task completion time is measured prior to training ($x_{\text{pre}}$) and again after completing a 4-week training regimen ($x_{\text{post}}$). The table below displays the completion times in seconds:

Subject 1 2 3 4 5 6 7 8 9 10
Pre-Training ($x_{\text{pre}}$) 45.2 38.0 52.1 41.5 49.8 44.0 51.3 39.7 46.5 43.0
Post-Training ($x_{\text{post}}$) 41.0 36.5 47.0 41.8 44.2 40.1 46.5 38.0 42.1 39.5

Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the training software reduces the mean completion time for this population?


Complete Step-by-Step Solution Checklist

Step 1: Identify Parameter and Hypotheses

Let $d_i = x_{\text{pre}, i} - x_{\text{post}, i}$ denote the pairwise difference in time (in seconds) for subject $i$.

Define $\mu_d$ as the true population mean difference in completion time (Pre $-$ Post) for adults aged 65 and older undergoing this training program.

(Note: Defining $d = \text{Pre} - \text{Post}$ means a reduction in time corresponds to $d > 0$.)


Step 2: Identify Procedure and Verify Conditions

Procedure Name: Matched Pairs $t$-test for a Mean Difference ($\mu_d$).

Calculate the 10 pairwise difference values $d_i = x_{\text{pre}, i} - x_{\text{post}, i}$:

$$\text{Differences } (d_i): {+4.2, +1.5, +5.1, -0.3, +5.6, +3.9, +4.8, +1.7, +4.4, +3.5}$$

Condition Verification: 1. Random Sample: The problem explicitly states that a random sample of 10 participants aged 65 or older was selected. 2. 10% Condition: $n_d = 10 \le 0.10 \times (\text{Total population of adults }\ge 65)$. This is reasonable to assume. 3. Normality of Differences: Since sample size $n_d = 10 < 30$, we must construct a plot of sample differences to verify the distribution shows no severe skewness or outliers.

       DOTPLOT OF DIFFERENCES (d_i)

       *                  *  *  *  *     *
───┼───┼───┼───┼───┼───┼───┼───┼───┼───┼───┼───
  -1   0   1   2   3   4   5   6
                     d_i (seconds)

Verification Statement: The dotplot of the 10 sample differences shows no extreme outliers or strong skewness. Therefore, the condition that the population of differences is approximately normal is satisfied.


Step 3: Compute Test Mechanics

Calculate sample statistics from the difference vector $d_i$:

$$\bar{x}_d = \frac{\sum d_i}{n_d} = \frac{34.4}{10} = 3.44 \text{ seconds}$$

$$s_d = \sqrt{\frac{\sum (d_i - \bar{x}_d)^2}{n_d - 1}} = 1.833 \text{ seconds}$$

$$SE_{\bar{x}_d} = \frac{s_d}{\sqrt{n_d}} = \frac{1.833}{\sqrt{10}} \approx 0.5796$$

Compute Test Statistic $t$:

$$t = \frac{\bar{x}_d - 0}{\frac{s_d}{\sqrt{n_d}}} = \frac{3.44 - 0}{0.5796} \approx 5.935$$

Degrees of Freedom:

$$df = n_d - 1 = 10 - 1 = 9$$

Calculate $P$-value:

$$P\text{-value} = P(T_9 \ge 5.935) < 0.0001$$

(Using $t$-table or calculator: $\text{tcdf}(5.935, \infty, 9) \approx 0.000112$)


Step 4: Write Conclusion in Context

Because the $P\text{-value} \approx 0.00011$ is less than the significance level $\alpha = 0.05$, we reject $H_0$.

There is convincing statistical evidence that the cognitive training software reduces the true mean task completion time (in seconds) for older adults aged 65 and older.

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Harvard University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断