Statistics • Score 5 Strategy

Two-Sample t-Test vs. Matched Pairs Inference & Verification Guide: AP Statistics Score 5 for UC Berkeley

AP Statistics Mastery Guide: Two-Sample $t$-Inference vs. Matched Pairs


1. Introduction & AP Exam Weight

In AP Statistics, few distinctions separate a Score 4 student from a Score 5 student as decisively as the structural differentiation between Two-Independent-Sample $t$-Inference and Matched Pairs $t$-Inference. Featured heavily across Unit 6 (Inference for Categorical Data/Means) and Unit 7 (Inference for Quantitative Data: Means), inference for quantitative data accounts for 10%–15% of the multiple-choice section and frequently anchors Question 4 or 5 on the Free-Response Section (FRQ).

The primary conceptual hurdle is recognizing experimental design and sampling architecture: * A Two-Sample $t$-test compares the central tendencies of two completely independent populations or treatments. * A Matched Pairs $t$-test reduces extraneous variability (nuisance variables) by pairing observational units or applying two treatments to the same unit, thereby transforming a two-variable problem into a single-sample inference on difference scores ($d$).

Misidentifying a matched pairs design as a two-sample design—or vice-versa—results in an incorrect standard error structure, incorrect degrees of freedom, and an automatic downgrade on the FRQ scoring rubric from Essentially Correct (E) to Partially Correct (P) or Incomplete (I).


2. Deep Concept Breakdown

A. Mathematical Foundations & Variance Derivation

To understand why misclassifying these tests breaks inference, we examine the variance of linear combinations of random variables.

1. Two-Independent-Sample Model

Let $X_1$ and $X_2$ be independent random variables representing populations 1 and 2, with means $\mu_1, \mu_2$ and variances $\sigma_1^2, \sigma_2^2$. For independent random samples of sizes $n_1$ and $n_2$:

$$\text{E}(\bar{X}_1 - \bar{X}_2) = \mu_1 - \mu_2$$

Because $X_1$ and $X_2$ are independent, $\text{Cov}(X_1, X_2) = 0$. Therefore, the variance of the difference between the sample means is purely additive:

$$\text{Var}(\bar{X}_1 - \bar{X}_2) = \text{Var}(\bar{X}_1) + \text{Var}(\bar{X}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$

The standard error ($SE$) is estimated using sample standard deviations $s_1$ and $s_2$:

$$SE_{(\bar{x}_1 - \bar{x}_2)} = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

Degrees of Freedom ($df$): Calculated via the Welch–Satterthwaite equation:

$$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{1}{n_1-1}\left(\frac{s_1^2}{n_1}\right)^2 + \frac{1}{n_2-1}\left(\frac{s_2^2}{n_2}\right)^2}$$

(Note: On the AP Exam, using $\min(n_1-1, n_2-1)$ as a conservative bound is acceptable if stated explicitly, though technology-generated Welch $df$ is preferred).


2. Matched Pairs Model (Dependent Samples)

Let $(X_{1,i}, X_{2,i})$ represent paired observations on the $i$-th subject or matched pair. We define the difference variable:

$$D_i = X_{1,i} - X_{2,i}$$

If $X_1$ and $X_2$ are dependent, $\text{Cov}(X_1, X_2) \neq 0$:

$$\text{Var}(D) = \text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$

When positive correlation exists within pairs ($\text{Cov}(X_1, X_2) > 0$), pairing reduces overall variance compared to the independent model.

We compress the data into a single sample of differences ${d_1, d_2, \dots, d_n}$ with sample mean $\bar{x}_d$ and sample standard deviation $s_d$.

The standard error of the mean difference $\mu_d$ is:

$$SE_{\bar{x}_d} = \frac{s_d}{\sqrt{n}}$$

Degrees of Freedom ($df$): Simply $df = n - 1$, where $n$ is the number of pairs (not total observations).


B. Condition Verification Matrix

Condition Two-Sample $t$-Test ($\mu_1 - \mu_2$) Matched Pairs $t$-Test ($\mu_d$)
Randomness Two independent Random Samples OR Random Assignment to 2 treatments. Random Sample of paired units OR Random Assignment of treatment order within units.
Independence 10% Condition: $n_1 \le 0.10 N_1$ AND $n_2 \le 0.10 N_2$ (if sampling without replacement). Samples must be mutually independent. 10% Condition: $n_{\text{pairs}} \le 0.10 N_{\text{pairs}}$ (if sampling without replacement). Differences must be independent.
Normal / Large Sample Both populations normal OR $n_1 \ge 30$ AND $n_2 \ge 30$ (CLT). If $n_1, n_2 < 30$, check both sample distributions for strong skewness or outliers. Population of differences normal OR $n_{\text{pairs}} \ge 30$ (CLT). If $n_{\text{pairs}} < 30$, plot the differences ($d_i$) to verify no strong skewness or outliers.

C. Computational Implementation: Identifying Misclassification Errors

The Python script below illustrates how misclassifying paired data as two independent samples drastically understates statistical power by ignoring intra-subject correlation ($\text{Cov}(X_1, X_2)$).

import numpy as np
from scipy import stats

# Seed for reproducibility
np.random.seed(42)

# Generate synthetic paired data: Pre-test vs Post-test score on the same 15 subjects
# Strong baseline subject variability (latent variable)
subject_baseline = np.random.normal(loc=100, scale=15, size=15)
pre_scores = subject_baseline + np.random.normal(loc=0, scale=3, size=15)
post_scores = subject_baseline + 5 + np.random.normal(loc=0, scale=3, size=15) # True effect = +5

# Calculate differences
differences = post_scores - pre_scores

print("=== CORRECT ANALYSIS: Matched Pairs t-Test ===")
t_stat_paired, p_val_paired = stats.ttest_rel(post_scores, pre_scores)
df_paired = len(differences) - 1
print(f"Mean Diff (x_bar_d): {np.mean(differences):.3f}")
print(f"Std Dev of Diff (s_d): {np.std(differences, ddof=1):.3f}")
print(f"t-statistic: {t_stat_paired:.4f} | df: {df_paired} | p-value: {p_val_paired:.5f}\n")

print("=== INCORRECT ANALYSIS: Independent Two-Sample t-Test ===")
t_stat_ind, p_val_ind = stats.ttest_ind(post_scores, pre_scores, equal_var=False)
print(f"t-statistic: {t_stat_ind:.4f} | p-value: {p_val_ind:.5f}")
print("Notice how misidentifying the test inflates the p-value across the critical threshold (0.05)!")

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

To secure a Score 5, your Free-Response responses must satisfy the rigorous rubric guidelines used by AP Readers.

+-----------------------------------------------------------------------------------+
|                                  SCORE LEVEL BREAKDOWN                            |
+-----------------------------------------------------------------------------------+
| SCORE 4 SOLUTION:                                                                 |
| - Identifies tests correctly but uses ambiguous notation (e.g., defines μ1, μ2     |
|   instead of μ_d for paired data).                                                |
| - Verifies conditions globally without plotting sample difference data when n < 30.|
| - Arrives at correct p-value via calculator command without showing formula.     |
+-----------------------------------------------------------------------------------+
| SCORE 5 SOLUTION:                                                                 |
| - Explicitly defines parameters in context (e.g., "μ_d = population mean          |
|   difference in reaction time (After - Before)").                                 |
| - Explicitly states the order of subtraction (d = X1 - X2).                       |
| - Draws a dotplot/boxplot of differences when n_pairs < 30 to verify normality.  |
| - States named test, displays test statistic formula with plugged-in values,      |
|   provides precise df, p-value, and frames conclusion in strict context.          |
+-----------------------------------------------------------------------------------+

Critical AP Rubric Pitfalls

  1. Parameter Mislabeling Trap:
  2. Incorrect: $H_0: \mu_1 = \mu_2$ vs $H_a: \mu_1 \neq \mu_2$ for paired data.
  3. Correct: $H_0: \mu_d = 0$ vs $H_a: \mu_d \neq 0$, where $\mu_d$ is defined as the true mean difference in [Variable] between [Condition A] and [Condition B].

  4. Graphing Condition Verification:

  5. When checking the Normal/Large Sample condition for a matched pairs test with $n < 30$, you cannot draw two separate graphs for Sample 1 and Sample 2. You must compute the vector of differences $d_i$ and sketch a dotplot, boxplot, or stem-and-leaf plot of the differences.

  6. Failure to Link $p$-value to Conclusion:

  7. Simply writing "Reject $H_0$" loses credit. You must write: > "Because $p\text{-value} = 0.023 < \alpha = 0.05$, we reject $H_0$. There is convincing statistical evidence that the true mean difference in [Variable] is [greater than/less than/different from] zero."

4. UC Berkeley Placement Pathway

Mastering AP Statistics to earn a Score 5 delivers tangible academic acceleration at UC Berkeley, particularly within the College of Computing, Data Science, and Society (CDSS) and the Haas School of Business.

                           [ AP Statistics Score 5 ]
                                       │
                                       ▼
                  [ Waive STAT 2 / STAT 20 (4 Semester Units) ]
                                       │
                                       ▼
                  [ Fulfills CDSS Core Statistics Prerequisite ]
                                       │
                      ┌────────────────┴────────────────┐
                      ▼                                 ▼
         [ DATA 8: Foundations of ]          [ CS 61A: Structure & Interp. ]
               Data Science                          of Computer Programs
                      │                                 │
                      └────────────────┬────────────────┘
                                       ▼
                       [ DATA 100: Principles & Techniques ]
                                of Data Science

Academic & Curriculum Impact


5. High-Yield Practice Problem & Step-by-Step Solution Checklist

The Problem (AP-Style Free Response)

An agricultural research station is evaluating whether a new biological leaf-coating spray reduces the population of a specific crop pest. Ten plots of land are selected. Each plot is divided into two equal sub-plots. One sub-plot is randomly assigned to receive the Control (no spray), and the other sub-plot receives the Biological Spray.

After three weeks, the pest counts per square meter are measured on both sub-plots:

Plot 1 2 3 4 5 6 7 8 9 10
Control ($X_C$) 45 52 38 67 41 59 48 55 62 50
Spray ($X_S$) 39 48 39 58 35 53 47 49 54 42

Is there convincing statistical evidence at the $\alpha = 0.05$ significance level that the Biological Spray reduces the mean pest count per square meter?


Step-by-Step Solution Checklist (4-Step AP Framework)

Step 1: STATE


Step 2: PLAN


Step 3: DO

$$t = \frac{\bar{x}_d - \mu_0}{\frac{s_d}{\sqrt{n}}} = \frac{5.30 - 0}{\frac{2.983}{\sqrt{10}}} = \frac{5.30}{0.9433} \approx 5.619$$

$$p\text{-value} = P(t_{9} \ge 5.619) \approx 0.00016$$


Step 4: CONCLUDE

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like UC Berkeley with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断