Statistics • Score 5 Strategy

Two-Sample t-Test vs. Matched Pairs Inference & Verification Guide: AP Statistics Score 5 for Caltech

AP Statistics Mastery Guide: Two-Sample $t$-Test vs. Matched Pairs Inference


1. Introduction & AP Exam Weight

In statistical inference, one of the most frequent points of points-loss on the AP Statistics Exam involves distinguishing between Independent Two-Sample $t$-Inference and Matched Pairs $t$-Inference. Governed primarily by Unit 7: Inference for Quantitative Data: Means (which constitutes 10–18% of the multiple-choice and free-response sections), this distinction tests your foundational understanding of experimental design, structural dependence, and variance reduction.

The core dilemma centers on how data are collected:

  1. Two Independent Samples: Two distinct, randomly selected groups (or two independent treatment arms) where observations in Sample 1 share no structural association with observations in Sample 2.
  2. Matched Pairs (Dependent Samples): A single sample measured under two distinct conditions (a "before-and-after" or repeated-measures design) or subject pairs linked by a confounding variable (e.g., identical twins, left vs. right eye, blocking on initial baseline scores).

On the AP Exam, misidentifying a Matched Pairs design as a Two-Sample design leads to an immediate loss of credit across all four scoring components (State, Plan, Do, Conclude). For a student aiming for a Score 5 and targeting elite quantitative benchmarks at institutions like Caltech, mastering this distinction is non-negotiable.


2. Deep Concept Breakdown

A. Mathematical Foundations & Variance Reduction Proof

The choice between a Two-Sample $t$-test and a Matched Pairs $t$-test alters how variance is partitioned.

Let $X_1$ and $X_2$ be random variables representing measurements under Condition 1 and Condition 2, respectively.

Two Independent Samples Strategy

If $X_1$ and $X_2$ are strictly independent, $\text{Cov}(X_1, X_2) = 0$. The sampling distribution of the difference between sample means $\bar{X}_1 - \bar{X}_2$ has expected value and variance:

$$\mathbb{E}[\bar{X}_1 - \bar{X}_2] = \mu_1 - \mu_2$$

$$\text{Var}(\bar{X}_1 - \bar{X}_2) = \text{Var}(\bar{X}_1) + \text{Var}(\bar{X}_2) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$

The standard error of the estimator is:

$$\text{SE}(\bar{X}_1 - \bar{X}_2) = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

The degrees of freedom $df$ are approximated using the Satterthwaite approximation:

$$df = \frac{\left(\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}\right)^2}{\frac{1}{n_1-1}\left(\frac{s_1^2}{n_1}\right)^2 + \frac{1}{n_2-1}\left(\frac{s_2^2}{n_2}\right)^2}$$

(Note: AP Statistics allows the conservative approximation $df = \min(n_1-1, n_2-1)$ when computing manually, though technology utilizes Satterthwaite).

Matched Pairs Strategy

When observations are paired, we define a new univariate random variable $D = X_1 - X_2$.

By the properties of linear combinations of random variables:

$$\text{Var}(D) = \text{Var}(X_1 - X_2) = \text{Var}(X_1) + \text{Var}(X_2) - 2\text{Cov}(X_1, X_2)$$

If $X_1$ and $X_2$ are positively correlated ($\text{Cov}(X_1, X_2) > 0$), which occurs naturally when measuring the same experimental unit twice or using homogeneous pairs:

$$\text{Var}(X_1 - X_2) < \text{Var}(X_1) + \text{Var}(X_2)$$

$$\implies \text{SE}(\bar{d}) = \frac{s_d}{\sqrt{n}} < \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

Mathematical Proof of Power Superiority: By reducing the noise (variance) introduced by subject-to-subject variability through pairing, $\text{SE}(\bar{d})$ shrinks relative to $\text{SE}(\bar{X}_1 - \bar{X}_2)$. For a given effect size $\delta = \mu_1 - \mu_2$, the test statistic:

$$t = \frac{\bar{d} - 0}{\frac{s_d}{\sqrt{n}}}$$

yields a higher magnitude $|t|$ than the independent two-sample equivalent, directly increasing the statistical power ($1 - \beta$) of the test to reject $H_0$ when a true effect exists.


B. Verification & Condition Checking Matrix

Condition Independent Two-Sample $t$-Test Matched Pairs $t$-Test
Randomness Two independent random samples OR random assignment of units to 2 treatment groups. Random sample of pairs OR random assignment of 2 treatments within each pair.
Independence $N_1 \ge 10n_1$ and $N_2 \ge 10n_2$ (if sampling without replacement). Samples must be independent of each other. $N_{\text{pairs}} \ge 10n_{\text{pairs}}$ (if sampling without replacement). Differences $d_i$ must be independent.
Normality / Sample Size Both populations normal, OR $n_1 \ge 30$ AND $n_2 \ge 30$ (CLT), OR both sample distributions are symmetric with no strong skewness/outliers. Population of differences is normal, OR $n_{\text{pairs}} \ge 30$ (CLT), OR sample of differences $d_i$ shows no strong skewness/outliers.

Critical Verification Rule: In matched pairs, you do not construct graphical displays for Sample 1 and Sample 2 separately. You must compute $d_i = x_{1i} - x_{2i}$ for each pair and construct a histogram, stem-and-leaf plot, or normal probability plot exclusively on the differences.


C. Computational Implementation: Power Analysis & Simulation

The following Python code simulates paired measurements subject to high inter-subject variance, contrasting the power and $p$-value computation between an independent two-sample $t$-test (which incorrectly ignores pairing) and a matched pairs $t$-test.

import numpy as np
from scipy import stats

# Seed for reproducibility
np.random.seed(42)

# Simulation Parameters
n_pairs = 20
baseline_variance = 100.0  # High subject-to-subject variance
treatment_effect = 2.5     # True underlying effect size
noise_variance = 1.0       # Low measurement noise

# Generate subject baseline traits (confounding subject effect)
subject_baselines = np.random.normal(loc=50.0, scale=np.sqrt(baseline_variance), size=n_pairs)

# Generate paired data: Condition 1 and Condition 2
cond1 = subject_baselines + np.random.normal(loc=0.0, scale=np.sqrt(noise_variance), size=n_pairs)
cond2 = subject_baselines + treatment_effect + np.random.normal(loc=0.0, scale=np.sqrt(noise_variance), size=n_pairs)

# 1. INCORRECT ANALYSIS: Two-Sample Independent t-Test
t_stat_ind, p_val_ind = stats.ttest_ind(cond2, cond1, equal_var=False)

# 2. CORRECT ANALYSIS: Matched Pairs t-Test
differences = cond2 - cond1
t_stat_paired, p_val_paired = stats.ttest_rel(cond2, cond1)

print(f"--- Independent Two-Sample t-Test ---")
print(f"SE Estimator : {np.sqrt(np.var(cond2, ddof=1)/n_pairs + np.var(cond1, ddof=1)/n_pairs):.4f}")
print(f"t-statistic  : {t_stat_ind:.4f}")
print(f"p-value      : {p_val_ind:.4f} (Fails to reject H0 at alpha = 0.05)\n")

print(f"--- Matched Pairs t-Test ---")
print(f"SE Estimator : {np.std(differences, ddof=1) / np.sqrt(n_pairs):.4f}")
print(f"t-statistic  : {t_stat_paired:.4f}")
print(f"p-value      : {p_val_paired:.6f} (Statistically Significant)")

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

Score 4 vs. Score 5 Performance Contrasts

On AP Statistics Free Response Questions (FRQs), the difference between a Score 4 (Essentially Correct with minor flaws) and a Score 5 (Complete, rigorous, contextually precise response) rests on three explicit nuances:

+-----------------------------------------------------------------------------------+
| SCORE 4 SOLUTION ATTRIBUTES                                                       |
| - Identifies test as "Two-Sample t-test" for paired data, but calculates correct t. |
| - Defines parameters as "mu_1 = mean after, mu_2 = mean before".                  |
| - Checks normality on Sample 1 and Sample 2 independently for paired data.        |
| - Concludes: "We reject H0. There is a difference between group 1 and group 2."    |
+-----------------------------------------------------------------------------------+
                                         vs
+-----------------------------------------------------------------------------------+
| SCORE 5 SOLUTION ATTRIBUTES                                                       |
| - Explicitly names the "Matched Pairs t-test for the mean difference".            |
| - Defines parameter: "mu_d = true mean difference in [measurement] (After - Before)"|
| - Computes differences d_i and plots histogram of DIFFERENCES to check skewness.  |
| - Concludes in context with explicit linkage: "Because p-value = 0.0003 < alpha,   |
|   we reject H0. There is convincing evidence that the true mean difference..."    |
+-----------------------------------------------------------------------------------+

Detailed AP Rubric Nuances

1. Parameter Definition Precise Language

2. The "Normality of Differences" Rule

3. Contextual Linkage in Conclusion


4. Caltech Placement Pathway

Quantitative Portfolio Evidence & Course Waiver

At Caltech, incoming freshmen who aim to waive introductory prerequisites and fast-track into advanced coursework present their AP scores alongside experimental research credentials.

[ AP Statistics Score 5 ] 
           +
[ Quantitative Portfolio ] ---> [ Waiver: Intro Modules ] ---> [ Accelerated Track: Ma 3 ]
(Error Propagation & Design)                                    (Inference & SURF Grant Prep)

Academic & Research Applications

  1. Experimental Rigor & Error Propagation: Caltech’s foundational laboratories (e.g., Physics 3, Chemistry 3) evaluate students on their ability to analyze empirical datasets with non-independent structure. Recognizing paired structures prevents erroneous estimation of experimental uncertainty.
  2. SURF (Summer Undergraduate Research Fellowships) Readiness: Caltech’s SURF program requires students to formulate hypotheses and evaluate signal-to-noise ratios in real experimental settings (e.g., LIGO interferometry noise subtraction, single-molecule fluorescence spectroscopy). Blocking and paired inference represent standard experimental techniques for controlling systematic variance across micro-arrays and detector channels.

5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Problem Statement

A Caltech materials science undergraduate lab is evaluating a new atomic layer deposition (ALD) coating designed to reduce the electrical resistivity of silicon micro-substrates. Twelve ($n = 12$) silicon wafers were sliced in half. One half of each wafer was randomly assigned to receive Treatment A (Standard Thermal Oxidation), while the matching half received Treatment B (ALD Coating). Resistivity ($\mu\Omega \cdot \text{cm}$) was measured for each half using a four-point microprobe.

The recorded measurements are presented below:

Wafer ID 1 2 3 4 5 6 7 8 9 10 11 12
Treatment A ($\text{Res}_{A}$) 14.2 15.8 13.9 16.5 14.8 15.1 16.0 14.5 15.3 14.9 15.7 16.2
Treatment B ($\text{Res}_{B}$) 13.5 15.0 13.2 15.9 14.1 14.7 15.2 14.1 14.6 14.3 15.0 15.4
Diff ($d_i = \text{Res}{A} - \text{Res}{B}$) 0.7 0.8 0.7 0.6 0.7 0.4 0.8 0.4 0.7 0.6 0.7 0.8

Task: Do these data provide convincing evidence at the $\alpha = 0.01$ significance level that the ALD Coating (Treatment B) yields a lower mean electrical resistivity than Standard Thermal Oxidation (Treatment A)? Execute a complete 4-step inference procedure.


Exemplary Step-by-Step Solution

Step 1: State

We wish to test the hypotheses:

$$H_0: \mu_d = 0$$ $$H_a: \mu_d > 0$$

where $\mu_d$ represents the true mean difference in electrical resistivity ($\mu\Omega \cdot \text{cm}$) between Treatment A and Treatment B ($\mu_d = \mu_{\text{Treatment A}} - \mu_{\text{Treatment B}}$) for silicon wafers of this type.

Significance level: $\alpha = 0.01$.


Step 2: Plan

Name of Procedure: Matched Pairs $t$-test for the mean difference $\mu_d$.

Conditions Verification: 1. Randomness: Halves of each wafer were randomly assigned to Treatment A and Treatment B, creating a matched-pairs experimental design. 2. Independence: Wafers are independently processed. The 12 sampled wafers represent less than 10% of all silicon wafers produced in this fabrication batch ($N \ge 120$). 3. Normality: Since $n = 12 < 30$, we must examine the distribution of sample differences $d_i$: $$\text{Differences}: {0.7, 0.8, 0.7, 0.6, 0.7, 0.4, 0.8, 0.4, 0.7, 0.6, 0.7, 0.8}$$

Dotplot Sketch of Differences: Value: 0.4 0.5 0.6 0.7 0.8 Dots : •• •• ••••• ••• Normality Evaluation Statement: The dotplot of the sample differences shows no extreme skewness and no outliers. Therefore, the conditions for performing a matched pairs $t$-test are met.


Step 3: Do

Sample Statistics: * Sample size of pairs: $n = 12$ * Mean of differences: $$\bar{d} = \frac{\sum d_i}{n} = \frac{7.5}{12} = 0.625\ \mu\Omega \cdot \text{cm}$$ * Standard deviation of differences: $$s_d = \sqrt{\frac{\sum (d_i - \bar{d})^2}{n - 1}} \approx 0.13568\ \mu\Omega \cdot \text{cm}$$

Standard Error Computation:

$$\text{SE}(\bar{d}) = \frac{s_d}{\sqrt{n}} = \frac{0.13568}{\sqrt{12}} \approx 0.03917$$

Test Statistic Calculation:

$$t = \frac{\bar{d} - \mu_0}{\text{SE}(\bar{d})} = \frac{0.625 - 0}{0.03917} \approx 15.957$$

Degrees of Freedom & $p$-value: * $df = n - 1 = 12 - 1 = 11$ * $p\text{-value} = P(t_{11} \ge 15.957) < 0.000001$ (specifically $p \approx 1.28 \times 10^{-8}$)


Step 4: Conclude

Because our $p$-value ($1.28 \times 10^{-8}$) is significantly less than the significance level $\alpha = 0.01$, we reject $H_0$.

There is statistically convincing evidence that the true mean difference in electrical resistivity ($\text{Treatment A} - \text{Treatment B}$) is greater than $0\ \mu\Omega \cdot \text{cm}$. Consequently, we conclude that the ALD coating (Treatment B) significantly reduces the mean electrical resistivity of silicon micro-substrates compared to Standard Thermal Oxidation (Treatment A).


Verification Checklist for AP Scoring Success

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Caltech with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断