Statistics • Score 5 Strategy

Two-Sample t-Test vs. Matched Pairs Inference & Verification Guide: AP Statistics Score 5 for Stanford University

AP Statistics Mastery Guide: Two-Sample $t$-Test vs. Matched Pairs Inference

Target Level: AP Score 5 Specialist
Target Institution: Stanford University
Exempted Course: STATS 60: Introduction to Statistical Methods (5 Quarter Units)
Accelerated Track: CS 109: Probability for Computer Scientists or STATS 116: Theory of Probability


1. Introduction & AP Exam Weight

In Unit 7 of the AP Statistics curriculum (Inference for Quantitative Data: Means), which accounts for 10–18% of the multiple-choice section and appears in virtually every Free-Response Question (FRQ) section, no conceptual trap is more frequently used to separate Score 4 students from Score 5 candidates than the distinction between:

  1. Two-Independent-Sample $t$-Inference (comparing two distinct, independent populations or randomized treatment groups).
  2. Matched Pairs $t$-Inference (analyzing paired observations from a single sample or dependent units to control for subject-to-subject variance).

On the AP Statistics exam, misidentifying a matched-pairs design as a two-sample $t$-test results in an immediate "Incomplete" (I) or "Incorrect" rating on Section 1 (Hypotheses and Identification) and cascades into systematic errors across Conditions, Calculations, and Conclusions—costing up to 4 full raw FRQ points.

For high-achieving students aiming for Stanford University, mastering this distinction is not merely about securing an AP 5; it establishes the foundational rigor required for experimental design, causal inference, and variance reduction techniques utilized in AI benchmark evaluations (e.g., comparing model performance across paired evaluation sets in CS 109) and bioengineering experimental protocols.


2. Deep Concept Breakdown

A. Mathematical Foundations and Variance Reduction

The structural difference between a Two-Sample $t$-Test and a Matched Pairs $t$-Test lies in the dependence structure of the random variables and its direct impact on standard error via variance propagation.

1. Two-Independent-Sample Design

Let $X_1, X_2, \dots, X_{n_1}$ be an i.i.d. random sample from a population with mean $\mu_1$ and variance $\sigma_1^2$.
Let $Y_1, Y_2, \dots, Y_{n_2}$ be an i.i.d. random sample from an independent population with mean $\mu_2$ and variance $\sigma_2^2$.

We define the estimator for the difference in population means as $\bar{X} - \bar{Y}$. Since $X$ and $Y$ are independent, $\text{Cov}(\bar{X}, \bar{Y}) = 0$.

$$\text{Var}(\bar{X} - \bar{Y}) = \text{Var}(\bar{X}) + \text{Var}(\bar{Y}) = \frac{\sigma_1^2}{n_1} + \frac{\sigma_2^2}{n_2}$$

The standard error is given by:

$$SE(\bar{x}_1 - \bar{x}_2) = \sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}$$

The test statistic follows a $t$-distribution with degrees of freedom calculated via the Welch–Satterthwaite equation (or conservatively $\text{df} = \min(n_1 - 1, n_2 - 1)$):

$$t = \frac{(\bar{x}_1 - \bar{x}_2) - (\mu_1 - \mu_2)_0}{\sqrt{\frac{s_1^2}{n_1} + \frac{s_2^2}{n_2}}}$$

2. Matched Pairs Design

Let $(X_1, Y_1), (X_2, Y_2), \dots, (X_n, Y_n)$ be $n$ paired observations. $X_i$ and $Y_i$ are measured on the same subject or closely matched units, introducing a correlation $\rho_{XY} = \text{Corr}(X, Y) > 0$.

Instead of treating $X$ and $Y$ as distinct samples, we transform the bivariate data into a single univariate sample of differences:

$$D_i = X_i - Y_i \quad \text{for } i = 1, 2, \dots, n$$

The population parameters reduce to a single parameter: $\mu_d = E[D] = \mu_X - \mu_Y$.
The sample variance of $D$ demonstrates why pairing reduces noise:

$$\text{Var}(D) = \text{Var}(X - Y) = \text{Var}(X) + \text{Var}(Y) - 2\text{Cov}(X, Y) = \sigma_X^2 + \sigma_Y^2 - 2\rho_{XY}\sigma_X\sigma_Y$$

When subject-to-subject variability is large, $\rho_{XY} \gg 0$. By subtracting $2\text{Cov}(X,Y)$, the variance of the difference $\sigma_d^2$ becomes significantly smaller than the combined variance $\sigma_1^2 + \sigma_2^2$ of an independent design.

The standard error for the mean difference $\bar{x}_d$ is:

$$SE(\bar{x}_d) = \frac{s_d}{\sqrt{n_d}}$$

The test statistic follows a standard single-sample $t$-distribution with $\text{df} = n_d - 1$:

$$t = \frac{\bar{x}d - \mu{d,0}}{\frac{s_d}{\sqrt{n_d}}}$$


B. Computational Implementation & Simulation

The following Python code simulates paired data with high subject variability ($\rho \approx 0.85$) to demonstrate how a Matched Pairs $t$-Test correctly detects a subtle treatment effect by controlling for baseline variance, whereas a Two-Sample $t$-Test fails (Type II error).

import numpy as np
from scipy import stats

def simulate_matched_vs_twosample(seed=42):
    np.random.seed(seed)
    n = 30

    # Baseline variability among subjects (Nuisance variable)
    subject_baseline = np.random.normal(loc=100, scale=15, size=n)

    # Treatment effect (True mu_d = 3.5)
    true_effect = 3.5

    # Generate paired observations with noise
    pre_treatment = subject_baseline + np.random.normal(loc=0, scale=3, size=n)
    post_treatment = subject_baseline + true_effect + np.random.normal(loc=0, scale=3, size=n)

    # 1. Matched Pairs t-Test
    differences = post_treatment - pre_treatment
    mean_diff = np.mean(differences)
    std_diff = np.std(differences, ddof=1)
    se_paired = std_diff / np.sqrt(n)
    t_paired = mean_diff / se_paired
    p_paired = 1 - stats.t.cdf(t_paired, df=n-1)

    # 2. Independent Two-Sample t-Test (Incorrectly assuming independence)
    mean_pre = np.mean(pre_treatment)
    mean_post = np.mean(post_treatment)
    s_pre = np.std(pre_treatment, ddof=1)
    s_post = np.std(post_treatment, ddof=1)
    se_twosample = np.sqrt((s_pre**2 / n) + (s_post**2 / n))
    t_twosample = (mean_post - mean_pre) / se_twosample

    # Welch-Satterthwaite df
    df_welch = ((s_pre**2/n + s_post**2/n)**2) / (
        ((s_pre**2/n)**2 / (n-1)) + ((s_post**2/n)**2 / (n-1))
    )
    p_twosample = 1 - stats.t.cdf(t_twosample, df=df_welch)

    print(f"=== SIMULATION RESULTS (n = {n}) ===")
    print(f"Mean Difference (d_bar): {mean_diff:.4f}")
    print(f"Matched Pairs SE: {se_paired:.4f} | t-stat: {t_paired:.4f} | p-value: {p_paired:.6f}")
    print(f"Two-Sample SE:    {se_twosample:.4f} | t-stat: {t_twosample:.4f} | p-value: {p_twosample:.6f}")

if __name__ == "__main__":
    simulate_matched_vs_twosample()

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

Matrix of Diagnostic Differences

Feature Two-Sample $t$-Test Matched Pairs $t$-Test
Experimental Design Completely Randomized Design (2 independent groups) Blocked Design / Repeated Measures / Twin Matching
Parameter Definition $\mu_1 - \mu_2$: Difference between two un-paired population means $\mu_d$: Mean of differences within matched pairs
Data Structure Two lists of potentially unequal length ($n_1 \neq n_2$ allowed) One list of differences of equal length ($n_1 = n_2 = n_d$)
Normality Condition Check Examine $X_1$ AND $X_2$ distributions separately for skewness/outliers Examine the distribution of DIFFERENCES ($d_i$) only
Degrees of Freedom Complex Welch-Satterthwaite or conservative $\min(n_1-1, n_2-1)$ $n_d - 1$

Scoring Rubric Nuances: Score 4 vs. Score 5 Contrasts

To earn an Essentially Correct (E) rating across all components of an Inference FRQ, candidates must adhere to precise Rubric criteria.

1. Parameter Definition & Hypotheses

2. Verification of Conditions

For Matched Pairs Inference, evaluating conditions separately on $Group_1$ and $Group_2$ rather than on $Difference = Group_1 - Group_2$ is a common reason students miss out on a 5 score.

       [INCORRECT CONDITION CHECK]
       Student constructs two separate boxplots for 'Pre' and 'Post' 
       and states "Both samples are roughly symmetric." 
       --> Score: Partially Correct (P)

       [CORRECT CONDITION CHECK]
       Student calculates Difference_i = Post_i - Pre_i, constructs 
       a SINGLE histogram/boxplot of the DIFFERENCES, and checks for 
       strong skewness or outliers.
       --> Score: Essentially Correct (E)

Complete Checklist for Matched Pairs Conditions: 1. Paired Data / Random Sampling: Data must arise from a random sample of paired units OR a randomized block experiment where treatment order is randomly assigned within each pair. 2. Independence: * $10\%$ Condition: If sampling without replacement, $n_d \le 0.10 \times N_{\text{pairs}}$. * Individual pair differences must be independent of one another. 3. Normal Distribution of Differences: * If $n_d \ge 30$, apply Central Limit Theorem ($\bar{d}$ is approximately Normal). * If $n_d < 30$, plot the sample differences $d_i$. State explicitly: "The graph of sample differences shows no strong skewness or outliers, so the sampling distribution of $\bar{x}_d$ can be assumed roughly Normal."


4. Stanford University Placement Pathway

Mastering rigorous statistical inference provides direct institutional advantages for students entering Stanford University.

┌─────────────────────────────────────────────────────────┐
│              AP Statistics Exam (Score 5)              │
└───────────────────────────┬─────────────────────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│        Waive STATS 60 (5 Quarter Units Awarded)         │
│         Exempts General Quantitative Reasoning          │
└───────────────────────────┬─────────────────────────────┘
                            │
            ┌───────────────┴───────────────┐
            ▼                               ▼
┌───────────────────────┐       ┌─────────────────────────┐
│       CS 109          │       │        STATS 116        │
│    Probability for    │  OR   │  Theory of Probability  │
│  Computer Scientists  │       │  (Math/Data Sci Track)  │
└───────────┬───────────┘       └───────────┬─────────────┘
            │                               │
            └───────────────┬───────────────┘
                            │
                            ▼
┌─────────────────────────────────────────────────────────┐
│         Advanced Stanford Research Pathways:            │
│   • CS 229 (Machine Learning) Model Evaluations         │
│   • BioE Capstone Causal Inference Protocols            │
│   • Stanford AI Lab (SAIL) Empirical Benchmarks         │
└─────────────────────────────────────────────────────────┘

Strategic Academic Advantages:

  1. Credit Exemption: A score of 5 on AP Statistics satisfies the Stanford General Education Requirement (WAY-AQR) and grants 5 quarter units, waiving STATS 60: Introduction to Statistical Methods.
  2. Direct Entry into CS 109 / STATS 116: Stanford STEM majors (Computer Science, Data Science, Bioengineering, Mathematical & Computational Science) bypass introductory descriptive statistics and enroll directly in CS 109 or STATS 116.
  3. Research Application Nuance: Stanford research labs demand immediate fluency in controlling for confounders. Using paired experimental designs (e.g., controlling for user variation in algorithmic A/B testing or subject baseline expressions in single-cell RNA sequencing) directly applies the matched-pairs variance reduction principles established in AP Statistics.

5. High-Yield Practice Problem & Step-by-Step Solution Checklist

The Problem

A biomedical engineering team at a Stanford spin-off is testing a new algorithmic sensor designed to lower peak blood pressure during exercise. Ten randomly selected cardiac patients are fitted with the sensor. Each patient performs a standardized treadmill protocol under two conditions: Condition A (standard control software) and Condition B (new algorithmic feedback software).

To eliminate order effects, the software sequence (A then B vs. B then A) is randomly assigned to each patient with adequate rest between trials. The maximum systolic blood pressure (mmHg) recorded for each patient under both conditions is presented below:

Patient 1 2 3 4 5 6 7 8 9 10
Control Software (A) 162 175 158 180 169 172 165 182 170 167
Algorithmic Software (B) 154 168 159 171 161 166 160 175 163 160

Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the new algorithmic software reduces the mean maximum systolic blood pressure in cardiac patients during this treadmill protocol?


Model Response Solution Checklist (Score 5 Standards)

Step 1: Identify Method & Define Parameters

We will conduct a One-Sample $t$-Test for the Mean Difference (Matched Pairs $t$-Test).

Let $d_i = \text{Control}_i - \text{Algorithmic}_i$ represent the difference in maximum systolic blood pressure for patient $i$.
Let $\mu_d$ represent the true mean difference in maximum systolic blood pressure (Control $-$ Algorithmic) for cardiac patients undergoing this treadmill protocol.

Hypotheses: * $H_0: \mu_d = 0$ (The algorithmic software has no effect on mean peak systolic blood pressure.) * $H_a: \mu_d > 0$ (The algorithmic software reduces mean peak systolic blood pressure.)


Step 2: Calculate Sample Differences & Summary Statistics

We compute $d_i = A_i - B_i$ for each patient:

$$\text{Differences } (d_i): [+8, +7, -1, +9, +8, +6, +5, +7, +7, +7]$$


Step 3: Verify Inference Conditions

  1. Paired Design & Random Assignment: Data are paired by patient. Treatment sequence order (A then B vs. B then A) was randomly assigned to eliminate order effects.
  2. Independence / 10% Condition: The $10$ cardiac patients can be assumed to be less than $10\%$ of all cardiac patients who could undergo this protocol ($N \ge 100$).
  3. Normality of Differences: Since $n_d = 10 < 30$, we must examine the sample differences for skewness or outliers.
  4. Data plot check: The differences $d_i$ range from $-1$ to $+9$. A dotplot of differences reveals a single mild negative value but no extreme skewness or severe outliers. Thus, it is reasonable to assume the population of differences is approximately Normally distributed.

Step 4: Perform Test Statistic and $p$-Value Calculations


Step 5: Conclude in Context

Because the $p$-value ($\approx 0.000028$) is significantly less than the significance level $\alpha = 0.05$, we reject the null hypothesis $H_0$.

There is convincing statistical evidence that the true mean difference in maximum systolic blood pressure (Control $-$ Algorithmic) is greater than zero. We conclude that the new algorithmic software effectively reduces the mean maximum systolic blood pressure in cardiac patients during this exercise protocol.


AP Scoring Key (FRQ Evaluation Criteria)

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Stanford University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断