Statistics • Score 5 Strategy

Chi-Square Tests: Goodness-of-Fit, Homogeneity & Independence Guide: AP Statistics Score 5 for Harvard University

AP Statistics Mastery Guide: Categorical Inference & $\chi^2$ Tests

Target Institution: Harvard University
Exempted Course: Stat 100 (Introduction to Quantitative Methods for the Social Sciences and Humanities)
Acceleration Track: Stat 111 (Introduction to Statistical Inference)
Target Score: 5


1. Introduction & AP Exam Weight

Categorical data analysis via Chi-Square ($\chi^2$) procedures represents one of the most conceptually nuanced sections of the AP Statistics curriculum. Covered in Unit 8 ($\chi^2$ Goodness-of-Fit) and Unit 9 ($\chi^2$ Tests for Two-Way Tables), these topics account for approximately 7%–11% of the total AP Exam weight.

On the AP Statistics exam, scoring a 5 requires more than mere formulaic computation. AP Readers rigorously evaluate your ability to justify sample design, verify explicit conditions, differentiate between study structures, and articulate contextually precise conclusions.

For aspiring Harvard undergraduates, mastering categorical inference serves a dual purpose: 1. It guarantees the empirical maturity needed to secure a Score 5. 2. It satisfies the foundational competencies evaluated by the Department of Statistics, allowing placement out of Stat 100 and immediate enrollment in Stat 111 (Introduction to Statistical Inference), a calculus-based treatment of statistical theory.


2. Deep Concept Breakdown

The family of Chi-Square tests evaluates discrepancies between Observed counts ($O$) and Expected counts ($E$) under a specified null model. The unified test statistic across all three procedures is:

$$\chi^2 = \sum \frac{(O - E)^2}{E}$$

Where the summation runs over all $k$ categories or $r \times c$ cells in a contingency table.

                  ┌─────────────────────────────────────────┐
                  │       Categorical Data Analysis         │
                  └────────────────────┬────────────────────┘
                                       │
            ┌──────────────────────────┴──────────────────────────┐
            ▼                                                     ▼
    Single Population                                    Multiple Populations
  (One Sample Drawn)                                   (or Stratified/Experimental)
            │                                                     │
    ┌───────┴────────┐                                            │
    ▼                ▼                                            ▼
One Variable    Two Variables                             One Variable Across Groups
    │                │                                            │
    ▼                ▼                                            ▼
 Goodness-of-Fit  Independence                               Homogeneity
(df = k - 1)      (df = (r-1)(c-1))                           (df = (r-1)(c-1))

A. The Three Essential Chi-Square Tests

1. Chi-Square Goodness-of-Fit ($\chi^2$ GoF)

2. Chi-Square Test for Homogeneity

3. Chi-Square Test for Independence


B. Fundamental Assumptions & Inference Conditions

To earn full credit (E) on AP Free-Response Questions, you must explicitly state and verify the following three conditions:

  1. Categorical Data Condition: Data must consist of observed counts (frequency data), not percentages, proportions, or continuous measurements.
  2. Random Sampling / Assignment Condition:
  3. Samples must be selected via Random Sampling (SRS, Stratified, Cluster) or treatments randomly assigned in an experiment.
  4. Large Sample Size Condition (Expected Cell Counts):
  5. All expected cell counts must be at least 5 ($\forall E_{ij} \ge 5$).
  6. Critical AP Pitfall: You must write out the computed matrix of expected counts—stating "all expected counts are $\ge 5$" without providing the actual values yields an Automatic Partial (P) or Incorrect (I).
  7. 10% Condition (Independence of Observations):
  8. When sampling without replacement, sample size $n \le 0.10 N$ (where $N$ is total population size).

C. Computational Implementation (Python Engine)

Below is an annotated Python implementation utilizing standard statistical libraries (scipy.stats and numpy) to perform contingency table analysis, verify expected counts, extract cell-by-cell contributions, and execute hypothesis testing.

import numpy as np
from scipy import stats

def analyze_chi_square_contingency(observed_matrix):
    """
    Executes a Chi-Square Test for Homogeneity or Independence.
    Calculates expected counts, individual cell contributions, df, and p-value.

    Parameters:
        observed_matrix (list or np.ndarray): 2D array of observed counts.
    """
    O = np.array(observed_matrix, dtype=np.float64)

    # Perform standard Chi-Square Contingency Test
    chi2_stat, p_val, dof, E = stats.chi2_contingency(O, correction=False)

    # Calculate cell-by-cell contributions: (O - E)^2 / E
    contributions = ((O - E) ** 2) / E

    print("=== CHI-SQUARE CONTINGENCY ANALYSIS ===")
    print(f"Observed Counts Matrix:\n{O}\n")
    print(f"Expected Counts Matrix:\n{np.round(E, 2)}\n")

    # Verify Condition: All expected counts >= 5
    condition_met = np.all(E >= 5)
    print(f"Condition Check (All E_ij >= 5): {condition_met}")
    if not condition_met:
        failing_cells = np.argwhere(E < 5)
        print(f"  --> WARNING: Expected count < 5 at indices: {failing_cells}")

    print(f"\nCell Contributions Matrix ((O - E)^2 / E):\n{np.round(contributions, 4)}\n")
    print(f"Chi-Square Test Statistic (\u03c7\u00b2): {chi2_stat:.4f}")
    print(f"Degrees of Freedom (df): {dof}")
    print(f"p-value: {p_val:.6e}")

    return {
        "chi2_stat": chi2_stat,
        "p_value": p_val,
        "dof": dof,
        "expected": E,
        "contributions": contributions
    }

# Example Usage: 2x3 Contingency Table
sample_observed = [
    [45, 30, 25],  # Group A
    [20, 40, 40]   # Group B
]

results = analyze_chi_square_contingency(sample_observed)

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

To secure a Score 5, candidates must avoid specific technical errors that reduce student scores from a 5 to a 4.

       ┌────────────────────────────────────────────────────────┐
       │   Score 4 Student          │   Score 5 Student         │
       ├────────────────────────────┼───────────────────────────┤
       │ Confuses Homogeneity &     │ Distinguishes test by     │
       │ Independence tests         │ sampling design           │
       │                            │                           │
       │ Asserts conditions:        │ Displays explicit expected│
       │ "All E_i >= 5"             │ matrix: E_11=12.5, etc.   │
       │                            │                           │
       │ Writes H_0 using           │ States H_0 in contextual  │
       │ symbols like mu or p       │ words or distribution     │
       │                            │                           │
       │ Stops at p-value;          │ Conducts component cell   │
       │ gives generic conclusion   │ analysis for insight      │
       └────────────────────────────┴───────────────────────────┘

Critical Pitfalls

1. Sampling Design Ambiguity (Homogeneity vs. Independence)

2. Parameterization Errors in Hypotheses

3. Incomplete Verification of Expected Counts

4. Omission of Component Contribution Analysis


Scoring Rubric Nuance Comparison (AP FRQ Style)

Scenario

A study selects a single random sample of 200 adults to determine if an association exists between Education Level (High School, College, Advanced Degree) and Primary Source of News (TV, Online, Print).

                             Observed Data
┌──────────────────┬──────────┬──────────┬───────────┬─────────┐
│ News Source      │ High Sch │ College  │ Adv Degree│ Total   │
├──────────────────┼──────────┼──────────┼───────────┼─────────┤
│ Television       │    40    │    20    │    10     │   70    │
│ Online           │    30    │    50    │    35     │  115    │
│ Print            │    10    │    10    │     5     │   25    │
├──────────────────┼──────────┼──────────┼───────────┼─────────┤
│ Total            │    80    │    80    │    40     │  200    │
└──────────────────┴──────────┴──────────┴───────────┴─────────┘

Score 4 vs. Score 5 Exemplar Comparison

Why this loses points: Hypotheses are poorly stated (uses "equal" for categorical association); expected counts are asserted without providing values; conclusion lacks full population context.


4. Harvard University Placement Pathway

Mastering categorical inference provides a direct path to satisfying statistical rigor requirements at top-tier institutions like Harvard University.

                    ┌───────────────────────────────────────┐
                    │ AP Statistics Score 5                 │
                    │ Demonstrates fundamental inference    │
                    └──────────────────┬────────────────────┘
                                       │
                                       ▼
                    ┌───────────────────────────────────────┐
                    │ Harvard Stat 100 Placement Waiver     │
                    │ Waives introductory requirement       │
                    └──────────────────┬────────────────────┘
                                       │
                                       ▼
                    ┌───────────────────────────────────────┐
                    │ Direct Acceleration into Stat 111     │
                    │ "Introduction to Statistical Inference"│
                    └───────────────────────────────────────┘

Strategic Placement & Academic Advantage

  1. Course Exemption Mechanics: A score of 5 on AP Statistics qualifies students to fulfill or waive introductory requirements like Stat 100 (Introduction to Quantitative Methods for the Social Sciences and Humanities).
  2. Acceleration into Stat 111: Waiving Stat 100 allows freshman students interested in Quantitative Social Sciences, Applied Mathematics, Economics, or Computer Science to enter Stat 111 (Introduction to Statistical Inference) in their first year.
  3. Mathematical Rigor in Stat 111: Stat 111 frames the AP Chi-Square test within mathematical frameworks, including:
  4. Likelihood Ratio Tests (LRT) & Wilks' Theorem: Proving that $-2 \ln \Lambda \xrightarrow{d} \chi^2_k$ as $n \to \infty$.
  5. Multinomial Maximum Likelihood Estimation (MLE): Deriving expected counts as vector projections under the constrained parameter space $\Theta_0$.
  6. Asymptotic Variance & Fisher Information Matrices: Linking categorical contingency tables to covariance structures in Generalized Linear Models (GLMs).

Demonstrating strong precision with categorical data—such as correctly identifying degrees of freedom vector dimensions and cell residual structures—reflects the empirical literacy evaluated by quantitative departments at Harvard.


5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Problem Statement

A research team at a university health center wants to evaluate the efficacy of three different stress-reduction interventions: Mindfulness Meditation, Aerobic Exercise, and Cognitive Behavioral Therapy (CBT).

A total of $n = 300$ undergraduate students are recruited and randomly assigned in equal numbers ($100$ per group) to one of the three interventions. At the end of an 8-week program, each participant's self-reported stress level reduction is categorized into one of three levels: Low Improvement, Moderate Improvement, or High Improvement.

The observed data are summarized in the contingency table below:

                            Observed Improvement Counts
┌─────────────────────────┬─────────────────┬──────────────────────┬──────────────────┬────────┐
│ Intervention Group      │ Low Improvement │ Moderate Improvement │ High Improvement │ Total  │
├─────────────────────────┼─────────────────┼──────────────────────┼──────────────────┼────────┤
│ Meditation              │       42        │          38          │        20        │  100   │
│ Aerobic Exercise        │       28        │          42          │        30        │  100   │
│ CBT                     │       20        │          40          │        40        │  100   │
├─────────────────────────┼─────────────────┼──────────────────────┼──────────────────┼────────┤
│ Total                   │       90        │         120          │        90        │  300   │
└─────────────────────────┴─────────────────┴──────────────────────┴──────────────────┴────────┘

Task: Perform the appropriate statistical test to determine whether there is evidence of a difference in the distribution of stress level improvements among the three interventions. Use $\alpha = 0.01$.


Step-by-Step Solution & AP Scoring Checklist

Step 1: Identify the Appropriate Test & State Hypotheses


Step 2: Check Conditions for Inference

  1. Categorical Data: The measured data represent counts of individuals in discrete categories.
  2. Randomization: Participants were randomly assigned to the three treatment groups.
  3. 10% Condition: Not required here because this is a randomized experiment, not sampling without replacement from a finite population.
  4. Expected Cell Counts Condition ($\ge 5$):
  5. Expected Count Formula: $E_{ij} = \frac{(\text{Row Total}_i) \times (\text{Column Total}_j)}{\text{Grand Total}}$
  6. Computed Expected Counts:

$$\begin{aligned} E_{\text{Med, Low}} &= \frac{100 \times 90}{300} = 30.0 & E_{\text{Med, Mod}} &= \frac{100 \times 120}{300} = 40.0 & E_{\text{Med, High}} &= \frac{100 \times 90}{300} = 30.0 \ E_{\text{Exe, Low}} &= \frac{100 \times 90}{300} = 30.0 & E_{\text{Exe, Mod}} &= \frac{100 \times 120}{300} = 40.0 & E_{\text{Exe, High}} &= \frac{100 \times 90}{300} = 30.0 \ E_{\text{CBT, Low}} &= \frac{100 \times 90}{300} = 30.0 & E_{\text{CBT, Mod}} &= \frac{100 \times 120}{300} = 40.0 & E_{\text{CBT, High}} &= \frac{100 \times 90}{300} = 30.0 \end{aligned}$$

                            Expected Counts Matrix
┌─────────────────────────┬─────────────────┬──────────────────────┬──────────────────┐
│ Intervention Group      │ Low Improvement │ Moderate Improvement │ High Improvement │
├─────────────────────────┼─────────────────┼──────────────────────┼──────────────────┤
│ Meditation              │      30.0       │         40.0         │       30.0       │
│ Aerobic Exercise        │      30.0       │         40.0         │       30.0       │
│ CBT                     │      30.0       │         40.0         │       30.0       │
└─────────────────────────┴─────────────────┴──────────────────────┴──────────────────┘

Verification Statement: All expected cell counts are equal to $30.0$ or $40.0$, which are all $\ge 5$. The condition is met.


Step 3: Perform Calculations (Mechanics)

$$\begin{aligned} \chi^2 &= \frac{(42 - 30)^2}{30} + \frac{(38 - 40)^2}{40} + \frac{(20 - 30)^2}{30} \ &+ \frac{(28 - 30)^2}{30} + \frac{(42 - 40)^2}{40} + \frac{(30 - 30)^2}{30} \ &+ \frac{(20 - 30)^2}{30} + \frac{(40 - 40)^2}{40} + \frac{(40 - 30)^2}{30} \end{aligned}$$

$$\begin{aligned} \chi^2 &= \frac{144}{30} + \frac{4}{40} + \frac{100}{30} + \frac{4}{30} + \frac{4}{40} + \frac{0}{30} + \frac{100}{30} + \frac{0}{40} + \frac{100}{30} \ &= 4.800 + 0.100 + 3.333 + 0.133 + 0.100 + 0.000 + 3.333 + 0.000 + 3.333 \ &= 15.232 \end{aligned}$$


Step 4: Contextual Conclusion & Decision


AP Scoring Checklist Review

Rubric Component Scoring Requirement Status
Component 1: Hypotheses Stated in terms of distributions across groups; no parameter symbols ($\mu, p$). E
Component 2: Name & Conditions Correct test named ($\chi^2$ Homogeneity); explicit expected matrix displayed; $E_{ij} \ge 5$ verified. E
Component 3: Calculations Correct degrees of freedom ($df=4$), test statistic ($\chi^2=15.232$), and $p$-value ($0.00424$) shown. E
Component 4: Conclusion Decision linked directly to $p$-value vs. $\alpha$; contextually stated with cell contribution nuance. E

Final Score: 4/4 (Essentially Correct - E) $\longrightarrow$ On track for Score 5 performance.

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Harvard University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断