Statistics • Score 5 Strategy

Chi-Square Tests: Goodness-of-Fit, Homogeneity & Independence Guide: AP Statistics Score 5 for Stanford University

AP Statistics Exam Mastery Guide: Chi-Square Tests

Target Topic: Inference for Categorical Data (Goodness-of-Fit, Homogeneity, and Independence)
Target Institution: Stanford University (Aiming for Score 5)


1. Introduction & AP Exam Weight

Categorical data analysis forms a cornerstone of statistical inference. On the AP Statistics Exam, Unit 8: Inference for Categorical Data: Chi-Square accounts for 7% to 11% of the multiple-choice section and routinely features in Free-Response Questions (FRQs)—either as a standalone 11-minute problem or integrated into Question 6 (the Investigation Task).

Mastery of Chi-Square ($\chi^2$) inferential procedures requires more than memorizing algorithms; it demands a precise understanding of study design, sampling mechanics, and probability distributions.

At Stanford University, achieving a score of 5 on AP Statistics awards 5 quarter-units of credit, waiving STATS 60: Introduction to Statistical Methods. This direct exemption allows prospective Computer Science, Data Science, Bioengineering, and Pre-Med students to accelerate into advanced coursework such as CS 109 (Probability for Computer Scientists) or STATS 141 (Biostatistics).


2. Deep Concept Breakdown

2.1 Mathematical Foundations of the Chi-Square Distribution

The Chi-Square distribution with $k$ degrees of freedom is defined as the distribution of a sum of the squares of $k$ independent standard normal random variables. If $Z_1, Z_2, \dots, Z_k \sim \text{N}(0, 1)$ independently, then:

$$Q = \sum_{i=1}^{k} Z_i^2 \sim \chi^2_k$$

The Probability Density Function (PDF) of a $\chi^2$ random variable with $k$ degrees of freedom is given by:

$$f(x; k) = \frac{x^{(k/2) - 1} e^{-x/2}}{2^{k/2} \Gamma\left(\frac{k}{2}\right)}, \quad x \ge 0$$

where $\Gamma(n)$ is the Gamma function. The mean of $\chi^2_k$ is $k$, and its variance is $2k$. As $k \to \infty$, by the Central Limit Theorem, the $\chi^2_k$ distribution approaches a Normal distribution $\text{N}(k, \sqrt{2k})$.


2.2 Derivation of the Pearson $\chi^2$ Test Statistic

For a sample divided into $k$ mutually exclusive categorical bins with observed counts $O_1, O_2, \dots, O_k$ and expected counts $E_1, E_2, \dots, E_k$ under a null model $H_0$, the sample count vector follows a Multinomial Distribution.

When the total sample size $n = \sum O_i$ is large, each observed count $O_i$ is approximately normally distributed:

$$O_i \sim \text{N}\left(E_i, \sqrt{E_i(1 - p_i)}\right)$$

Standardizing the residual yields $Z_i = \frac{O_i - E_i}{\sqrt{E_i}}$. Squaring and summing these standardized residuals over all $k$ categories yields the classic Pearson $\chi^2$ test statistic:

$$\chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i}$$

Because the constraint $\sum_{i=1}^{k} O_i = n$ removes one degree of freedom, the statistic $\chi^2$ asymptotically follows a $\chi^2$ distribution with $df = k - 1$.


2.3 Comparative Analysis of the Three Chi-Square Tests

The AP Statistics exam evaluates three distinct $\chi^2$ tests. Choosing the correct test depends on the sampling strategy and the number of variables/populations.

Structural Metric $\chi^2$ Goodness-of-Fit Test $\chi^2$ Test for Homogeneity $\chi^2$ Test for Independence
Number of Categorical Variables $1$ variable $1$ variable $2$ variables
Number of Samples / Populations $1$ sample vs. theoretical model $2+$ independent samples/groups $1$ single sample
Null Hypothesis ($H_0$) Formulation $H_0$: The population distribution matches specified probabilities $p_1, p_2, \dots, p_k$. $H_0$: The distribution of [Variable] is the same across all populations. $H_0$: Variable A and Variable B are independent in the population.
Degrees of Freedom ($df$) $df = k - 1$ $df = (r - 1)(c - 1)$ $df = (r - 1)(c - 1)$
Expected Cell Count Formula $E_i = n \cdot p_i$ $E_{i,j} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Table Total}}$ $E_{i,j} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Table Total}}$

2.4 Necessary Conditions for Statistical Inference

Before executing any $\chi^2$ test, all three of the following conditions must be explicitly verified:

  1. Random Condition: Data must come from a well-designed random sample (SRS, stratified, cluster) or a randomized experiment.
  2. 10% Condition (Independence): When sampling without replacement from a finite population, the total sample size must satisfy $n \le 0.10 N$. (Note: This condition does NOT apply to randomized experiments).
  3. Large Counts Condition: All expected counts must be at least 5 ($E_i \ge 5$ or $E_{i,j} \ge 5$).
  4. Critical Pitfall: This condition applies to expected cell counts, not observed counts.

2.5 Python Implementation for Computation and Verification

import numpy as np
from scipy.stats import chi2_contingency, chisquare

def run_chi_square_tests():
    # -------------------------------------------------------------
    # 1. Chi-Square Goodness-of-Fit Test
    # -------------------------------------------------------------
    # Observed counts of genomic mutations across 4 loci
    observed_gof = np.array([32, 28, 15, 25])
    # Expected proportions under Mendelian inheritance (3:3:2:2 ratio -> 0.3, 0.3, 0.2, 0.2)
    expected_props = np.array([0.30, 0.30, 0.20, 0.20])
    total_n = np.sum(observed_gof)
    expected_gof = total_n * expected_props

    gof_stat, gof_p_val = chisquare(f_obs=observed_gof, f_exp=expected_gof)

    print(f"--- Goodness-of-Fit Test ---")
    print(f"Chi2 Stat: {gof_stat:.4f} | p-value: {gof_p_val:.4f} | df: {len(observed_gof)-1}\n")

    # -------------------------------------------------------------
    # 2. Chi-Square Test for Homogeneity / Independence
    # -------------------------------------------------------------
    # Two-way Contingency Table (Rows: Treatment Groups, Cols: Patient Outcomes)
    # Row 1: Drug A, Row 2: Placebo
    contingency_table = np.array([
        [45, 35, 20],  # Drug A: [Remission, Stable, Progression]
        [25, 40, 35]   # Placebo: [Remission, Stable, Progression]
    ])

    chi2_stat, p_val, dof, expected_matrix = chi2_contingency(contingency_table)

    # Calculate Standardized Residuals: (O - E) / sqrt(E * (1 - row_prop) * (1 - col_prop))
    # AP Statistics uses Unstandardized/Standardized Residuals: (O - E) / sqrt(E)
    raw_residuals = (contingency_table - expected_matrix) / np.sqrt(expected_matrix)

    print(f"--- Two-Way Contingency Analysis ---")
    print(f"Chi2 Stat: {chi2_stat:.4f}")
    print(f"p-value:   {p_val:.4f}")
    print(f"Degrees of Freedom: {dof}")
    print("Expected Counts Matrix:\n", np.round(expected_matrix, 2))
    print("Standardized Residuals (O - E)/sqrt(E):\n", np.round(raw_residuals, 2))

if __name__ == "__main__":
    run_chi_square_tests()

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

On the AP Statistics exam, earning full credit (a "Score 5 performance") requires precise terminology, thorough condition checking, and contextual conclusions. Below are common student errors and how AP Readers evaluate them.

       [Student Performance Breakdown]
  Score 4 Student               Score 5 Student
  ┌──────────────────────────┐  ┌──────────────────────────┐
  │ • Correct calculation    │  │ • Explicit study-design  │
  │ • Checks observed counts │  │   identification         │
  │   instead of expected    │  │ • Shows expected table   │
  │ • Generic decision       │  │ • Explicitly states      │
  │   ("Reject H0")          │  │   E_ij >= 5 for ALL cells│
  │                          │  │ • Contextual conclusion  │
  │                          │  │   in terms of Ha         │
  └──────────────────────────┘  └──────────────────────────┘

3.1 Matrix of Score 4 vs. Score 5 Performance

Rubric Component Score 4 Student Response (Partial Credit) Score 5 Student Response (Complete Credit)
1. Hypothesis Formulation Writes $H_0: \text{Variables are independent}$ without defining context or variable names. States $H_0: \text{There is no association between patient dosage level}$ $\text{and recovery status in the population of clinical trial participants}$. Defines variables clearly.
2. Test Identification States "Chi-Square Test" without specifying which type (GoF, Homogeneity, or Independence). States "Chi-Square Test for Independence" and explicitly justifies choice based on single-sample two-variable design.
3. Condition Verification Writes "Expected counts $\ge 5$ holds" without showing expected values, or checks observed counts. Computes and displays the complete matrix of Expected Counts, then explicitly states: "All expected counts are $\ge 5$, as the minimum expected count is $8.42 \ge 5$."
4. Calculations Displays final $\chi^2$ and $p$-value from calculator without formula or degrees of freedom. Displays formula setup $\chi^2 = \sum \frac{(O-E)^2}{E}$, explicit degrees of freedom ($df = (r-1)(c-1) = 2$), test statistic $\chi^2 = 8.74$, and $p$-value $= 0.0126$.
5. Conclusion Writes "Reject $H_0$ because $p < 0.05$. Therefore, there is a relationship." Writes "Because $p = 0.0126 < \alpha = 0.05$, we reject $H_0$. There is statistically convincing evidence of an association between treatment dosage and patient recovery."

4. Stanford University Placement Pathway

4.1 Credit Mechanics & Course Exemption

A score of 5 on AP Statistics grants 5 units for STATS 60: Introduction to Statistical Methods.

                   AP Statistics Exam (Score 5)
                                │
                                ▼
                   Stanford STATS 60 Exemption
                       (5 Quarter Units)
                                │
        ┌───────────────────────┴───────────────────────┐
        ▼                                               ▼
   CS 109 Track                                   STATS 141 Track
(Probability for CS)                       (Biostatistics / Genomics)
  • Machine Learning                         • Medical Informatics
  • AI Models & Algorithms                   • Genomic Association Studies

4.2 Acceleration Tracks

Track A: CS 109 (Probability for Computer Scientists)

Track B: STATS 141 / BIOE 141 (Biostatistics & Medical Informatics)


5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Problem Statement

A clinical researcher at a Stanford-affiliated hospital investigates whether three alternative therapeutic regimes ($A$, $B$, and $C$) yield different distributions of side-effect severity in patients undergoing immunotherapy.

Independent random samples of patients were randomly assigned to one of the three regimes. After 12 weeks, side-effect severity was recorded as None/Mild, Moderate, or Severe. The observed counts are summarized below:

Therapeutic Regime None / Mild Moderate Severe Total
Regime A $40$ $30$ $10$ $80$
Regime B $50$ $25$ $5$ $80$
Regime C $30$ $35$ $15$ $80$
Total $120$ $90$ $30$ $240$

Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the distribution of side-effect severity differs among the three therapeutic regimes?


Exemplary Step-by-Step Solution (4-Step AP Format)

Step 1: State

We test the following hypotheses at $\alpha = 0.05$:


Step 2: Plan

Expected Counts Matrix: * Regime A: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$ * Regime B: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$ * Regime C: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$

Expected Counts None / Mild Moderate Severe
Regime A $40.0$ $30.0$ $10.0$
Regime B $40.0$ $30.0$ $10.0$
Regime C $40.0$ $30.0$ $10.0$

Verification: All expected cell counts are $\ge 5$ (minimum expected count is $10.0 \ge 5$). The condition is satisfied.


Step 3: Do

$$\chi^2 = \frac{(40 - 40)^2}{40} + \frac{(30 - 30)^2}{30} + \frac{(10 - 10)^2}{10}$$ $$+ \frac{(50 - 40)^2}{40} + \frac{(25 - 30)^2}{30} + \frac{(5 - 10)^2}{10}$$ $$+ \frac{(30 - 40)^2}{40} + \frac{(35 - 30)^2}{30} + \frac{(15 - 10)^2}{10}$$

$$\chi^2 = 0 + 0 + 0 + \frac{100}{40} + \frac{25}{30} + \frac{25}{10} + \frac{100}{40} + \frac{25}{30} + \frac{25}{10}$$ $$\chi^2 = 0 + 2.5 + 0.8333 + 2.5 + 2.5 + 0.8333 + 2.5 = 11.6667$$


Step 4: Conclude

Because the $p$-value ($0.0199$) is less than the significance level $\alpha = 0.05$, we reject $H_0$.

There is statistically convincing evidence that the distribution of side-effect severity differs among the three therapeutic regimes ($A$, $B$, and $C$).


Scoring Rubric Checklist for Exam Mastery

[AP Scoring Rubric Check]
✔ Hypotheses: Both H0 and Ha correctly stated in terms of distribution equality.
✔ Plan: Named Chi-Square Test for Homogeneity.
✔ Conditions: Explicitly calculated and displayed Expected Counts matrix; checked E_ij >= 5.
✔ Calculations: Correct df (4), Chi-Square statistic (11.67), and p-value (0.0199).
✔ Conclusion: Linked p-value to alpha, made explicit rejection decision, stated context.

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Stanford University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断