AP Statistics Exam Mastery Guide: Chi-Square Tests
Target Topic: Inference for Categorical Data (Goodness-of-Fit, Homogeneity, and Independence)
Target Institution: Stanford University (Aiming for Score 5)
1. Introduction & AP Exam Weight
Categorical data analysis forms a cornerstone of statistical inference. On the AP Statistics Exam, Unit 8: Inference for Categorical Data: Chi-Square accounts for 7% to 11% of the multiple-choice section and routinely features in Free-Response Questions (FRQs)—either as a standalone 11-minute problem or integrated into Question 6 (the Investigation Task).
Mastery of Chi-Square ($\chi^2$) inferential procedures requires more than memorizing algorithms; it demands a precise understanding of study design, sampling mechanics, and probability distributions.
At Stanford University, achieving a score of 5 on AP Statistics awards 5 quarter-units of credit, waiving STATS 60: Introduction to Statistical Methods. This direct exemption allows prospective Computer Science, Data Science, Bioengineering, and Pre-Med students to accelerate into advanced coursework such as CS 109 (Probability for Computer Scientists) or STATS 141 (Biostatistics).
2. Deep Concept Breakdown
2.1 Mathematical Foundations of the Chi-Square Distribution
The Chi-Square distribution with $k$ degrees of freedom is defined as the distribution of a sum of the squares of $k$ independent standard normal random variables. If $Z_1, Z_2, \dots, Z_k \sim \text{N}(0, 1)$ independently, then:
$$Q = \sum_{i=1}^{k} Z_i^2 \sim \chi^2_k$$
The Probability Density Function (PDF) of a $\chi^2$ random variable with $k$ degrees of freedom is given by:
$$f(x; k) = \frac{x^{(k/2) - 1} e^{-x/2}}{2^{k/2} \Gamma\left(\frac{k}{2}\right)}, \quad x \ge 0$$
where $\Gamma(n)$ is the Gamma function. The mean of $\chi^2_k$ is $k$, and its variance is $2k$. As $k \to \infty$, by the Central Limit Theorem, the $\chi^2_k$ distribution approaches a Normal distribution $\text{N}(k, \sqrt{2k})$.
2.2 Derivation of the Pearson $\chi^2$ Test Statistic
For a sample divided into $k$ mutually exclusive categorical bins with observed counts $O_1, O_2, \dots, O_k$ and expected counts $E_1, E_2, \dots, E_k$ under a null model $H_0$, the sample count vector follows a Multinomial Distribution.
When the total sample size $n = \sum O_i$ is large, each observed count $O_i$ is approximately normally distributed:
$$O_i \sim \text{N}\left(E_i, \sqrt{E_i(1 - p_i)}\right)$$
Standardizing the residual yields $Z_i = \frac{O_i - E_i}{\sqrt{E_i}}$. Squaring and summing these standardized residuals over all $k$ categories yields the classic Pearson $\chi^2$ test statistic:
$$\chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i}$$
Because the constraint $\sum_{i=1}^{k} O_i = n$ removes one degree of freedom, the statistic $\chi^2$ asymptotically follows a $\chi^2$ distribution with $df = k - 1$.
2.3 Comparative Analysis of the Three Chi-Square Tests
The AP Statistics exam evaluates three distinct $\chi^2$ tests. Choosing the correct test depends on the sampling strategy and the number of variables/populations.
| Structural Metric | $\chi^2$ Goodness-of-Fit Test | $\chi^2$ Test for Homogeneity | $\chi^2$ Test for Independence |
|---|---|---|---|
| Number of Categorical Variables | $1$ variable | $1$ variable | $2$ variables |
| Number of Samples / Populations | $1$ sample vs. theoretical model | $2+$ independent samples/groups | $1$ single sample |
| Null Hypothesis ($H_0$) Formulation | $H_0$: The population distribution matches specified probabilities $p_1, p_2, \dots, p_k$. | $H_0$: The distribution of [Variable] is the same across all populations. | $H_0$: Variable A and Variable B are independent in the population. |
| Degrees of Freedom ($df$) | $df = k - 1$ | $df = (r - 1)(c - 1)$ | $df = (r - 1)(c - 1)$ |
| Expected Cell Count Formula | $E_i = n \cdot p_i$ | $E_{i,j} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Table Total}}$ | $E_{i,j} = \frac{(\text{Row } i \text{ Total}) \times (\text{Column } j \text{ Total})}{\text{Table Total}}$ |
2.4 Necessary Conditions for Statistical Inference
Before executing any $\chi^2$ test, all three of the following conditions must be explicitly verified:
- Random Condition: Data must come from a well-designed random sample (SRS, stratified, cluster) or a randomized experiment.
- 10% Condition (Independence): When sampling without replacement from a finite population, the total sample size must satisfy $n \le 0.10 N$. (Note: This condition does NOT apply to randomized experiments).
- Large Counts Condition: All expected counts must be at least 5 ($E_i \ge 5$ or $E_{i,j} \ge 5$).
- Critical Pitfall: This condition applies to expected cell counts, not observed counts.
2.5 Python Implementation for Computation and Verification
import numpy as np
from scipy.stats import chi2_contingency, chisquare
def run_chi_square_tests():
# -------------------------------------------------------------
# 1. Chi-Square Goodness-of-Fit Test
# -------------------------------------------------------------
# Observed counts of genomic mutations across 4 loci
observed_gof = np.array([32, 28, 15, 25])
# Expected proportions under Mendelian inheritance (3:3:2:2 ratio -> 0.3, 0.3, 0.2, 0.2)
expected_props = np.array([0.30, 0.30, 0.20, 0.20])
total_n = np.sum(observed_gof)
expected_gof = total_n * expected_props
gof_stat, gof_p_val = chisquare(f_obs=observed_gof, f_exp=expected_gof)
print(f"--- Goodness-of-Fit Test ---")
print(f"Chi2 Stat: {gof_stat:.4f} | p-value: {gof_p_val:.4f} | df: {len(observed_gof)-1}\n")
# -------------------------------------------------------------
# 2. Chi-Square Test for Homogeneity / Independence
# -------------------------------------------------------------
# Two-way Contingency Table (Rows: Treatment Groups, Cols: Patient Outcomes)
# Row 1: Drug A, Row 2: Placebo
contingency_table = np.array([
[45, 35, 20], # Drug A: [Remission, Stable, Progression]
[25, 40, 35] # Placebo: [Remission, Stable, Progression]
])
chi2_stat, p_val, dof, expected_matrix = chi2_contingency(contingency_table)
# Calculate Standardized Residuals: (O - E) / sqrt(E * (1 - row_prop) * (1 - col_prop))
# AP Statistics uses Unstandardized/Standardized Residuals: (O - E) / sqrt(E)
raw_residuals = (contingency_table - expected_matrix) / np.sqrt(expected_matrix)
print(f"--- Two-Way Contingency Analysis ---")
print(f"Chi2 Stat: {chi2_stat:.4f}")
print(f"p-value: {p_val:.4f}")
print(f"Degrees of Freedom: {dof}")
print("Expected Counts Matrix:\n", np.round(expected_matrix, 2))
print("Standardized Residuals (O - E)/sqrt(E):\n", np.round(raw_residuals, 2))
if __name__ == "__main__":
run_chi_square_tests()
3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances
On the AP Statistics exam, earning full credit (a "Score 5 performance") requires precise terminology, thorough condition checking, and contextual conclusions. Below are common student errors and how AP Readers evaluate them.
[Student Performance Breakdown]
Score 4 Student Score 5 Student
┌──────────────────────────┐ ┌──────────────────────────┐
│ • Correct calculation │ │ • Explicit study-design │
│ • Checks observed counts │ │ identification │
│ instead of expected │ │ • Shows expected table │
│ • Generic decision │ │ • Explicitly states │
│ ("Reject H0") │ │ E_ij >= 5 for ALL cells│
│ │ │ • Contextual conclusion │
│ │ │ in terms of Ha │
└──────────────────────────┘ └──────────────────────────┘
3.1 Matrix of Score 4 vs. Score 5 Performance
| Rubric Component | Score 4 Student Response (Partial Credit) | Score 5 Student Response (Complete Credit) |
|---|---|---|
| 1. Hypothesis Formulation | Writes $H_0: \text{Variables are independent}$ without defining context or variable names. | States $H_0: \text{There is no association between patient dosage level}$ $\text{and recovery status in the population of clinical trial participants}$. Defines variables clearly. |
| 2. Test Identification | States "Chi-Square Test" without specifying which type (GoF, Homogeneity, or Independence). | States "Chi-Square Test for Independence" and explicitly justifies choice based on single-sample two-variable design. |
| 3. Condition Verification | Writes "Expected counts $\ge 5$ holds" without showing expected values, or checks observed counts. | Computes and displays the complete matrix of Expected Counts, then explicitly states: "All expected counts are $\ge 5$, as the minimum expected count is $8.42 \ge 5$." |
| 4. Calculations | Displays final $\chi^2$ and $p$-value from calculator without formula or degrees of freedom. | Displays formula setup $\chi^2 = \sum \frac{(O-E)^2}{E}$, explicit degrees of freedom ($df = (r-1)(c-1) = 2$), test statistic $\chi^2 = 8.74$, and $p$-value $= 0.0126$. |
| 5. Conclusion | Writes "Reject $H_0$ because $p < 0.05$. Therefore, there is a relationship." | Writes "Because $p = 0.0126 < \alpha = 0.05$, we reject $H_0$. There is statistically convincing evidence of an association between treatment dosage and patient recovery." |
4. Stanford University Placement Pathway
4.1 Credit Mechanics & Course Exemption
A score of 5 on AP Statistics grants 5 units for STATS 60: Introduction to Statistical Methods.
AP Statistics Exam (Score 5)
│
▼
Stanford STATS 60 Exemption
(5 Quarter Units)
│
┌───────────────────────┴───────────────────────┐
▼ ▼
CS 109 Track STATS 141 Track
(Probability for CS) (Biostatistics / Genomics)
• Machine Learning • Medical Informatics
• AI Models & Algorithms • Genomic Association Studies
4.2 Acceleration Tracks
Track A: CS 109 (Probability for Computer Scientists)
- Focus: Algorithmic probability, Naive Bayes models, stochastic processes, maximum likelihood estimation (MLE), cross-entropy metrics.
- Chi-Square Connection: categorical distributions, feature independence assertions in vector models, discrete goodness-of-fit testing for probabilistic graphical models.
Track B: STATS 141 / BIOE 141 (Biostatistics & Medical Informatics)
- Focus: Contingency tables, survival analysis, log-linear models, logistic regression, bioengineering trial design.
- Chi-Square Connection: Genome-wide Association Studies (GWAS) evaluate single nucleotide polymorphism (SNP) frequency shifts across disease phenotypes using $2 \times 2$ chi-square independence frameworks.
5. High-Yield Practice Problem & Step-by-Step Solution Checklist
Problem Statement
A clinical researcher at a Stanford-affiliated hospital investigates whether three alternative therapeutic regimes ($A$, $B$, and $C$) yield different distributions of side-effect severity in patients undergoing immunotherapy.
Independent random samples of patients were randomly assigned to one of the three regimes. After 12 weeks, side-effect severity was recorded as None/Mild, Moderate, or Severe. The observed counts are summarized below:
| Therapeutic Regime | None / Mild | Moderate | Severe | Total |
|---|---|---|---|---|
| Regime A | $40$ | $30$ | $10$ | $80$ |
| Regime B | $50$ | $25$ | $5$ | $80$ |
| Regime C | $30$ | $35$ | $15$ | $80$ |
| Total | $120$ | $90$ | $30$ | $240$ |
Do these data provide convincing statistical evidence at the $\alpha = 0.05$ significance level that the distribution of side-effect severity differs among the three therapeutic regimes?
Exemplary Step-by-Step Solution (4-Step AP Format)
Step 1: State
We test the following hypotheses at $\alpha = 0.05$:
- $H_0$: The distribution of side-effect severity (None/Mild, Moderate, Severe) is the same across the three therapeutic regimes ($A$, $B$, and $C$).
- $H_a$: The distribution of side-effect severity differs for at least one of the therapeutic regimes.
Step 2: Plan
- Test Name: $\chi^2$ Test for Homogeneity.
-
Justification: We have three independent sample groups (Therapeutic Regimes $A$, $B$, $C$) and one categorical response variable (Side-Effect Severity).
-
Check Conditions:
- Random: Patients were randomly assigned to the three therapeutic regimes.
- 10% Condition: Not applicable because this is a randomized experiment rather than sampling without replacement from a finite population.
- Large Counts Condition: Calculate Expected Counts using: $$E_{i,j} = \frac{(\text{Row Total}) \times (\text{Column Total})}{\text{Table Total}}$$
Expected Counts Matrix: * Regime A: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$ * Regime B: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$ * Regime C: $$\text{None/Mild: } \frac{80 \times 120}{240} = 40.0, \quad \text{Moderate: } \frac{80 \times 90}{240} = 30.0, \quad \text{Severe: } \frac{80 \times 30}{240} = 10.0$$
| Expected Counts | None / Mild | Moderate | Severe |
|---|---|---|---|
| Regime A | $40.0$ | $30.0$ | $10.0$ |
| Regime B | $40.0$ | $30.0$ | $10.0$ |
| Regime C | $40.0$ | $30.0$ | $10.0$ |
Verification: All expected cell counts are $\ge 5$ (minimum expected count is $10.0 \ge 5$). The condition is satisfied.
Step 3: Do
-
Degrees of Freedom: $$df = (r - 1)(c - 1) = (3 - 1)(3 - 1) = 2 \times 2 = 4$$
-
Test Statistic Calculation: $$\chi^2 = \sum \frac{(O - E)^2}{E}$$
$$\chi^2 = \frac{(40 - 40)^2}{40} + \frac{(30 - 30)^2}{30} + \frac{(10 - 10)^2}{10}$$ $$+ \frac{(50 - 40)^2}{40} + \frac{(25 - 30)^2}{30} + \frac{(5 - 10)^2}{10}$$ $$+ \frac{(30 - 40)^2}{40} + \frac{(35 - 30)^2}{30} + \frac{(15 - 10)^2}{10}$$
$$\chi^2 = 0 + 0 + 0 + \frac{100}{40} + \frac{25}{30} + \frac{25}{10} + \frac{100}{40} + \frac{25}{30} + \frac{25}{10}$$ $$\chi^2 = 0 + 2.5 + 0.8333 + 2.5 + 2.5 + 0.8333 + 2.5 = 11.6667$$
- $p$-value: Using $df = 4$ and $\chi^2 = 11.6667$: $$p\text{-value} = P(\chi^2_4 \ge 11.6667) \approx 0.0199$$
Step 4: Conclude
Because the $p$-value ($0.0199$) is less than the significance level $\alpha = 0.05$, we reject $H_0$.
There is statistically convincing evidence that the distribution of side-effect severity differs among the three therapeutic regimes ($A$, $B$, and $C$).
Scoring Rubric Checklist for Exam Mastery
[AP Scoring Rubric Check]
✔ Hypotheses: Both H0 and Ha correctly stated in terms of distribution equality.
✔ Plan: Named Chi-Square Test for Homogeneity.
✔ Conditions: Explicitly calculated and displayed Expected Counts matrix; checked E_ij >= 5.
✔ Calculations: Correct df (4), Chi-Square statistic (11.67), and p-value (0.0199).
✔ Conclusion: Linked p-value to alpha, made explicit rejection decision, stated context.