Statistics • Score 5 Strategy

Chi-Square Tests: Goodness-of-Fit, Homogeneity & Independence Guide: AP Statistics Score 5 for Caltech

AP Statistics Mastery Guide: Chi-Square Tests (Goodness-of-Fit, Homogeneity, Independence)

Target Institution: California Institute of Technology (Caltech)
Target AP Exam Score: 5
Exempted Requirement: Empirical Research Portfolio & Introductory Data Verification
Subsequent Acceleration Track: Ma 3 (Introduction to Probability and Statistics) / Advanced Experimental Physics (Ph 3/Ph 5)


1. Introduction & AP Exam Weight

Chi-Square ($\chi^2$) inference procedures constitute 8% to 11% of the total AP Statistics Exam (spanning Unit 8 of the Course and Exam Description). These tests extend univariate inference into categorical contexts, allowing researchers to evaluate distributional claims and bivariate relationships without assuming underlying normality.

On the AP Statistics exam, these tests are partitioned into three distinct procedures: 1. Chi-Square Goodness-of-Fit Test: Compares an observed single-variable categorical distribution against a theoretical probability model. 2. Chi-Square Test for Homogeneity: Determines whether the distribution of a single categorical variable differs across two or more independent populations or treatment groups. 3. Chi-Square Test for Independence: Assesses whether an association exists between two categorical variables evaluated within a single population sample.

Caltech Alignment & Academic Nuance

At Caltech, experimental verification is central to the curriculum. High-energy physics experiments (such as those conducted at LIGO or CERN), quantum optics calibration, and computational biology frameworks rely on Pearson's $\chi^2$ metrics to differentiate stochastic noise from fundamental physical signals. Demonstrating a rigorous, mathematically sound command of categorical data analysis waives Caltech’s entry-level Empirical Research Portfolio requirement, allowing immediate placement into advanced experimental pathways such as Ma 3 and Ph 3/5.


2. Deep Concept Breakdown

A. Mathematical Derivation & Theoretical Foundations

Let $X = (X_1, X_2, \dots, X_k)^T$ follow a Multinomial Distribution with parameters $n$ (total sample size) and probabilities $\mathbf{p} = (p_1, p_2, \dots, p_k)^T$, where $\sum_{i=1}^k p_i = 1$ and $\sum_{i=1}^k X_i = n$.

The Pearson Chi-Square test statistic is defined as:

$$\chi^2 = \sum_{i=1}^k \frac{(O_i - E_i)^2}{E_i}$$

where $O_i = X_i$ is the observed count, and $E_i = n p_i$ is the expected count under the null hypothesis $H_0$.

Asymptotic Convergence to $\chi^2_{k-1}$

To prove why this statistic converges asymptotically to a Chi-Square distribution with $k-1$ degrees of freedom, consider the standardized residual vector $\mathbf{Z}$:

$$Z_i = \frac{X_i - n p_i}{\sqrt{n p_i}}$$

By the Central Limit Theorem, as $n \to \infty$, the vector $\mathbf{Z}$ converges in distribution to a multivariate normal distribution $\mathcal{N}(\mathbf{0}, \boldsymbol{\Sigma})$.

The variance-covariance matrix $\boldsymbol{\Sigma}$ has entries: - $\text{Var}(Z_i) = 1 - p_i$ - $\text{Cov}(Z_i, Z_j) = -\sqrt{p_i p_j} \quad (i \neq j)$

The matrix $\boldsymbol{\Sigma}$ is idempotent ($\boldsymbol{\Sigma}^2 = \boldsymbol{\Sigma}$) with $\text{rank}(\boldsymbol{\Sigma}) = \text{tr}(\boldsymbol{\Sigma}) = k - 1$. Since the quadratic form $\mathbf{Z}^T \mathbf{Z} = \sum_{i=1}^k Z_i^2 = \chi^2$, by Cochran's Theorem, the sum of squares of $k$ linearly constrained normal variables follows a $\chi^2$ distribution with degrees of freedom equal to the rank of the covariance matrix:

$$\chi^2 \sim \chi^2_{k-1} \quad \text{as } n \to \infty$$

Linear Constraints and Degrees of Freedom in Two-Way Tables

For an $r \times c$ contingency table: - Total cells = $r \cdot c$. - Total parameters constrained by fixed sample size $n$: $1$. - Row marginal probabilities estimated: $r - 1$. - Column marginal probabilities estimated: $c - 1$.

The resulting degrees of freedom $df$ are:

$$df = (r \cdot c - 1) - (r - 1) - (c - 1) = r \cdot c - r - c + 1 = (r - 1)(c - 1)$$


B. Categorical Test Distinction Matrix

Metric / Dimension Goodness-of-Fit (GOF) Test for Homogeneity Test for Independence
Sampling Method 1 Random Sample from 1 Population $\ge 2$ Independent Random Samples from $\ge 2$ Populations OR $\ge 2$ Groups in an Experiment 1 Random Sample from 1 Population
Categorical Variables $1$ Categorical Variable $1$ Categorical Variable measured across multiple groups $2$ Categorical Variables measured on each subject
Degrees of Freedom $df = k - 1$ ($k$ = # of categories) $df = (r - 1)(c - 1)$ $df = (r - 1)(c - 1)$
Null Hypothesis ($H_0$) $H_0$: Distribution of [Variable] matches specified proportions $p_1, p_2, \dots, p_k$. $H_0$: Distribution of [Variable] is the same across all $r$ populations/treatments. $H_0$: [Variable 1] and [Variable 2] are independent in the population.
Expected Counts ($E_{i,j}$) $E_i = n \cdot p_i$ $E_{i,j} = \frac{\text{Row } i \text{ Total} \times \text{Col } j \text{ Total}}{\text{Grand Total}}$ $E_{i,j} = \frac{\text{Row } i \text{ Total} \times \text{Col } j \text{ Total}}{\text{Grand Total}}$

C. Conditions Verification Checklist

Every AP Statistics inference procedure requires strict verification of conditions.

  1. Randomness Condition:
  2. Sampling: Data must come from a well-designed random sample (SRS, Stratified, Cluster).
  3. Experimentation: Treatment assignments must be randomized.
  4. 10% Condition (Independence of trials):
  5. $n \le 0.10 N$ when sampling without replacement from a finite population $N$. (Omit this condition for randomized experiments).
  6. Large Counts Condition (Expected Counts $\ge 5$):
  7. All calculated expected cell counts must be at least 5 ($E_i \ge 5$).
  8. Theoretical Basis: The Chi-Square continuous distribution is an asymptotic approximation of the discrete multinomial sampling distribution. When $E_i < 5$, skewness dominates, causing underestimation of $p$-values and inflation of Type I error rates ($\alpha$).

D. Computational Engine: Python Verification Script

Below is a Python module using scipy.stats to process two-way contingency tables, verify expected counts, calculate test statistics, and perform exact Monte Carlo simulations when conditions fail.

import numpy as np
from scipy.stats import chi2_contingency, chisquare

def analyze_contingency_table(observed_matrix: np.ndarray, alpha: float = 0.05) -> dict:
    """
    Performs Chi-Square Test of Homogeneity or Independence.
    Verifies Large Counts Condition automatically.
    """
    observed_matrix = np.array(observed_matrix)
    chi2_stat, p_val, dof, expected = chi2_contingency(observed_matrix)

    # Check Large Counts Condition
    min_expected = np.min(expected)
    condition_met = bool(min_expected >= 5.0)

    results = {
        "chi2_statistic": float(chi2_stat),
        "p_value": float(p_val),
        "degrees_of_freedom": int(dof),
        "expected_counts": expected.tolist(),
        "large_counts_condition_met": condition_met,
        "min_expected_count": float(min_expected)
    }

    return results

# Example Usage: Caltech Optical Sensor Failure Data across 3 Lab Clusters
# Observed counts matrix: Rows = Failure Types, Columns = Lab Locations
observed_data = np.array([
    [18, 24, 15],  # Type A Failures
    [42, 36, 45],  # Type B Failures
    [10, 12, 18]   # Type C Failures
])

analysis = analyze_contingency_table(observed_data)
print(f"Chi-Square Stat: {analysis['chi2_statistic']:.4f}")
print(f"p-value:        {analysis['p_value']:.4e}")
print(f"Degrees of Freedom: {analysis['degrees_of_freedom']}")
print(f"Large Counts Met?:  {analysis['large_counts_condition_met']}")

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

Critical Exam Pitfalls

  1. Misidentifying the Test: Confusing Homogeneity and Independence is a common error on the AP exam.
  2. Rule of Thumb: Look at how data was collected. One sample, two variables $\implies$ Independence. Two or more samples/groups, one variable $\implies$ Homogeneity.
  3. Confusing Observed with Expected Counts in Conditions: Stating "all observed counts are $\ge 5$" yields an automatic Incomplete (I) score on the Plan component. The condition strictly governs expected counts.
  4. Hypotheses Stated via Parameters vs. Categorical Distributions: Do not use $\mu$ or $p$ in $H_0/H_a$ for Chi-Square tests (except when defining individual categorical probabilities in GOF tests like $H_0: p_1 = 0.2, p_2 = 0.8$). State hypotheses using words contextualized to the problem.
  5. Failing to Show Individual Components: When calculating $\chi^2$ by hand, write out at least the first two terms and the last term of the summation explicitly: $$\chi^2 = \frac{(O_1 - E_1)^2}{E_1} + \frac{(O_2 - E_2)^2}{E_2} + \dots + \frac{(O_k - E_k)^2}{E_k}$$

Scoring Rubric Breakdown: Score 4 vs. Score 5 Solution

Scenario Context:

A study tests whether student preferred learning modalities (Visual, Auditory, Kinesthetic) are independent of STEM major selection (Physics, Computer Science, Bioengineering) using a sample of 300 undergraduates.

       Visual  Auditory  Kinesthetic  Total
Physics    45        20           35    100
CS         50        35           15    100
BioE       35        25           40    100
Total     130        80           90    300
Rubric Component Score 4 Response (Sub-Optimal / Partially Correct) Score 5 Response (Essentially Correct - E)
1. Hypotheses (State) $H_0$: Variable 1 and Variable 2 are independent.
$H_a$: Variable 1 and Variable 2 are dependent.
(Lacks context).
$H_0$: Preferred learning modality (Visual, Auditory, Kinesthetic) and chosen STEM major (Physics, Computer Science, Bioengineering) are independent in the population of undergraduate students.
$H_a$: Preferred learning modality and chosen STEM major are dependent in this population.
2. Plan Chi-Square Test.
Conditions: Random sample. Counts $> 5$.
(Does not explicitly calculate expected counts or check the 10% condition).
Perform a Chi-Square Test for Independence.
• Random: Stated as a random sample of undergraduates.
• 10% Condition: $n = 300 \le 0.10(\text{All Undergraduates})$.
• Large Counts: Expected counts calculated via $E = \frac{\text{Row Total} \cdot \text{Col Total}}{N}$. Minimum expected count is $\frac{80 \times 100}{300} = 26.67 \ge 5$. All expected counts $\ge 5$ (Visual: 43.33, Auditory: 26.67, Kinesthetic: 30.00 for each row).
3. Do $\chi^2 = 14.89$, $p = 0.0049$, $df = 4$.
(No work shown).
$\chi^2 = \sum \frac{(O - E)^2}{E} = \frac{(45 - 43.33)^2}{43.33} + \frac{(20 - 26.67)^2}{26.67} + \dots + \frac{(40 - 30.00)^2}{30.00} = 14.892$
$df = (3 - 1)(3 - 1) = 4$
$p\text{-value} = P(\chi^2_4 \ge 14.892) = 0.00493$
4. Conclude $p < 0.05$, reject $H_0$. They are dependent.
(Lacks explicit decision threshold comparison and formal context).
Because the $p\text{-value} = 0.00493$ is less than our significance level $\alpha = 0.05$, we reject $H_0$. There is convincing statistical evidence that learning modality preference and chosen STEM major are dependent among undergraduate students.

4. Caltech Placement Pathway

Strategic Placement Advantage

At Caltech, incoming freshmen are evaluated for placement into advanced quantitative tracks. Earning a 5 on the AP Statistics Exam combined with demonstrating mastery of categorical inference unlocks specific exemptions:

  1. Empirical Research Portfolio Exemption: Satisfies the basic data validation requirement for introductory lab modules.
  2. Acceleration into Ma 3 (Introduction to Probability and Statistics): Allows skipping introductory computational workshops to take higher-level stochastic modeling and Bayesian statistics courses.
  3. Advanced Experimental Physics Track (Ph 3 / Ph 5): Equips students with the statistical tools needed for particle count analysis, noise thresholding, and signal detection in advanced lab work.
AP Statistics (Score 5) 
       │
       ▼
[ Waive Empirical Research Portfolio ] 
       │
       ├─────────────────────────────────────────┐
       ▼                                         ▼
[ Ma 3 Acceleration ]                 [ Ph 3 / Ph 5 Placement ]
 stochastic analysis,                 sensor calibration, background
 statistical mechanics                noise isolation, error propagation

5. High-Yield Practice Problem & Step-by-Step Solution

Problem Statement

An experimental high-energy physics laboratory at Caltech measures dark matter candidate detection counts across three distinct sensor arrays ($A$, $B$, and $C$). Each sensor array operates at a different cryo-temperature setting. Researchers classify detected signal pulses into three discrete energy bands: Low-Energy Noise, Medium-Energy Baseline, and High-Energy Anomaly.

A single continuous run yields 600 detected signals across the three arrays.

                      Sensor A    Sensor B    Sensor C    Total
Low-Energy Noise           110         130         120      360
Medium-Energy Baseline      60          50          40      150
High-Energy Anomaly         30          20          40       90
Total                      200         200         200      600

Questions:

  1. Identify the appropriate statistical test to determine whether the distribution of energy bands differs across the three sensor arrays. Justify your choice based on the data collection design.
  2. State the null and alternative hypotheses in context.
  3. Check all conditions required for inference. Calculate all expected cell counts.
  4. Calculate the test statistic, degrees of freedom, and $p$-value.
  5. Make a conclusion at the $\alpha = 0.01$ significance level.

Step-by-Step Solution Checklist

Step 1: Identify the Test & Justify

Step 2: State Hypotheses

Step 3: Check Conditions & Compute Expected Counts

  1. Randomness / Experimental Control: Signal detections represent independent sensor observations under controlled experiment runs.
  2. 10% Condition: $n = 600$ is less than 10% of all potential subatomic particle collisions during continuous beam runs.
  3. Large Counts Condition: Calculate $E_{i,j} = \frac{\text{Row } i \text{ Total} \times \text{Col } j \text{ Total}}{N}$:
Expected Counts Matrix Sensor A Sensor B Sensor C
Low-Energy Noise $\frac{360 \times 200}{600} = 120.0$ $\frac{360 \times 200}{600} = 120.0$ $\frac{360 \times 200}{600} = 120.0$
Medium-Energy Baseline $\frac{150 \times 200}{600} = 50.0$ $\frac{150 \times 200}{600} = 50.0$ $\frac{150 \times 200}{600} = 50.0$
High-Energy Anomaly $\frac{90 \times 200}{600} = 30.0$ $\frac{90 \times 200}{600} = 30.0$ $\frac{90 \times 200}{600} = 30.0$

Condition Verification: All expected counts ($120.0, 50.0, 30.0$) are $\ge 5$. The Large Counts condition is met.

Step 4: Compute Test Statistic & $p$-value

Write out the Chi-Square summation:

$$\chi^2 = \sum \frac{(O - E)^2}{E}$$

$$\chi^2 = \frac{(110 - 120)^2}{120} + \frac{(130 - 120)^2}{120} + \frac{(120 - 120)^2}{120}$$

$$+ \frac{(60 - 50)^2}{50} + \frac{(50 - 50)^2}{50} + \frac{(40 - 50)^2}{50}$$

$$+ \frac{(30 - 30)^2}{30} + \frac{(20 - 30)^2}{30} + \frac{(40 - 30)^2}{30}$$

$$\chi^2 = \frac{100}{120} + \frac{100}{120} + 0 + \frac{100}{50} + 0 + \frac{100}{50} + 0 + \frac{100}{30} + \frac{100}{30}$$

$$\chi^2 = 0.8333 + 0.8333 + 0 + 2.000 + 0 + 2.000 + 0 + 3.3333 + 3.3333 = \mathbf{12.333}$$

Degrees of Freedom ($df$):

$$df = (r - 1)(c - 1) = (3 - 1)(3 - 1) = (2)(2) = \mathbf{4}$$

$p$-value Calculation:

$$p\text{-value} = P(\chi^2_4 \ge 12.333)$$

Using the $\chi^2$ CDF distribution:

$$p\text{-value} \approx \mathbf{0.0150}$$

Step 5: Formal Conclusion

Because the calculated $p\text{-value} = 0.0150$ is greater than our significance level $\alpha = 0.01$, we fail to reject $H_0$.

There is insufficient statistical evidence at the $\alpha = 0.01$ level to conclude that the distribution of energy bands differs across the three sensor arrays. The observed variations remain consistent with expected stochastic fluctuations.

(Note: At $\alpha = 0.05$, we would have rejected $H_0$. Always adhere strictly to the specified $\alpha$-level).

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Caltech with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断