Statistics • Score 5 Strategy

Chi-Square Tests: Goodness-of-Fit, Homogeneity & Independence Guide: AP Statistics Score 5 for MIT

AP Statistics Mastery Guide: Chi-Square Inference

Target Exam Score: 5
Target Institution: Massachusetts Institute of Technology (MIT)
Academic Pathway: Quantitative Rigor Portfolio Satisfier $\rightarrow$ Acceleration into 18.650 (Fundamentals of Statistics)


1. Introduction & AP Exam Weight

Chi-Square ($\chi^2$) inference procedures constitute 7% to 11% of the AP Statistics Exam (spanning Units 8 and 9). On the exam, these concepts appear in both Section I (Multiple-Choice) and Section II (Free-Response Questions).

While standard high school curricula treat Chi-Square procedures as simple calculator routines, top-tier performance—and placement into MIT’s advanced quantitative tracks—requires an understanding of categorical data analysis. At MIT, categorical inference, contingency table analysis, and goodness-of-fit testing are foundational for empirical research in machine learning, computational biology (e.g., genomic sequence alignment), and microeconomics. Mastering this topic demonstrates the mathematical rigor needed to pass the Quantitative Rigor Portfolio review and enter 18.650 (Fundamentals of Statistics).

                      ┌─────────────────────────────────────────┐
                      │    Categorical Inference Framework     │
                      └────────────────────┬────────────────────┘
                                           │
        ┌──────────────────────────────────┼──────────────────────────────────┐
        ▼                                  ▼                                  ▼
┌──────────────┐                   ┌──────────────┐                   ┌──────────────┐
│  Goodness-   │                   │ Test for     │                   │ Test for     │
│  of-Fit      │                   │ Homogeneity  │                   │ Independence │
└───────┬──────┘                   └───────┬──────┘                   └───────┬──────┘
        │                                  │                                  │
  1 Population                       2+ Populations                     1 Population
  1 Categorical Variable             1 Categorical Variable             2 Categorical Variables
  $df = k - 1$                       $df = (r-1)(c-1)$                  $df = (r-1)(c-1)$

2. Deep Concept Breakdown

Mathematical Foundations of the Chi-Square Distribution

Let $X_1, X_2, \dots, X_k$ be independent standard normal random variables, where $X_i \sim N(0, 1)$. The sum of their squares follows a Chi-Square distribution with $k$ degrees of freedom:

$$V = \sum_{i=1}^{k} X_i^2 \sim \chi^2_k$$

The probability density function (PDF) of a $\chi^2$ distribution with $k$ degrees of freedom is:

$$f(x; k) = \frac{x^{(k/2) - 1} e^{-x/2}}{2^{k/2} \Gamma\left(\frac{k}{2}\right)}, \quad x \ge 0$$

where $\Gamma(z) = \int_0^\infty t^{z-1} e^{-t} dt$ is the Gamma function.

Derivation of the Pearson Chi-Square Test Statistic

For categorical counts $O_i$ (Observed) under a multinomial model with expected values $E_i = n p_i$, the sample counts converge asymptotically to a multivariate normal distribution via the Central Limit Theorem. Standardizing these counts yields:

$$Z_i = \frac{O_i - E_i}{\sqrt{E_i(1 - p_i)}}$$

Under the null hypothesis, summing the standardized squared deviations yields Pearson's test statistic:

$$\chi^2 = \sum_{i=1}^{k} \frac{(O_i - E_i)^2}{E_i}$$

As $n \to \infty$, this statistic converges in distribution to $\chi^2_{df}$.


Comparative Architecture of Chi-Square Tests

Dimension Goodness-of-Fit (GOF) Test for Homogeneity Test for Independence
Number of Samples/Populations 1 Sample from 1 Population $2+$ Independent Samples from $2+$ Populations/Treatments 1 Sample from 1 Population
Variables Evaluated 1 Categorical Variable (compared against a hypothesized distribution) 1 Categorical Variable measured across different groups 2 Categorical Variables measured simultaneously on each subject
Degrees of Freedom ($df$) $df = k - 1$
(where $k$ = number of categories)
$df = (r - 1)(c - 1)$
(where $r$ = rows, $c$ = columns)
$df = (r - 1)(c - 1)$
(where $r$ = rows, $c$ = columns)
Expected Cell Count ($E$) $E_i = n \cdot p_i$ $E_{i,j} = \frac{\text{Row Total}_i \times \text{Column Total}_j}{\text{Grand Total}}$ $E_{i,j} = \frac{\text{Row Total}_i \times \text{Column Total}_j}{\text{Grand Total}}$
Primary Research Question Does the categorical distribution match a theoretical model? Do two or more populations have identical categorical proportions? Is there an association between two categorical variables in a single population?

Verification of Technical Conditions

To earn full credit on the AP exam, conditions must be explicitly verified, not merely stated:

  1. Random Sampling / Assignment: Data must originate from a simple random sample (SRS) or a randomized experiment.
  2. 10% Condition (Independence): When sampling without replacement, the sample size must satisfy $n \le 0.10 N$ to preserve near-independence of Bernoulli trials. (Note: This condition is not required for randomized experiments).
  3. Large Counts Condition (Expected Counts): All expected cell counts must satisfy $E \ge 5$.
  4. Critical Distinction: This condition applies to expected counts ($E_i$), never to observed counts ($O_i$).

Python Computation & Analysis Framework

The script below performs a complete Chi-Square analysis, including model identification, expected count generation, test-statistic computation, and standardized residual evaluation:

import numpy as np
from scipy import stats

def analyze_contingency_table(observed_matrix, row_labels, col_labels):
    """
    Performs full Chi-Square Test of Independence/Homogeneity with 
    Expected Counts and Standardized Residuals analysis.
    """
    obs = np.array(observed_matrix)
    chi2_stat, p_val, dof, expected = stats.chi2_contingency(obs)

    print("=" * 60)
    print(f"Chi-Square Test Statistic (χ²): {chi2_stat:.4f}")
    print(f"Degrees of Freedom (df):       {dof}")
    print(f"p-value:                        {p_val:.4e}")
    print("=" * 60)

    # Check Large Counts Condition
    all_valid = np.all(expected >= 5)
    print(f"Condition Check - All Expected Counts >= 5: {all_valid}")
    if not all_valid:
        print(" WARNING: Large Counts Condition violated. Chi-Square approximation may fail.")

    print("\nExpected Cell Counts Matrix:")
    print(np.round(expected, 2))

    # Calculate Standardized Residuals: (O - E) / sqrt(E * (1 - row_prop) * (1 - col_prop))
    # Adjusted Pearson Residuals:
    row_sums = obs.sum(axis=1, keepdims=True)
    col_sums = obs.sum(axis=0, keepdims=True)
    total = obs.sum()

    margin_adjust = (1 - row_sums / total) * (1 - col_sums / total)
    std_residuals = (obs - expected) / np.sqrt(expected * margin_adjust)

    print("\nAdjusted Standardized Residuals:")
    print(np.round(std_residuals, 3))

    return {
        "chi2": chi2_stat,
        "p_value": p_val,
        "df": dof,
        "expected": expected,
        "std_residuals": std_residuals
    }

# Example: AI Failure Modes across 3 Model Architectures
# Rows: Transformer, CNN, Mamba
# Cols: Out-of-Distribution, Adversarial, Latent Collapse
data = [
    [120, 85, 45],  # Transformer
    [90, 110, 50],  # CNN
    [60, 45, 95]    # Mamba
]
r_names = ["Transformer", "CNN", "Mamba"]
c_names = ["OOD", "Adversarial", "Latent Collapse"]

analyze_contingency_table(data, r_names, c_names)

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

Score 4 vs. Score 5 Performance Matrix

┌────────────────────────────────────────────────────────────────────────────────────────┐
│                                   RESPONSE COMPARISON                                  │
├──────────────────────────────┬──────────────────────────────┬──────────────────────────┤
│ Section                      │ Score 4 Response (Competent) │ Score 5 Response (MIT)   │
├──────────────────────────────┼──────────────────────────────┼──────────────────────────┤
│ Hypotheses Identification    │ Uses vague, uncontextualized │ Defines non-association  │
│                              │ parameter statements:        │ or structural identity   │
│                              │ "H0: Variables are           │ within the population    │
│                              │ independent."                │ using clear context.     │
├──────────────────────────────┼──────────────────────────────┼──────────────────────────┤
│ Checking Conditions          │ Writes "Expected counts > 5  │ Explicitly displays the  │
│                              │ check out" without listing   │ matrix of all calculated │
│                              │ calculated values.           │ expected values >= 5.    │
├──────────────────────────────┼──────────────────────────────┼──────────────────────────┤
│ Mechanics & Communication    │ Reports χ², df, p-value from │ Shows formula set up,    │
│                              │ calculator without setup.    │ inputs, $df$, and exact  │
│                              │                              │ tail-area expression.    │
├──────────────────────────────┼──────────────────────────────┼──────────────────────────┤
│ Decision & Context           │ "Reject H0. Variables are    │ Links $p$-value to $\alpha$,│
│                              │ related."                    │ frames decision without  │
│                              │                              │ "accepting $H_0$", and   │
│                              │                              │ evaluates residuals.     │
└──────────────────────────────┴──────────────────────────────┴──────────────────────────┘

Detailed Exam Pitfalls

Pitfall 1: Confusing Test for Homogeneity with Test for Independence

Pitfall 2: Confusing Observed and Expected Counts in Conditions

Pitfall 3: Failing to Interpret Standardized Residuals

$$R = \frac{O - E}{\sqrt{E}}$$


4. MIT Placement Pathway

Satisfying the Quantitative Rigor Portfolio

At MIT, incoming students with AP Statistics scores of 5 are evaluated for advanced quantitative placement. Demonstrating mastery of categorical inference helps satisfy the Quantitative Rigor Portfolio, allowing students to bypass introductory statistics requirements.

┌─────────────────────────────────────────┐
│     AP Statistics (Score 5 Target)      │
└────────────────────┬────────────────────┘
                     │
                     ▼
┌─────────────────────────────────────────┐
│     Quantitative Rigor Portfolio        │
│  (Validation of Empirical Data Analysis)│
└────────────────────┬────────────────────┘
                     │
                     ▼
┌─────────────────────────────────────────┐
│     Advanced Placement Granted:         │
│     Bypass Introductory Statistics      │
└────────────────────┬────────────────────┘
                     │
                     ▼
┌─────────────────────────────────────────┐
│      Accelerate Directly Into:          │
│  18.650: Fundamentals of Statistics     │
└─────────────────────────────────────────┘

Direct Acceleration: Course 18.650 (Fundamentals of Statistics)

18.650 is MIT’s rigorous, calculus-based introduction to mathematical statistics. Mastering Chi-Square inference prepares you for advanced topics in 18.650, such as: 1. Likelihood Ratio Tests (LRT): Proving that Pearson's $\chi^2$ test is an asymptotic equivalent of the LRT via Wilks' Theorem:

$$-2 \ln \Lambda = 2 \sum_{i} O_i \ln\left(\frac{O_i}{E_i}\right) \xrightarrow{d} \chi^2_{df}$$

  1. Generalized Linear Models (GLMs): Extending categorical analysis to logistic regression and Poisson log-linear models.
  2. High-Dimensional Categorical Analysis: Managing large sparse matrices in machine learning applications where standard Chi-Square assumptions break down (requiring Fisher's Exact Test or permutation tests).

5. High-Yield Practice Problem & Step-by-Step Solution

Scenario

An MIT CSAIL research team is evaluating three different large language model (LLM) alignment strategies: RLHF (Reinforcement Learning from Human Feedback), DPO (Direct Preference Optimization), and KTO (Kahneman-Tversky Optimization).

The researchers randomly sample $n = 600$ model outputs evaluated under high-temperature sampling and classify each output into one of three response categories: Fully Compliant, Hallucination, or Refusal.

The observed counts are recorded in the contingency table below:

Alignment Strategy Fully Compliant Hallucination Refusal Total
RLHF 135 35 30 200
DPO 150 20 30 200
KTO 115 45 40 200
Total 400 100 100 600

Questions

  1. Identify the appropriate inference procedure and state the null and alternative hypotheses in context.
  2. Verify all necessary theoretical conditions for performing this test.
  3. Calculate the matrix of expected counts, the test statistic ($\chi^2$), degrees of freedom ($df$), and the corresponding $p$-value. Show all work.
  4. Based on your $p$-value, render a conclusion using $\alpha = 0.01$.
  5. Calculate the standardized residuals for the DPO / Hallucination and KTO / Hallucination cells. Interpret these values to explain which alignment strategy deviates most significantly from expected behavior.

Comprehensive Solution & Grading Rubric

Step 1: Identify Test & Hypotheses

Step 2: Verification of Conditions

  1. Randomization: Models/outputs were randomly sampled for evaluation across alignment strategies.
  2. Independence: Outputs are independently generated. Sampling without replacement is not applicable here if generation is an infinite process, or $N > 6000$ generations easily satisfies $n \le 0.10 N$.
  3. Large Expected Counts: Expected counts are calculated using $E_{i,j} = \frac{\text{Row Total}_i \times \text{Column Total}_j}{\text{Grand Total}}$:

$$\text{For all rows, Row Total} = 200$$

Matrix of Expected Counts ($E_{i,j}$):

Strategy Fully Compliant Hallucination Refusal
RLHF 133.33 33.33 33.33
DPO 133.33 33.33 33.33
KTO 133.33 33.33 33.33

Condition Met: All expected cell counts are $133.33 \text{ or } 33.33$, which are all $\ge 5$.


Step 3: Mechanics & Calculations

$$df = (r - 1)(c - 1) = (3 - 1)(3 - 1) = 2 \times 2 = 4$$

Calculate Pearson's $\chi^2$ Statistic:

$$\chi^2 = \sum \frac{(O_{i,j} - E_{i,j})^2}{E_{i,j}}$$

$$\chi^2 = \frac{(135 - 133.33)^2}{133.33} + \frac{(35 - 33.33)^2}{33.33} + \frac{(30 - 33.33)^2}{33.33}$$

$$+ \frac{(150 - 133.33)^2}{133.33} + \frac{(20 - 33.33)^2}{33.33} + \frac{(30 - 33.33)^2}{33.33}$$

$$+ \frac{(115 - 133.33)^2}{133.33} + \frac{(45 - 33.33)^2}{33.33} + \frac{(40 - 33.33)^2}{33.33}$$

Evaluating individual term values: * $\text{RLHF}: 0.021 + 0.084 + 0.333 = 0.438$ * $\text{DPO}: 2.083 + 5.330 + 0.333 = 7.746$ * $\text{KTO}: 2.520 + 4.085 + 1.333 = 7.938$

$$\chi^2 = 0.438 + 7.746 + 7.938 = 16.122$$

$p$-value calculation:

$$p\text{-value} = P(\chi^2_4 \ge 16.122)$$

Using the continuous $\chi^2$ integral with $df=4$:

$$p\text{-value} \approx 0.00286$$


Step 4: Decision and Justification


Step 5: Standardized Residual Analysis

Standardized Residual formula:

$$R_{i,j} = \frac{O_{i,j} - E_{i,j}}{\sqrt{E_{i,j}}}$$

  1. DPO / Hallucination: $$R_{\text{DPO, Hall}} = \frac{20 - 33.33}{\sqrt{33.33}} = \frac{-13.33}{5.773} \approx -2.31$$

  2. KTO / Hallucination: $$R_{\text{KTO, Hall}} = \frac{45 - 33.33}{\sqrt{33.33}} = \frac{11.67}{5.773} \approx +2.02$$

Interpretation: * DPO produces significantly fewer hallucinations than expected under the null hypothesis of homogeneity ($R = -2.31$), falling more than 2 standard deviations below expectation. * KTO produces significantly more hallucinations than expected ($R = +2.02$), falling more than 2 standard deviations above expectation. * These two cells are the primary contributors to the overall test statistic ($\chi^2 = 16.122$). DPO demonstrates a clear advantage over KTO in reducing output hallucinations.


6. Target Score 5 Checklist for the AP Exam

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like MIT with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断