Statistics • Score 5 Strategy

Binomial vs. Geometric Probability Models & Expected Values Guide: AP Statistics Score 5 for Stanford University

AP Statistics Exam Mastery Guide: Discrete Probability Models & Expected Value

Target Exam: AP Statistics (Score 5 Focus)
Target Institution: Stanford University
Equivalent Course & Credit: STATS 60: Introduction to Statistical Methods (5 Quarter Units)
Accelerated Sequence: CS 109 (Probability for Computer Scientists) / CME 106


1. Introduction & AP Exam Weight

Discrete probability models—specifically the Binomial and Geometric distributions—form the mathematical bridge between basic probability rules and advanced statistical inference. Within the AP Statistics Course and Exam Description (CED), these topics reside in Unit 4: Random Variables, Sets, and Probability Distributions (covering Topic 4.10 through Topic 4.12).

While Unit 4 directly constitutes 10–20% of the AP Exam multiple-choice section, discrete models are foundational to performance across the entire exam: * Binomial models underpin the sampling distribution of a sample proportion ($\hat{p}$), setting up $z$-procedures, $1$-prop $z$-tests, and $2$-prop $z$-intervals. * Geometric models test discrete infinite sequences, expectation, and conditional probability bounds. * Expected value ($E[X]$) calculations appear regularly in Free Response Questions (FRQs) requiring decision analysis and expected profit/loss evaluations.

At elite institutions like Stanford University, a Score of 5 waives STATS 60 and grants 5 quarter units toward graduation. More importantly, it unlocks early enrollment into CS 109: Probability for Computer Scientists. CS 109 assumes an immediate, rigorous intuition for indicator random variables, expectations, variance, and memoryless properties—concepts derived directly from mastering Binomial and Geometric distributions.


2. Deep Concept Breakdown

Model Disambiguation: BINS vs. BITS

To secure full credit on AP FRQs, you must explicitly state conditions in context before applying formulas.

Characteristic Binomial Model ($X \sim \text{Binom}(n, p)$) Geometric Model ($Y \sim \text{Geom}(p)$)
Definition Counts number of successes in $n$ trials. Counts number of trials until the first success.
Acronym BINS BITS
Binary Success / Failure outcomes. Success / Failure outcomes.
Independance Trials are independent (check $10\%$ condition if sampling without replacement). Trials are independent (check $10\%$ condition if sampling without replacement).
N / T Number of trials fixed ($n$). Trials variable (counting until 1st success).
Success $p$ Probability of success $p$ is constant per trial. Probability of success $p$ is constant per trial.
Support $k \in {0, 1, 2, \dots, n}$ $k \in {1, 2, 3, \dots, \infty}$

Rigorous Mathematical Foundations

1. Binomial Distribution

Let $X$ be a binomial random variable representing the number of successes in $n$ independent Bernoulli trials with success probability $p$.


2. Geometric Distribution

Let $Y$ be a geometric random variable representing the number of trials required to observe the first success.

Solving algebraically for $E[Y]$: $$E[Y] - (1-p)E[Y] = 1$$ $$EY = 1 \implies p \cdot E[Y] = 1 \implies E[Y] = \frac{1}{p}$$


Computational Verification (Python)

To verify discrete distribution behavior and simulate convergence to mathematical expectation (Monte Carlo Method):

import numpy as np
from scipy import stats

def simulate_discrete_models(p=0.25, n_trials=10, n_simulations=1_000_000):
    """
    Simulates Binomial and Geometric distributions to mathematically verify
    theoretical expectations and variances.
    """
    # Seed for reproducibility
    np.random.seed(42)

    # 1. Binomial Simulation: X ~ Binom(n=10, p=0.25)
    binom_sim = np.random.binomial(n=n_trials, p=p, size=n_simulations)

    theoretical_binom_mean = n_trials * p
    theoretical_binom_var = n_trials * p * (1 - p)

    # 2. Geometric Simulation: Y ~ Geom(p=0.25)
    geom_sim = np.random.geometric(p=p, size=n_simulations)

    theoretical_geom_mean = 1 / p
    theoretical_geom_var = (1 - p) / (p**2)

    print("--- BINOMIAL MODEL METRICS ---")
    print(f"Theoretical E[X]: {theoretical_binom_mean:.4f} | Simulated: {np.mean(binom_sim):.4f}")
    print(f"Theoretical Var(X): {theoretical_binom_var:.4f} | Simulated: {np.var(binom_sim):.4f}\n")

    print("--- GEOMETRIC MODEL METRICS ---")
    print(f"Theoretical E[Y]: {theoretical_geom_mean:.4f} | Simulated: {np.mean(geom_sim):.4f}")
    print(f"Theoretical Var(Y): {theoretical_geom_var:.4f} | Simulated: {np.var(geom_sim):.4f}")

if __name__ == "__main__":
    simulate_discrete_models()

3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances

To secure a Score 5 (specifically an "Essentially Correct" (E) rating on discrete probability FRQs), students must avoid common presentation errors that downgrade responses from Score 4 level to Score 3 or below.

       [ Score 4 Student Response ]                    [ Score 5 Student Response ]
 ┌──────────────────────────────────────┐        ┌──────────────────────────────────────┐
 │ Uses calculator syntax directly:     │        │ Formulates probability mathematically│
 │   binomcdf(10, 0.3, 4)               │        │   P(X <= 4) = sum(...)                │
 ├──────────────────────────────────────┤        ├──────────────────────────────────────┤
 │ Omits variable definitions and       │   VS   │ Explicitly defines random variable:  │
 │ conditional checks.                  │        │   Let X = count of success...        │
 ├──────────────────────────────────────┤        ├──────────────────────────────────────┤
 │ Assumes independence automatically   │        │ Verifies 10% condition for            │
 │ without verifying sample size bounds.│        │ sampling without replacement.        │
 └──────────────────────────────────────┘        └──────────────────────────────────────┘

Pitfall 1: Unlabeled Calculator Syntax ("Calculator Dump")

Pitfall 2: Confusing Cumulative Bounds ($P(X < k)$ vs $P(X \le k)$)

Because these are discrete distributions, continuous probability logic does not apply: $$P(X < 3) \neq P(X \le 3)$$ $$P(X < 3) = P(X \le 2)$$ For geometric models, "takes more than $k$ attempts" MUST be written as: $$P(Y > k) = (1-p)^k$$ Do not evaluate $P(Y > k)$ as $1 - P(Y \le k-1)$ due to common indexing off-by-one errors.

Pitfall 3: Failing the $10\%$ Condition Check

Sampling without replacement breaks independence. If $N$ is the population size and $n$ is the sample size, you must explicitly state: $$\text{"Sampling without replacement is acceptable because } n = 50 \le 0.10 N \text{ (population of items } \ge 500\text{)."}$$


4. Stanford University Placement Pathway

At Stanford University, performance on the AP Statistics exam dictates placement relative to the Department of Statistics and the School of Engineering.

   [ AP Statistics Score 5 ]
               │
               ▼
   Exempts STATS 60 (5 Units)
               │
               ├─────────────────────────────────────────┐
               ▼                                         ▼
   CS 109: Probability for CS            CME 106: Continuous Math/Stats
   (Core for AI/ML/Data Science)          (Engineering Optimization Track)

Direct Course Credit & Acceleration Dynamics

  1. STATS 60 Waiver: Earning a 5 on the AP Statistics exam awards 5 quarter units toward degree completion and satisfies the introductory statistics prerequisite across all majors (including Computer Science, Symbolic Systems, and Management Science & Engineering [MS&E]).
  2. Acceleration into CS 109: CS 109 (Probability for Computer Scientists) bypasses introductory statistical formulas to focus on probabilistic modeling, indicator variables, expectations, joint continuous distributions, and Maximum Likelihood Estimation (MLE).
  3. The Discrete Probability Advantage:
  4. In CS 109, discrete models are generalized immediately into Poisson Distributions, Negative Binomial Distributions, and Hypergeometric Models.
  5. Mastery of the indicator variable proof for $E[X] = np$ learned in AP Statistics serves as the fundamental building block for evaluating algorithm runtime complexities (e.g., Randomized QuickSort analysis) in Stanford's core CS architecture.

5. High-Yield Practice Problem & Step-by-Step Solution Checklist

Part I: The Scenario

An automated semiconductor manufacturing line produces microchips. Historically, $8\%$ of the chips produced possess micro-fissures (defects). Microchips are inspected individually by an automated optical system. Assume inspection results for individual chips are mutually independent.


Part II: Step-by-Step Solution & AP Rubric Alignment

Part (a) Solution

Model: Let $X$ represent the number of defective microchips in a sample of $n = 20$. $X$ follows a Binomial Probability Distribution: $X \sim \text{Binom}(n = 20, p = 0.08)$.

Condition Verification (BINS): 1. B (Binary): Each chip is classified as either defective or not defective. 2. I (Independent): Inspection outcomes for individual chips are stated as independent. 3. N (Number of trials fixed): The sample size is fixed at $n = 20$ microchips. 4. S (Success probability constant): The probability of selecting a defective chip is constant at $p = 0.08$.


Part (b) Solution

We must find $P(X \ge 3)$ for $X \sim \text{Binom}(20, 0.08)$.

$$P(X \ge 3) = 1 - P(X \le 2)$$ $$P(X \le 2) = P(X = 0) + P(X = 1) + P(X = 2)$$

Using the Binomial PMF $P(X = k) = \binom{20}{k} (0.08)^k (0.92)^{20-k}$:

$$P(X = 0) = \binom{20}{0} (0.08)^0 (0.92)^{20} \approx 0.1887$$ $$P(X = 1) = \binom{20}{1} (0.08)^1 (0.92)^{19} \approx 0.3282$$ $$P(X = 2) = \binom{20}{2} (0.08)^2 (0.92)^{18} \approx 0.2711$$

$$P(X \le 2) = 0.1887 + 0.3282 + 0.2711 = 0.7880$$ $$P(X \ge 3) = 1 - 0.7880 = 0.2120$$

(Note: Stating $1 - \text{binomcdf}(20, 0.08, 2) = 0.2120$ with parameters defined earns full credit).


Part (c) Solution

Let $Y$ represent the number of microchips inspected until the first defective chip is found. $Y$ follows a Geometric Probability Distribution: $Y \sim \text{Geom}(p = 0.08)$.

We need $P(Y = 5)$: $$P(Y = 5) = (1 - p)^{5-1} \cdot p = (0.92)^4 \cdot (0.08)$$ $$P(Y = 5) = (0.7164) \cdot (0.08) \approx 0.0573$$


Part (d) Solution

We evaluate the probability of requiring $22$ or more inspections to find the first defective chip, assuming $p = 0.08$:

$$P(Y \ge 22) = P(\text{first } 21 \text{ chips are not defective}) = (1 - p)^{21}$$ $$P(Y \ge 21 + 1) = (0.92)^{21} \approx 0.1708$$

Conclusion / Decision Inference: Because $P(Y \ge 22) \approx 0.1708$ is greater than standard significance levels ($\alpha = 0.05$), this outcome is not unusually rare under the baseline assumption that $p = 0.08$. Roughly $17\%$ of the time, a process with an $8\%$ defect rate will take $22$ or more trials to observe a failure simply due to random variation. Therefore, there is insufficient statistical evidence to conclude that the defect rate has decreased below $0.08$.


Scoring Checklist for Student Verification

Part Score Element Requirement for "Essentially Correct" (E)
(a) Model & Conditions Names Binomial model ($n=20, p=0.08$) and verifies BINS in context.
(b) Calculation Shows correct setup ($1 - P(X \le 2)$) and yields $0.2120$.
(c) Geometric PMF Shows correct geometric calculation $(0.92)^4 (0.08) = 0.0573$.
(d) Inference Logic Correctly computes $P(Y \ge 22) = (0.92)^{21} \approx 0.1708$ and links the high probability ($0.1708 > 0.05$) to a lack of evidence for process improvement.

Aiming for a Score 5 in Statistics?

Secure admission and advanced standing at top institutions like Stanford University with elite 1-on-1 AP STEM mentorship.

無料相談・学習プラン診断