AP Statistics Exam Mastery Guide: Discrete Probability Models & Expected Value
Target Exam: AP Statistics (Score 5 Focus)
Target Institution: Stanford University
Equivalent Course & Credit: STATS 60: Introduction to Statistical Methods (5 Quarter Units)
Accelerated Sequence: CS 109 (Probability for Computer Scientists) / CME 106
1. Introduction & AP Exam Weight
Discrete probability models—specifically the Binomial and Geometric distributions—form the mathematical bridge between basic probability rules and advanced statistical inference. Within the AP Statistics Course and Exam Description (CED), these topics reside in Unit 4: Random Variables, Sets, and Probability Distributions (covering Topic 4.10 through Topic 4.12).
While Unit 4 directly constitutes 10–20% of the AP Exam multiple-choice section, discrete models are foundational to performance across the entire exam: * Binomial models underpin the sampling distribution of a sample proportion ($\hat{p}$), setting up $z$-procedures, $1$-prop $z$-tests, and $2$-prop $z$-intervals. * Geometric models test discrete infinite sequences, expectation, and conditional probability bounds. * Expected value ($E[X]$) calculations appear regularly in Free Response Questions (FRQs) requiring decision analysis and expected profit/loss evaluations.
At elite institutions like Stanford University, a Score of 5 waives STATS 60 and grants 5 quarter units toward graduation. More importantly, it unlocks early enrollment into CS 109: Probability for Computer Scientists. CS 109 assumes an immediate, rigorous intuition for indicator random variables, expectations, variance, and memoryless properties—concepts derived directly from mastering Binomial and Geometric distributions.
2. Deep Concept Breakdown
Model Disambiguation: BINS vs. BITS
To secure full credit on AP FRQs, you must explicitly state conditions in context before applying formulas.
| Characteristic | Binomial Model ($X \sim \text{Binom}(n, p)$) | Geometric Model ($Y \sim \text{Geom}(p)$) |
|---|---|---|
| Definition | Counts number of successes in $n$ trials. | Counts number of trials until the first success. |
| Acronym | BINS | BITS |
| Binary | Success / Failure outcomes. | Success / Failure outcomes. |
| Independance | Trials are independent (check $10\%$ condition if sampling without replacement). | Trials are independent (check $10\%$ condition if sampling without replacement). |
| N / T | Number of trials fixed ($n$). | Trials variable (counting until 1st success). |
| Success $p$ | Probability of success $p$ is constant per trial. | Probability of success $p$ is constant per trial. |
| Support | $k \in {0, 1, 2, \dots, n}$ | $k \in {1, 2, 3, \dots, \infty}$ |
Rigorous Mathematical Foundations
1. Binomial Distribution
Let $X$ be a binomial random variable representing the number of successes in $n$ independent Bernoulli trials with success probability $p$.
-
Probability Mass Function (PMF): $$P(X = k) = \binom{n}{k} p^k (1-p)^{n-k}, \quad \text{where } \binom{n}{k} = \frac{n!}{k!(n-k)!}$$
-
Expected Value Derivation via Indicator Variables:
Let $I_i$ be an indicator random variable for the $i$-th trial, where $I_i = 1$ if success, $0$ if failure. $$X = \sum_{i=1}^{n} I_i$$ By linearity of expectation: $$E[X] = E\left[\sum_{i=1}^{n} I_i\right] = \sum_{i=1}^{n} E[I_i]$$ Since $E[I_i] = 1 \cdot p + 0 \cdot (1-p) = p$: $$E[X] = \sum_{i=1}^{n} p = np$$ -
Variance Derivation: Because the trials are independent, $\text{Var}(X) = \sum_{i=1}^{n} \text{Var}(I_i)$. $$\text{Var}(I_i) = E[I_i^2] - (E[I_i])^2 = p - p^2 = p(1-p)$$ $$\text{Var}(X) = np(1-p) \implies \sigma_X = \sqrt{np(1-p)}$$
2. Geometric Distribution
Let $Y$ be a geometric random variable representing the number of trials required to observe the first success.
-
Probability Mass Function (PMF): $$P(Y = k) = (1-p)^{k-1} p$$
-
Cumulative Distribution Function (CDF): The probability that it takes more than $k$ trials to see a success is the probability of $k$ consecutive failures: $$P(Y > k) = (1-p)^k \implies P(Y \le k) = 1 - (1-p)^k$$
-
Expected Value Derivation via Conditional Expectation: Using the Law of Total Expectation: $$E[Y] = 1 + (1-p)E[Y]$$ (Interpretation: 1 trial is always used. If it fails [prob $1-p$], the process restarts identically.)
Solving algebraically for $E[Y]$: $$E[Y] - (1-p)E[Y] = 1$$ $$EY = 1 \implies p \cdot E[Y] = 1 \implies E[Y] = \frac{1}{p}$$
- Variance: $$\text{Var}(Y) = \frac{1-p}{p^2} \implies \sigma_Y = \frac{\sqrt{1-p}}{p}$$
Computational Verification (Python)
To verify discrete distribution behavior and simulate convergence to mathematical expectation (Monte Carlo Method):
import numpy as np
from scipy import stats
def simulate_discrete_models(p=0.25, n_trials=10, n_simulations=1_000_000):
"""
Simulates Binomial and Geometric distributions to mathematically verify
theoretical expectations and variances.
"""
# Seed for reproducibility
np.random.seed(42)
# 1. Binomial Simulation: X ~ Binom(n=10, p=0.25)
binom_sim = np.random.binomial(n=n_trials, p=p, size=n_simulations)
theoretical_binom_mean = n_trials * p
theoretical_binom_var = n_trials * p * (1 - p)
# 2. Geometric Simulation: Y ~ Geom(p=0.25)
geom_sim = np.random.geometric(p=p, size=n_simulations)
theoretical_geom_mean = 1 / p
theoretical_geom_var = (1 - p) / (p**2)
print("--- BINOMIAL MODEL METRICS ---")
print(f"Theoretical E[X]: {theoretical_binom_mean:.4f} | Simulated: {np.mean(binom_sim):.4f}")
print(f"Theoretical Var(X): {theoretical_binom_var:.4f} | Simulated: {np.var(binom_sim):.4f}\n")
print("--- GEOMETRIC MODEL METRICS ---")
print(f"Theoretical E[Y]: {theoretical_geom_mean:.4f} | Simulated: {np.mean(geom_sim):.4f}")
print(f"Theoretical Var(Y): {theoretical_geom_var:.4f} | Simulated: {np.var(geom_sim):.4f}")
if __name__ == "__main__":
simulate_discrete_models()
3. Common AP Exam Pitfalls & Score 5 Scoring Rubric Nuances
To secure a Score 5 (specifically an "Essentially Correct" (E) rating on discrete probability FRQs), students must avoid common presentation errors that downgrade responses from Score 4 level to Score 3 or below.
[ Score 4 Student Response ] [ Score 5 Student Response ]
┌──────────────────────────────────────┐ ┌──────────────────────────────────────┐
│ Uses calculator syntax directly: │ │ Formulates probability mathematically│
│ binomcdf(10, 0.3, 4) │ │ P(X <= 4) = sum(...) │
├──────────────────────────────────────┤ ├──────────────────────────────────────┤
│ Omits variable definitions and │ VS │ Explicitly defines random variable: │
│ conditional checks. │ │ Let X = count of success... │
├──────────────────────────────────────┤ ├──────────────────────────────────────┤
│ Assumes independence automatically │ │ Verifies 10% condition for │
│ without verifying sample size bounds.│ │ sampling without replacement. │
└──────────────────────────────────────┘ └──────────────────────────────────────┘
Pitfall 1: Unlabeled Calculator Syntax ("Calculator Dump")
- The Error: Writing
binompdf(8, 0.15, 3) = 0.0839or1 - geometcdf(0.2, 4)as primary work on the exam booklet. - Rubric Impact: Automatic downgrade to Partially Correct (P) or Incomplete (I). College Board rubrics mandate clear mathematical notation or parameter definitions.
- Score 5 Requirement: If you use calculator commands, you MUST explicitly define parameters: $$\text{Binomial with } n = 8, p = 0.15, k = 3 \implies P(X = 3) = \binom{8}{3}(0.15)^3(0.85)^5 \approx 0.0839$$
Pitfall 2: Confusing Cumulative Bounds ($P(X < k)$ vs $P(X \le k)$)
Because these are discrete distributions, continuous probability logic does not apply: $$P(X < 3) \neq P(X \le 3)$$ $$P(X < 3) = P(X \le 2)$$ For geometric models, "takes more than $k$ attempts" MUST be written as: $$P(Y > k) = (1-p)^k$$ Do not evaluate $P(Y > k)$ as $1 - P(Y \le k-1)$ due to common indexing off-by-one errors.
Pitfall 3: Failing the $10\%$ Condition Check
Sampling without replacement breaks independence. If $N$ is the population size and $n$ is the sample size, you must explicitly state: $$\text{"Sampling without replacement is acceptable because } n = 50 \le 0.10 N \text{ (population of items } \ge 500\text{)."}$$
4. Stanford University Placement Pathway
At Stanford University, performance on the AP Statistics exam dictates placement relative to the Department of Statistics and the School of Engineering.
[ AP Statistics Score 5 ]
│
▼
Exempts STATS 60 (5 Units)
│
├─────────────────────────────────────────┐
▼ ▼
CS 109: Probability for CS CME 106: Continuous Math/Stats
(Core for AI/ML/Data Science) (Engineering Optimization Track)
Direct Course Credit & Acceleration Dynamics
- STATS 60 Waiver: Earning a 5 on the AP Statistics exam awards 5 quarter units toward degree completion and satisfies the introductory statistics prerequisite across all majors (including Computer Science, Symbolic Systems, and Management Science & Engineering [MS&E]).
- Acceleration into CS 109: CS 109 (Probability for Computer Scientists) bypasses introductory statistical formulas to focus on probabilistic modeling, indicator variables, expectations, joint continuous distributions, and Maximum Likelihood Estimation (MLE).
- The Discrete Probability Advantage:
- In CS 109, discrete models are generalized immediately into Poisson Distributions, Negative Binomial Distributions, and Hypergeometric Models.
- Mastery of the indicator variable proof for $E[X] = np$ learned in AP Statistics serves as the fundamental building block for evaluating algorithm runtime complexities (e.g., Randomized QuickSort analysis) in Stanford's core CS architecture.
5. High-Yield Practice Problem & Step-by-Step Solution Checklist
Part I: The Scenario
An automated semiconductor manufacturing line produces microchips. Historically, $8\%$ of the chips produced possess micro-fissures (defects). Microchips are inspected individually by an automated optical system. Assume inspection results for individual chips are mutually independent.
- (a) Identify the probability model for the number of defective chips found in a random sample of $20$ chips. Justify your model by checking all required conditions in context.
- (b) Calculate the probability that at least $3$ chips in a sample of $20$ are defective. Show your work clearly.
- (c) Calculate the probability that the first defective chip is identified on the $5\text{th}$ chip inspected.
- (d) An engineer optimizes the optical sensor, claiming it reduces the defect detection threshold. In a new test run, the first defective chip is found on the $22\text{nd}$ chip inspected. Calculate $P(Y \ge 22)$ under the assumption that the defect rate remains $p = 0.08$. Based on this result, is there compelling evidence that the defect rate $p$ has decreased below $0.08$? Justify your answer.
Part II: Step-by-Step Solution & AP Rubric Alignment
Part (a) Solution
Model: Let $X$ represent the number of defective microchips in a sample of $n = 20$. $X$ follows a Binomial Probability Distribution: $X \sim \text{Binom}(n = 20, p = 0.08)$.
Condition Verification (BINS): 1. B (Binary): Each chip is classified as either defective or not defective. 2. I (Independent): Inspection outcomes for individual chips are stated as independent. 3. N (Number of trials fixed): The sample size is fixed at $n = 20$ microchips. 4. S (Success probability constant): The probability of selecting a defective chip is constant at $p = 0.08$.
Part (b) Solution
We must find $P(X \ge 3)$ for $X \sim \text{Binom}(20, 0.08)$.
$$P(X \ge 3) = 1 - P(X \le 2)$$ $$P(X \le 2) = P(X = 0) + P(X = 1) + P(X = 2)$$
Using the Binomial PMF $P(X = k) = \binom{20}{k} (0.08)^k (0.92)^{20-k}$:
$$P(X = 0) = \binom{20}{0} (0.08)^0 (0.92)^{20} \approx 0.1887$$ $$P(X = 1) = \binom{20}{1} (0.08)^1 (0.92)^{19} \approx 0.3282$$ $$P(X = 2) = \binom{20}{2} (0.08)^2 (0.92)^{18} \approx 0.2711$$
$$P(X \le 2) = 0.1887 + 0.3282 + 0.2711 = 0.7880$$ $$P(X \ge 3) = 1 - 0.7880 = 0.2120$$
(Note: Stating $1 - \text{binomcdf}(20, 0.08, 2) = 0.2120$ with parameters defined earns full credit).
Part (c) Solution
Let $Y$ represent the number of microchips inspected until the first defective chip is found. $Y$ follows a Geometric Probability Distribution: $Y \sim \text{Geom}(p = 0.08)$.
We need $P(Y = 5)$: $$P(Y = 5) = (1 - p)^{5-1} \cdot p = (0.92)^4 \cdot (0.08)$$ $$P(Y = 5) = (0.7164) \cdot (0.08) \approx 0.0573$$
Part (d) Solution
We evaluate the probability of requiring $22$ or more inspections to find the first defective chip, assuming $p = 0.08$:
$$P(Y \ge 22) = P(\text{first } 21 \text{ chips are not defective}) = (1 - p)^{21}$$ $$P(Y \ge 21 + 1) = (0.92)^{21} \approx 0.1708$$
Conclusion / Decision Inference: Because $P(Y \ge 22) \approx 0.1708$ is greater than standard significance levels ($\alpha = 0.05$), this outcome is not unusually rare under the baseline assumption that $p = 0.08$. Roughly $17\%$ of the time, a process with an $8\%$ defect rate will take $22$ or more trials to observe a failure simply due to random variation. Therefore, there is insufficient statistical evidence to conclude that the defect rate has decreased below $0.08$.
Scoring Checklist for Student Verification
| Part | Score Element | Requirement for "Essentially Correct" (E) |
|---|---|---|
| (a) | Model & Conditions | Names Binomial model ($n=20, p=0.08$) and verifies BINS in context. |
| (b) | Calculation | Shows correct setup ($1 - P(X \le 2)$) and yields $0.2120$. |
| (c) | Geometric PMF | Shows correct geometric calculation $(0.92)^4 (0.08) = 0.0573$. |
| (d) | Inference Logic | Correctly computes $P(Y \ge 22) = (0.92)^{21} \approx 0.1708$ and links the high probability ($0.1708 > 0.05$) to a lack of evidence for process improvement. |