The one-proportion z-test
◈ 5 cardsz = (p̂ − p₀)/√(p₀(1 − p₀)/n): the test’s SE is built from the null value p₀, not from p̂ — the line the exam marks — then the critical value or the p-value from the cumulative table.
The SE comes from the null
A test of asks how far sits from if is true. If it is true, the SE of is known exactly — it is , with no estimate needed. So the statistic is
This is the one place proportions differ from means in shape. The interval (L12.1) has no null value and must use in its SE; the test has and uses it. Writing in a test’s denominator is the error the marker looks for. Conditions, under the null: and .
Worked example — does the error rate exceed the tolerable 10 %?
Maple Ledger’s tolerable invoice error rate is 10 %. The audit found 60 errors in 500 invoices. Is the true rate above 10 %? .
Step 1. , , where is the error rate for all invoices. Upper-tailed, from exceed.
Step 2. .
Step 3 — critical value. Upper tail at 5 %: . Step 3 — p-value. .
Step 4. , equivalently : do not reject at the 5 % level. There is insufficient evidence, at the 5 % level, that the error rate exceeds 10 %.
Say what that does and does not mean: the sample rate of 12 % is above the benchmark, and the interval from L12.1 reaches almost to 15 %. The rate may well exceed 10 %; the test simply did not establish it with 500 invoices. At the same would reject (; ) — the p-value lets the reader see that.
Two-sided and lower-tailed versions
Does the rate differ from 10 %? — , compare with 1.96 and double the tail: . Is the rate below 10 %? — ; with the estimate leans the wrong way, , and stands without further arithmetic. The tail comes from the question, as always.
The interval and the test disagree slightly, by design
The interval’s SE was 0.01453 (from ); the test’s is 0.01342 (from ). They differ because they answer different questions — how precise is my estimate? versus how surprising is my estimate if is true? — so Module 10’s exact duality holds only approximately for proportions. Here both agree that 0.10 is plausible at 5 % two-sided; near the boundary they can differ, and then the test is the answer to a test question and the interval to an estimation question.