Independent or paired
◈ 6 cardsTwo disjoint groups of units give independent samples and a question about μ₁ − μ₂; the same unit measured twice, or units matched in pairs, gives paired data and a question about μ_D — whatever the question calls it.
The first decision is about the design
Every two-mean question on the paper starts with a decision that is made before any arithmetic: are the two samples independent, or are the data paired? The wrong answer here is not a lost mark — it is a wrong procedure, a wrong df and a wrong conclusion, all marked as one error.
Independent samples arise when two separate groups of units are measured, and nothing links a unit in one group to a unit in the other: 30 clients on the old invoice template and 30 different clients on the new one. The parameter is the difference of two population means, , estimated by .
Paired data arise when each observation in one sample is naturally matched with exactly one in the other — most often because they are the same unit measured twice. The parameter is the mean of the differences, , estimated by , and the analysis is a one-sample problem on the differences (L11.6).
Worked example — five designs
- Maple Ledger draws 25 clients from its Kitchener office and 25 from its Guelph office and compares mean days-to-pay. Two disjoint groups → independent.
- The same eight client stores report weekly sales before a promotion and again after it. Each store appears twice → paired.
- Twenty bank branches are sorted by size and split into ten pairs of similar size; in each pair one branch gets the new queuing system. The pairs were built on purpose → paired (matched pairs).
- A study of spending habits measures one twin from each of 40 twin pairs under scheme A and the other under scheme B → paired.
- A random sample of 50 invoices from 2024 and another random sample of 50 from 2025 → independent (different invoices, no link).
The test in every case: can I write the data as one column of differences, each difference belonging to one unit or one pair? If yes, paired. If the two columns could be shuffled independently without losing anything, independent.
The wording can mislead
The paper may describe design 2 as "two groups of measurements" or "two samples of eight" — it is still paired, because the eight stores are the same eight stores. Conversely, equal sample sizes do not make data paired: design 5 has and nothing pairs invoice 7 of 2024 with invoice 7 of 2025. Look for the unit, not the count.
Why pairing is worth the trouble
Stores differ enormously in size; a difference of a few thousand dollars between two different stores says nothing about the promotion. The same store before and after removes the store-to-store variability from the comparison, so the promotion effect stands out against a much smaller background noise. The price is that the analysis must honour the pairing: eight pairs give , not , and the SD is that of the eight differences, not of the two columns.