Memra

Target population, study population, and bias

◈ 5 cards

Three layers between the question and the data — and why a large sample does not cure a flaw at any of them.

Three layers, three places to go wrong

Between the question a business asks and the numbers it ends up with sit three sets of units.

  1. The target population — the group the question is actually about.
  2. The study population (the sampling frame) — the group you can actually list and draw from.
  3. The sample — the units you measure.

Each boundary is a place where bias enters, and each kind of bias has a name.

  • Study bias (undercoverage) lives between the target and the study population: the frame misses part of the target.
  • Sample bias lives between the study population and the sample: the units that end up measured differ systematically from the frame — non-response and self-selection are the usual culprits.
  • Measurement bias lives inside the sample: the recorded value differs systematically from the true one — a leading question, a faulty instrument, a respondent who shades the truth.

Worked example — one survey, three flaws

A consultancy wants the average tax overpayment among Ontario small-business owners (the target). It samples from a Chamber of Commerce mailing list — but Chamber members are larger and more established than the typical small business, so the frame undercovers the target. Study bias.

Of the 2,000 owners mailed, 22 % reply. Owners who feel they overpaid are the ones motivated to answer; those with nothing to complain about bin the letter. The 440 replies are a self-selected slice of the frame. Sample bias.

The question reads: "How much did you overpay in tax last year?" It presumes an overpayment and invites an inflated figure. Measurement bias.

Three biases, three layers, and the reported average is wrong in the same direction three times.

Why 40,000 responses do not help

Suppose the consultancy had 40,000 replies instead of 440. A bigger sample shrinks the random wobble around whatever the sampling process is centred on — and this process is centred on the wrong number. More data makes a biased answer more precise, not more accurate: it converges, confidently, on the mean of established, aggrieved, prompted respondents. Representativeness comes from how units are chosen, not how many. A random sample of 500 Ontario small businesses with follow-up of non-responders beats 40,000 volunteers.

Diagnosing a described survey

On the paper you are given a survey and asked to name the flaw and its layer. Ask three questions in order: Who is the question about, and who could actually be drawn? (target → study). Of those who could be drawn, who actually answered? (study → sample). Did the measurement itself push the answers? (measurement). Name the layer, name the bias, and say what would fix it — usually a proper frame, random selection, and follow-up.

Target populationall Ontario small-business owners↓ undercoverage (study bias)Study populationChamber of Commerce mailing list↓ non-response / self-selection (sample bias)Sample440 replies; leading question = measurement biasthe questionthe data
The three layers of the tax-overpayment survey. Each boundary is where one bias enters; measurement bias sits inside the sample itself. Sample size shrinks only the random error, never any of these.
NORMAL ~/memra/learn/afm-113/target-population-study-population-and-bias utf-8 LF