Non-probability designs and sampling error
◈ 5 cardsSelf-selection, convenience, quota, judgement — what each bakes in, and why n cannot fix it; sampling error versus non-sampling error.
Four designs the paper wants you to name
A non-probability design is any way of choosing units in which the selection probabilities are unknown — usually because a person, or the units themselves, did the choosing. Four are examined.
- Self-selection (voluntary response). Units opt in. An open web poll; a comment card. It bakes in the motivation to respond: the aggrieved and the enthusiastic answer, the indifferent majority does not.
- Convenience. Whoever is easiest to reach. The first 50 people through the door; your own co-op cohort. It bakes in whatever made them convenient — time of day, location, the fact that they know you.
- Quota. Fill fixed counts per group (30 women, 30 men; 20 per age band) with whoever the interviewer can find. It looks stratified, but there is no random draw inside the quota, so it bakes in the interviewer's choices — and if the quota proportions are wrong, that too.
- Judgement (purposive). An expert picks the "representative" units. It bakes in the expert's model of representativeness, which is exactly what the study was supposed to test.
Worked example — recruiting for a co-op employer survey
A student society wants employers' views on co-op students. Four recruitment methods were proposed. (1) Post the survey link on the society's LinkedIn page and take whoever answers — self-selection. (2) Survey the employers at next week's networking night — convenience. (3) Ask a committee member to fill 10 responses from each of banking, accounting firms, tech and government, choosing firms they know — quota. (4) Have the co-op office nominate the 25 "most typical" employers — judgement.
None of the four lets you say how far the resulting average is likely to be from the true average across all co-op employers, because none has a known selection probability to compute it from.
Two kinds of error
Every estimate misses the parameter for two reasons that behave completely differently.
Sampling error is the random gap between a statistic and the parameter that exists because you measured a sample rather than everyone. Draw another random sample and it changes sign and size. It shrinks with — Module 8 will show it shrinks like — and for a probability sample its typical size can be computed. It is the only error the margin of error in a confidence interval accounts for.
Non-sampling error is everything else: undercoverage, non-response, a leading question, a data-entry slip, a convenience frame. It is systematic — the same sign every time — and does not shrink with . The convenience sample of networking-night employers stays a sample of employers who go to networking nights however many of them you interview.
On the paper: "the survey had only 40 respondents" is a sampling-error concern; "the survey was posted on a fan page" is a non-sampling-error concern. Name which one you are being asked about.