Statistical Sampling
Statistical sampling is the selection of a subset of a population in order to say something about the whole. The part that decides what you may conclude is not the size of the subset but the method of selection. Definitions here follow Statistics Canada’s survey methodology material, checked in September 2026.
Three things have to be named before a number means anything: the population you intend to describe, the frame you actually selected from, and the selection method. Most misreadings come from the second and third being different from what the reader assumes.
Probability and non-probability selection
The dividing line is whether chance did the selecting. Probability sampling “refers to the selection of a sample from a population, when this selection is based on the principle of randomization, that is, random selection or chance,” and it “requires that each member of the survey population has a known probability of being included in the sample.” Those probabilities need not be equal — they need to be known, because that is what lets an estimate be weighted back to the population.
Non-probability sampling “is a method of selecting units from a population using a subjective (i.e. non-random) method.” The consequence is stated plainly in the same documentation: “it is impossible to determine the probability that a unit in the population is selected for the sample, so reliable estimates and estimates of sampling error cannot be computed.”
| Probability selection | Non-probability selection | |
|---|---|---|
| Who gets in | Decided by chance, with a known inclusion probability | Decided by a person, a rule of convenience, or by who showed up |
| What you can compute | An estimate for the population, and a stated uncertainty for it | A description of the units you happened to collect |
| What it costs | “More complex, more time-consuming and usually more costly” | Cheap, fast, and frequently already sitting in a table |
| What it needs you to assume | That the frame covers the population | That the sample “is representative of the population. This is often a risky assumption to make” |
The guidance where generalization is the point is direct: use probability sampling instead. That does not make non-probability data useless — it makes the claim you can attach to it smaller.
The frame is where quiet errors live
Selection happens from a list, and that list is rarely the population. Coverage is “the completeness of the information for the target population that would be derived if all of the frame units were to be surveyed,” and it fails in two directions: undercoverage “when units are erroneously omitted from the frame file,” overcoverage “when units are incorrectly included on the frame file (e.g. dead units).”
Neither announces itself in the output. The documented effect is that frame imperfections “are likely to bias or diminish the reliability of the survey estimates,” and that under- and overcoverage “can undermine both the relevance and the accuracy of survey results.” So the frame deserves the same scrutiny as the analysis: what is on this list, what should be and is not, and what is on it twice.
Why this applies to data you did not collect
Nothing above is limited to surveys, which is the part most easily missed in a company with large event tables. A log of application activity looks like a census and behaves like a sample with an unexamined frame. Four routes produce one without anyone choosing it.
- Instrumentation coverage. The rows describe the platforms, versions, and regions where the tracking works. A client that fails to report is absent, not zero.
- Who transacts at all. Customers who never got far enough to be recorded are outside the frame, which is exactly the group most questions about conversion are about.
- Who responds. Opt-in feedback, reviews, and support tickets are selected by willingness to act, and that willingness correlates with the sentiment being measured.
- What survived. Retention windows, deletions, and pipelines that dropped rows mean the table describes what is still there — a survivorship bias in the plainest form.
Each of those makes the data a non-probability sample of the population someone will name in the conclusion. The data is still worth using; the sentence attached to it has to match it.
Size does not fix selection
This is the practical error worth being blunt about. Collecting more rows narrows the uncertainty that comes from having only a subset. It does nothing to the difference between the group you selected from and the group you meant, because that difference is not random — it does not shrink as n grows.
A hundred thousand responses from people who chose to respond estimate the views of people who choose to respond, very precisely. Precision on the wrong quantity reads as authority, which is why large convenience samples mislead more effectively than small ones.
Two habits make this tractable without a survey programme. Write the population next to the number — one sentence, in the report, saying which units this describes. And when the selection is not random, say what would have to be true for the conclusion to transfer, since that assumption is the actual claim being made.
How sampling fits with uncertainty, significance, and the other judgments behind reading a result is worked through in Reading a Result Honestly.
References: Statistics Canada, Probability sampling and Non-probability sampling; Statistics Canada, Quality Guidelines: Coverage and frames.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
