Random Sampling, Selection Bias, Nonresponse, and Stratified Samples
A sample should represent the population from which researchers want to draw conclusions. A large sample alone does not guarantee accurate results. If the sampling method systematically favors certain people, the resulting estimate may be biased.
Suppose the objective is to estimate support for a political candidate among all eligible US voters. Surveying 1,000 people may sound substantial, but the result is meaningful only if those people were selected appropriately.
Population, Sample, and Sampling Frame
Three concepts must be distinguished.
Population
The population is the complete group about which conclusions are desired.
In a voter survey, the target population might be:\[ \text{all eligible US voters}. \]
The exact population must be defined carefully. Eligible voters, registered voters, and likely voters are not identical groups.
Sample
The sample is the subset from which data are actually collected:\[ S=\{i_1,i_2,\ldots,i_n\}. \]
If 1,000 voters are surveyed, then:\[ n=1000. \]
Sampling frame
The sampling frame is the operational list or mechanism used to reach members of the population.
Examples include:
- a voter-registration list;
- an address-based list;
- a telephone-number frame;
- a membership database;
- a household registry.
A sampling frame may fail to cover the entire target population. The difference between the population and the frame is a major potential source of error.
The Danger of Convenience Samples
A convenience sample consists of people who are easy to reach.
Examples include:
- residents of the researcher’s hometown;
- people leaving one shopping center;
- members of one social-media group;
- coworkers;
- visitors to one website;
- customers at one location.
Convenience sampling is fast and inexpensive, but it usually does not give every population member a known or comparable chance of selection.
Suppose a researcher surveys 1,000 voters from one city. The local sample may differ from the national electorate in:
- political preference;
- age;
- income;
- education;
- race and ethnicity;
- urbanization;
- occupation;
- regional issues.
Increasing the local sample from 1,000 to 10,000 reduces random sampling variability within that city, but it does not make the city representative of the country.
A larger biased sample can produce a more precise estimate of the wrong population.
What Is Bias?
In sampling, bias is a systematic tendency for a procedure to produce estimates that are too high, too low, or otherwise unrepresentative.
Let \(\theta\) be a population quantity and \(\hat{\theta}\) its estimator. Statistical bias is defined as:\[ \operatorname{Bias}(\hat{\theta}) = E[\hat{\theta}]-\theta. \]
If\[ E[\hat{\theta}]=\theta, \]
the estimator is unbiased under the assumed sampling design and response process.
In practice, the word bias is also used more broadly for systematic distortions caused by coverage, selection, nonresponse, measurement, or data-processing problems.
Bias differs from random error:
- random error varies unpredictably from sample to sample;
- bias systematically pushes results in a particular direction.
A large sample generally reduces random error but does not automatically remove bias.
Selection Bias
Selection bias occurs when the process used to obtain the sample makes some types of population members more likely to appear than others in a way related to the outcome.
A hometown convenience sample creates selection bias because geography is related to voting behavior.
Other examples include:
- estimating employee satisfaction using only managers;
- measuring public transportation quality by surveying only car owners;
- estimating fitness using members of a gym;
- studying internet access with an online-only survey;
- evaluating a product using only current subscribers.
The central question is:
Does the selection process systematically favor people whose responses differ from those who are less likely to be selected?
If the answer is yes, the estimate may be biased.
Undercoverage Bias
Undercoverage occurs when some members of the target population are absent or inadequately represented in the sampling frame.
Suppose a telephone survey reaches only one type of telephone service. People without that service have no chance of selection.
Other examples include:
- a housing survey that omits people without fixed addresses;
- an email survey that excludes people without reliable internet access;
- a voter survey based on an outdated registration list;
- a workplace survey that omits night-shift employees.
Undercoverage is related to selection bias, but identifying it separately is useful because the problem begins with the sampling frame.
Nonresponse Bias
Nonresponse occurs when selected individuals do not provide usable responses.
Nonresponse bias arises when:
- response rates differ among groups; and
- those groups differ on the variable being studied.
Suppose parents of young children are less likely to answer a survey around dinner time. If parenting status is related to the survey topic, their lower response rate can distort the result.
Let:\[ R_i= \begin{cases} 1, & \text{if person }i\text{ responds}\\ 0, & \text{otherwise}. \end{cases} \]
Nonresponse becomes problematic when the response probability\[ P(R_i=1) \]
is associated with the person’s answer \(Y_i\), even after accounting for observed characteristics.
A low response rate increases concern, but response rate alone does not determine bias. A survey can have a modest response rate with limited bias if respondents and nonrespondents are similar on the outcome. Conversely, even a seemingly strong response rate can be biased if the missing respondents differ systematically.
Voluntary Response Bias
A voluntary response sample allows people to decide for themselves whether to participate.
Examples include:
- public website polls;
- call-in surveys;
- unsolicited product reviews;
- open comment forms;
- social-media polls.
People with strong opinions are often more motivated to respond than people with moderate or neutral experiences.
For business reviews, customers who had extremely positive or extremely negative experiences may be more likely to write a review. Consequently, the observed reviews may overrepresent the tails of the satisfaction distribution.
Conceptually:\[ P(\text{response}\mid\text{extreme opinion}) > P(\text{response}\mid\text{moderate opinion}). \]
Voluntary response bias can be viewed as a form of self-selection bias. It is related to nonresponse, but the initial sample is not selected by a controlled probability procedure.
Response and Measurement Bias
Even a well-selected sample can produce misleading data if questions or measurements influence the answers.
Response bias can result from:
- leading wording;
- sensitive questions;
- interviewer influence;
- social desirability;
- inaccurate memory;
- confusing response categories;
- question order;
- measurement errors.
For example, these questions may produce different answers:
“Do you support the proposed policy?”
and
“Do you support the costly proposed policy?”
Random sampling addresses who is surveyed. It does not automatically correct how questions are asked or how responses are measured.
Comparing Major Sources of Bias
| Bias | Primary cause | Example |
|---|---|---|
| Selection bias | Sample acquisition favors certain groups | Surveying only one hometown |
| Undercoverage | Sampling frame omits population members | Excluding households without covered contact methods |
| Nonresponse bias | Selected people respond at different rates | Busy households respond less often |
| Voluntary response bias | Participants select themselves | Open online ratings |
| Response bias | Answers are systematically distorted | Leading survey wording |
These problems can occur simultaneously.
Probability Sampling
Probability sampling uses a planned random mechanism to select participants.
Its essential properties are:
- selection probabilities are determined by design;
- randomness is introduced deliberately;
- sampling uncertainty can be quantified;
- population estimates can be constructed using known selection probabilities.
Probability sampling does not guarantee a perfect sample. Nonresponse, undercoverage, and measurement error may remain. However, it provides a defensible foundation for inference.
Simple Random Sampling
A simple random sample without replacement of size \(n\) from a population of size \(N\) is a design in which every possible subset of \(n\) distinct population members has the same probability of selection.
The number of possible samples is:\[ \binom{N}{n}. \]
Each possible sample has probability:\[ \frac{1}{\binom{N}{n}}. \]
Every individual has the same inclusion probability:\[ \pi_i=\frac{n}{N}. \]
The phrase “without replacement” means that once a person is selected, the same person cannot be selected again.
Selecting a Simple Random Sample
If a complete population list is available:
- assign every population member a unique identifier;
- use a reliable random-number generator;
- select \(n\) distinct identifiers;
- contact the corresponding individuals;
- record and address nonresponse according to a predefined protocol.
For example, if a population contains 100,000 voters, label them:\[ 1,2,\ldots,100000. \]
Then select 1,000 distinct identifiers uniformly at random.
A systematic computer procedure is preferable to informal methods such as choosing names that “look random.”
Simple Random Sampling Example
Suppose a population contains:\[ N=10{,}000 \]
voters, and researchers select:\[ n=1{,}000. \]
Each voter’s inclusion probability is:\[ \pi_i = \frac{1{,}000}{10{,}000} = 0.10. \]
Thus, each voter has a 10% probability of appearing in the sample.
For a binary response \(Y_i\), such as candidate support, the sample proportion is:\[ \hat{p} = \frac{1}{n} \sum_{i\in S}Y_i. \]
Under a true simple random design and ideal response conditions:\[ E[\hat{p}]=p, \]
where \(p\) is the population proportion.
Sampling Variability
Different random samples produce different estimates.
If the true proportion is \(p\), the approximate standard error of a sample proportion under independent sampling is:\[ \operatorname{SE}(\hat{p}) \approx \sqrt{\frac{p(1-p)}{n}}. \]
When sampling without replacement from a finite population, a finite-population correction may be used:\[ \operatorname{SE}(\hat{p}) = \sqrt{ \frac{p(1-p)}{n} \cdot \frac{N-n}{N-1} }. \]
The factor\[ \sqrt{\frac{N-n}{N-1}} \]
reduces uncertainty when the sample is a substantial fraction of the population.
Random Sampling Versus Random Assignment
These concepts serve different purposes.
Random sampling
Random sampling selects people from a population. It supports generalization from the sample to the population.
Random assignment
Random assignment places participants into treatment groups. It supports causal comparisons between treatments when other design requirements are satisfied.
A study may have:
- random sampling without random assignment;
- random assignment without random sampling;
- both;
- neither.
The conclusions supported by the study depend on which forms of randomization were used.
Random-Digit Dialing
Random-digit dialing selects telephone numbers using a probability-based procedure rather than relying only on a fixed directory.
Historically, it helped reach listed and unlisted landline households. Modern implementation is more complicated because communication patterns and coverage have changed.
Challenges include:
- mobile-only households;
- multiple numbers per person;
- geographic mobility;
- spam filtering;
- low contact rates;
- unknown eligibility;
- unequal probabilities of selection;
- legal and operational restrictions.
Modern surveys may combine several sampling frames or contact modes. The statistical principle remains the same: the coverage and selection probability of each population member must be understood as well as possible.
Stratified Random Sampling
Stratified random sampling divides the population into nonoverlapping groups called strata. A probability sample is then drawn independently from every stratum.
If the population is divided into \(H\) strata:\[ U=U_1\cup U_2\cup\cdots\cup U_H, \]
with\[ U_h\cap U_k=\varnothing \quad\text{for }h\neq k. \]
Possible voter strata include:
- urban, suburban, and rural;
- geographic regions;
- age groups;
- registration categories;
- combinations of selected characteristics.
The strata should cover the entire population, and every population member should belong to exactly one stratum.
Steps in Stratified Sampling
- Define the target population.
- Select variables for forming strata.
- Divide the population into mutually exclusive strata.
- determine the population size \(N_h\) of each stratum.
- Select a random sample of size \(n_h\) within each stratum.
- Collect responses using consistent procedures.
- Combine stratum estimates using appropriate weights.
Randomness is applied within each stratum rather than only across the undivided population.
Combining Stratum Estimates
Suppose stratum \(h\) contains \(N_h\) population members. Its population weight is:\[ W_h=\frac{N_h}{N}. \]
If \(\hat{\theta}_h\) is the estimate within stratum \(h\), the combined estimator is:\[ \boxed{ \hat{\theta}_{\text{strat}} = \sum_{h=1}^{H} W_h\hat{\theta}_h } \]
For a population proportion:\[ \hat{p}_{\text{strat}} = \sum_{h=1}^{H} \frac{N_h}{N}\hat{p}_h. \]
The strata should not simply be averaged unless they contain equal population proportions or the design specifically justifies equal weighting.
Stratified Sampling Example
Suppose the voter population is:
| Stratum | Population share | Sample size | Estimated support |
|---|---|---|---|
| Urban | 40% | 400 | 60% |
| Suburban | 35% | 350 | 50% |
| Rural | 25% | 250 | 35% |
The combined estimate is:\[ \hat{p} = (0.40)(0.60) + (0.35)(0.50) + (0.25)(0.35). \]
Therefore:\[ \hat{p} = 0.24+0.175+0.0875 = 0.5025. \]
The estimated support is:\[ 50.25\%. \]
This weighted calculation preserves each stratum’s actual population share.
Proportional and Disproportional Allocation
Proportional allocation
The sample fraction is the same in every stratum:\[ n_h = n\frac{N_h}{N}. \]
A stratum containing 25% of the population receives approximately 25% of the sample.
This design is simple and often produces self-weighting observations.
Disproportional allocation
Some strata are intentionally oversampled.
This may be useful when:
- a stratum is small but analytically important;
- reliable estimates are needed for every stratum;
- data collection costs differ;
- variability differs substantially among strata.
Oversampling does not mean counting every sampled person equally in the final population estimate. Sampling weights must correct for unequal inclusion probabilities.
Why Stratification Can Improve Precision
Stratification can improve precision when:
- members within each stratum are relatively similar;
- the strata differ meaningfully from one another;
- the stratifying variable is related to the outcome.
The total variability can be viewed as containing:
- variability within strata;
- variability between strata.
If strata capture important differences, estimates can be constructed from more homogeneous groups, reducing sampling variance.
However, stratification is not automatically more precise. Poorly chosen strata, incorrect weights, nonresponse, or operational errors can eliminate its advantages.
Sampling Weights
If person \(i\) has inclusion probability \(\pi_i\), a basic design weight is:\[ w_i=\frac{1}{\pi_i}. \]
A person selected with probability 0.05 represents approximately:\[ \frac{1}{0.05}=20 \]
population members.
A person selected with probability 0.20 represents approximately:\[ \frac{1}{0.20}=5 \]
population members.
Weighted estimators account for unequal sampling probabilities:\[ \hat{\mu}_w = \frac{\sum_{i\in S}w_i y_i} {\sum_{i\in S}w_i}. \]
Practical survey weights may also include adjustments for:
- nonresponse;
- frame imperfections;
- known population totals;
- oversampling;
- multiple selection pathways.
Weighting can reduce certain distortions, but it cannot reconstruct information about completely unobserved groups without credible assumptions and auxiliary data.
Other Probability Sampling Designs
Cluster sampling
The population is divided into natural groups, such as schools, neighborhoods, or households. Some clusters are randomly selected, and observations are collected within them.
Cluster sampling can reduce cost but may decrease precision when members of the same cluster are similar.
Multistage sampling
Sampling occurs in stages. For example:
- select counties;
- select neighborhoods within counties;
- select households within neighborhoods;
- select one adult within each household.
This is common when a complete national list of individuals is unavailable.
Systematic sampling
Choose a random starting point and then select every \(k\)th member from an ordered list.
This can be efficient, but hidden periodic patterns in the list can cause problems.
These designs require analysis methods that respect their actual sampling structures.
Margin of Error Does Not Include Every Error
A reported margin of error usually describes uncertainty from random sampling under a specified design.
It generally does not automatically account for:
- undercoverage;
- nonresponse bias;
- voluntary response bias;
- misleading wording;
- inaccurate answers;
- data-processing mistakes;
- an incorrect model of likely participation.
A survey with a small reported margin of error can still be seriously biased.
Sampling variability is only one part of total survey error.
Evaluating a Survey
Before trusting a survey result, ask:
Population
- Who is the target population?
- Are registered, eligible, and likely voters being distinguished?
Sampling frame
- How were people reached?
- Who could not be reached through that method?
Selection
- Was a probability-based design used?
- Were inclusion probabilities known?
- Were any groups oversampled?
Nonresponse
- How many selected people responded?
- Did response rates differ across groups?
- How were missing responses handled?
Measurement
- How were questions worded?
- In what order were they asked?
- Was the survey conducted by phone, online, mail, or in person?
Weighting
- Were estimates weighted?
- Which population benchmarks were used?
- Were design effects reported?
Timing
- When was the survey conducted?
- Could an event during that period have influenced responses?
Common Misconceptions
“A sample of 1,000 is always representative”
Representativeness depends more on the selection process than on sample size.
“Random means haphazard”
Probability sampling uses a documented random mechanism, not informal or arbitrary selection.
“Every person must respond”
Complete response is ideal but rarely achieved. The important issue is whether nonresponse creates systematic distortion and how it is addressed.
“Weighting fixes every bias”
Weights can adjust for measured differences and known selection probabilities. They cannot reliably correct for every omitted or unmeasured group.
“Online reviews show typical customer experience”
They describe the people who chose to post. Those people may differ systematically from the complete customer population.
“Simple random sampling is always the best design”
It is conceptually clean, but stratified, cluster, multistage, or mixed-frame designs may be more precise or practical.
Key Takeaway
Reliable inference begins with how observations are selected. Convenience samples, incomplete coverage, nonresponse, voluntary participation, and response errors can systematically distort results, regardless of sample size. A simple random sample gives every possible sample of a fixed size an equal probability of selection. Stratified sampling draws random samples within meaningful population groups and combines them using population weights, often improving subgroup representation and statistical precision. Random selection reduces bias only when the sampling frame, response process, measurement method, and analysis are also handled carefully.
