Selection and Nonresponse Bias
Selection bias arises when the way records enter a dataset is related to what is being measured, so the records describe a different group from the one the question is about. Nonresponse bias is the survey form of it: some of the people asked do not answer, and if they differ from those who do, the answers describe respondents rather than the population. Both are systematic; Statistics Canada notes that such errors accumulate and are not reduced by a larger sample.
Eight of twelve is not twelve
A fictional support team surveys its twelve agents; eight respond, a response rate of 8 / 12 ≈ 67%. Five of the eight say they lack time for billing questions. If the four who did not answer were the busiest agents, the survey understates the problem; if the most frustrated agents were the keenest to respond, it overstates it. The honest sentence is “five of eight respondents,” with the four nonrespondents reported, not “most agents.” The U.S. Census Bureau defines a response rate against the units that should have been interviewed and treats it as a direct indicator of data quality for the same reason.
Operational data has its own selection. A ticket log contains customers who wrote in and says nothing about those who gave up first; a voluntary review site contains the pleased and the angry; a study of employees who stayed cannot see the ones who left before the study began.
What to do about it
Name the population the question is about and the group the data actually contains, and state how they differ. Report response rates and who is absent. Where possible, compare respondents with nonrespondents on something known about both, such as team or tenure. Follow-up that identifies individuals may break a promise of anonymity, so the remedy is sometimes to report the limit rather than to chase the missing answers. Weighting and adjustment methods exist, but they rest on assumptions that must be stated; they do not turn respondents into the whole group. Survey sampling and the oversampling or subsampling used to balance machine-learning training sets are different procedures with different purposes, despite the shared word.
References: Statistics Canada: Non-sampling error, U.S. Census Bureau: Response rates definitions. Examples here are illustrative.
Discover more from Insightful Data Lab
Subscribe to get the latest posts sent to your email.
