Sampling error
Error from estimating population parameters using a sample.
Last updated
Bob K · CC0
Sampling error is a concept in statistics that arises when the characteristics of a population are estimated from a subset, or sample, of that population. It is defined as the difference between a sample statistic (such as a mean or quartile) and the actual but unknown population parameter. Because sampling is almost always done to estimate unknown parameters, exact measurement of sampling errors is usually not possible, though they can often be estimated using methods like bootstrapping or assumptions about the population distribution.
Effective sampling
The fundamental cause of sampling error is that the sample does not include every member of the population. For instance, measuring the height of a thousand individuals from a population of one million will typically yield an average height that differs from the true average of the entire million. This discrepancy is the sampling error.
Bootstrapping and standard error
While the exact error cannot be known (since the true population parameter is unknown), its likely magnitude can be estimated. One general estimation method is bootstrapping, which involves comparing many samples or splitting a larger sample into smaller overlapping ones to gauge the spread of sample statistics, thereby estimating the standard error. Alternatively, specific methods may incorporate assumptions or guesses about the true population distribution and its parameters.
The size of the sampling error is influenced by sample size. A very small sample, such as measuring only two or three individuals, will produce wildly varying results each time.
Sample size determination
Increasing the sample size generally reduces the likely error. However, the cost of obtaining a larger sample may be prohibitive. Consequently, sample size determination methods are used to balance the predicted accuracy of an estimator against the predicted cost of taking a larger sample.
It is crucial to distinguish sampling error from sampling bias. A truly random sample requires selecting individuals from a population with equivalent probability, avoiding systematic bias. Failing to do so—for example, measuring the average height of the entire human population using a sample from only one country—can dramatically increase the error in a systematic way.
Quick Facts
- Field
- Statistics
- Known for
- Difference between sample statistic and population parameter
- Related concepts
- Sampling bias
- sample size determination
- bootstrapping
- standard error
Facts from the source article.
Lore & Background
Sampling error is the discrepancy that arises when the statistical properties of a population are estimated from a subset, or sample, rather than from every member of the population. The statistics calculated from a sample—such as means or quartiles, known as estimators—generally differ from the true population values, called parameters. This difference is the sampling error.
For instance, if the average height of one thousand individuals is measured from a population of one million, that sample average will typically not equal the average height of all one million people. Because sampling is almost always performed to estimate unknown population parameters, the exact measurement of sampling error is usually impossible. However, its magnitude can often be estimated through general methods like bootstrapping, or through specific methods that incorporate assumptions about the true population distribution. A truly random sample, where each individual is selected with equal probability, is essential to avoid sampling bias, which can systematically increase error.
Even with a perfect unbiased sample, sampling error persists due to random statistical variation; measuring only two or three individuals, for example, will produce wildly varying results. The likely size of this error can generally be reduced by increasing the sample size, though the cost of doing so may be prohibitive. In genetics, the term "sampling error" has been used in a related but distinct sense to describe effects like the bottleneck or founder effect, where a dramatic reduction in population size (e.g., from natural disasters or migrations) creates a smaller population that may not fairly represent the original, contributing to genetic drift.
The Fundamental Gap Between Sample and Population
Sampling error sits at the very heart of inferential statistics, emerging from a simple but unavoidable reality: we almost never have access to an entire population. When researchers select a subset of individuals and compute statistics such as means or quartiles, those figures—commonly called estimators—will almost certainly diverge from the true values that describe the full population, known as parameters. The numerical gap between the two is what statisticians label the sampling error.
A concrete illustration makes this tangible: if you record the height of one thousand people drawn from a national population of one million, the average you calculate will typically not match the average height of all one million residents. Because the parameters we seek are, by definition, unknown, pinning down the exact magnitude of the sampling error is generally impossible. Nevertheless, statisticians have developed routes to approximate it, ranging from assumption-free resampling techniques to methods that layer in educated guesses about the underlying population distribution.
Bias, Randomness, and the Inevitable Residual
Even when a researcher goes to great lengths to construct a genuinely random sample—selecting every individual with an equal probability and eliminating systematic favoritism—a residual sampling error persists. This irreducible component stems from the sheer act of observing a finite subset rather than the whole. Picture measuring just two or three people's heights and averaging them; the result would swing dramatically from one draw to the next. The remedy is straightforward in principle: enlarge the sample, and the likely magnitude of the error shrinks.
However, the far more dangerous threat is sampling bias. If the selection process inadvertently favors certain subgroups—say, drawing a global height sample exclusively from a single nation—the resulting estimate can be skewed in a systematic direction, producing a large over- or under-estimation. In practice, many variables such as country of origin, age, and gender can each introduce distortion, and ensuring none of them creeps into the selection mechanism is a genuinely difficult task.
Quantifying Uncertainty and Managing Cost
Because any single sample statistic—be it a mean, a proportion, or a quartile—will fluctuate from one draw to the next, statisticians need a way to gauge how much that fluctuation matters. One powerful approach is bootstrapping: by repeatedly resampling from the data at hand, or by partitioning a larger sample into overlapping sub-samples, the spread of the resulting statistics provides a practical estimate of the standard error. More targeted methods exist as well, though they typically require assumptions about the true population distribution. A separate but equally critical concern is cost.
In real-world settings, enlarging a sample to squeeze out more error can be prohibitively expensive in time, money, or logistics. Fortunately, the anticipated sampling error can often be projected in advance as a function of sample size. This allows practitioners to apply formal sample-size determination techniques that balance the predicted precision of an estimator against the predicted expense of collecting additional observations, arriving at a defensible compromise.
A Borrowed Term in Genetics
Outside the strict statistical framework, the phrase 'sampling error' has been adopted in genetics to describe a phenomenon that, while conceptually analogous, is not an error in the conventional sense. When a natural disaster or a migration event drastically shrinks a population—what biologists call a bottleneck effect or a founder effect—the surviving or migrating group is a small, random subset of the original gene pool. That subset may or may not faithfully mirror the allele frequencies of the population it came from.
The consequence is genetic drift: certain genetic variants become more prevalent while others fade, simply because of chance in who survived or who moved. Researchers in this field label the process 'sampling error' because the reduced population is, in a sense, a random draw from a larger genetic reservoir. Yet unlike the statistical concept, there is no true parameter being estimated and no deliberate measurement being made; the term captures the stochastic reduction in genetic diversity rather than a gap between an estimator and a parameter.
Reader's Guide
Sampling error is a fundamental concept in statistics, representing the inevitable discrepancy between a sample statistic and the true population parameter it estimates. Its significance lies in the fact that most statistical inference relies on samples, not full populations. The article also distinguishes sampling error from sampling bias, which arises from non-random selection and can systematically increase error.
Sample size determination methods weigh predicted accuracy against cost, as larger samples reduce error but may be expensive. In genetics, the term has been used in a related but different sense, referring to population bottlenecks or founder effects that cause genetic drift, though this is not an 'error' in the statistical sense. Understanding sampling error is crucial for interpreting margins of error and propagation of uncertainty in research.
More in Probability & Statistics
Sources
Compiled from Wikipedia and the sources listed below. Text from Wikipedia is available under CC BY-SA 4.0; this entry is adapted from it.
- Wikipedia: Sampling error (CC BY-SA 4.0).
- Word definitions: the Codexery glossary, each quoted from its Wikipedia article.
Spotted an error? Know more?
Reader corrections go straight into our review queue. Suggest an edit · How this site is sourced