Updated for the 2026-2027 CFA® Level I curriculum.
Any estimate based on a sample carries some uncertainty. A properly selected sample may differ from the full population through ordinary random variation, while problems in selection or data collection can push the estimate consistently in one direction.
For CFA Level I, you should understand how sampling error differs from standard error and bias. You should also be able to explain why a large sample can still produce a misleading investment conclusion.
Quick Answer
Sampling error is the difference between a statistic calculated from one sample and the corresponding population parameter. Standard error estimates how much that statistic would vary across repeated samples. Sampling bias comes from a selection process that systematically overrepresents or excludes certain observations. Larger random samples usually reduce standard error, while selection-related bias can remain.
Key Takeaways About Sampling Error and Bias
Sampling error occurs because an analyst observes a sample rather than the full population.
The realized sampling error can be positive or negative.
Standard error measures the expected spread of a statistic across repeated samples.
Larger properly selected samples generally have lower standard errors.
Sampling bias pushes an estimate systematically away from the population value.
A larger sample drawn through the same biased process may remain inaccurate.
Small standard error indicates precision, not necessarily accuracy.
Reliable estimates require both reasonable precision and a representative sampling process.
What You Need to Know for CFA Level I
For CFA Level I, focus on:
Defining sampling error.
Distinguishing realized sampling error from standard error.
Explaining how sample size affects standard error.
Separating random sampling variation from systematic bias.
Identifying common types of sampling bias.
Recognizing non-sampling errors in data collection and processing.
Explaining why a precise estimate can still be inaccurate.
What Is Sampling Error?
Sampling error is the difference between a statistic calculated from a sample and the corresponding parameter for the full population.
The general relationship is:
For a sample mean:
Where:
= sample mean
= population mean
A positive sampling error means the sample statistic is above the population parameter. A negative sampling error means it is below the population parameter.
Suppose the true average annual return for a population of funds is 7.0%, while a random sample produces an average of 8.2%. The realized sampling error is:
The sample mean is 1.2 percentage points above the population mean.
In practice, analysts usually do not know the true population parameter. They therefore cannot observe the exact sampling error directly. Standard error helps them estimate how much the statistic is likely to vary across repeated samples.
Sampling Error vs Standard Error
Sampling error and standard error describe related ideas, but they refer to different quantities.
Measure | Meaning |
|---|---|
Sampling error | The actual difference between one sample statistic and the population parameter |
Standard error | The standard deviation of the statistic across repeated samples |
What it describes | The result from one particular sample |
Can it usually be observed? | Only when the population parameter is known |
When the population standard deviation is known, the standard error of the sample mean is:
When the population standard deviation is unknown, the sample standard deviation is commonly used:
Where:
= standard error of the sample mean
= estimated standard error of the sample mean
= population standard deviation
= sample standard deviation
= sample size
A lower standard error means sample means tend to cluster more closely around their expected value. This indicates greater precision.
Precision alone does not confirm that the sampling process represents the target population properly.
What Is Sampling Bias?
Sampling bias occurs when the selection process systematically makes certain population members more or less likely to appear in the sample.
The resulting sample may consistently overstate or understate the population characteristic being studied.
At the estimator level, bias can be expressed as:
Where:
= estimator calculated from a sample
= expected value of the estimator across repeated samples
= true population parameter
A positive bias means the estimator tends to be too high. A negative bias means it tends to be too low.
Sampling bias often reflects a problem with the sampling frame, participation process, or analyst selection criteria. Repeating the same flawed process with more observations can produce a very precise estimate of the wrong population value.
Common Types of Sampling Bias
Selection Bias
Selection bias occurs when the process used to choose observations systematically favors certain members of the population.
Investment example: An analyst studies only companies with complete and easily available financial records. Distressed firms with poor reporting may be underrepresented.
Undercoverage
Undercoverage occurs when the sampling frame excludes part of the target population.
Investment example: A database containing only listed companies cannot represent a population that also includes private companies.
Random selection within the incomplete database will not restore the missing firms.
Survivorship Bias
Survivorship bias occurs when the sample contains entities that remained active while excluding entities that failed, closed, merged, or disappeared.
Investment example: Measuring historical mutual fund performance using only funds that still operate may overstate the average return because poorly performing funds are more likely to have closed.
Nonresponse Bias
Nonresponse bias occurs when selected participants do not respond and their characteristics differ meaningfully from those who do respond.
Investment example: Investors who experienced very poor service may be less willing to complete a satisfaction survey, causing the average response to appear more positive.
Self-Selection Bias
Self-selection bias occurs when people decide for themselves whether to enter the sample, and that decision is related to the subject being measured.
Investment example: A voluntary survey about active trading may attract investors who trade more often and feel more strongly about the topic than the typical investor.
Sampling Error vs Non-Sampling Error
Sampling error arises because an analyst studies only part of the population. Non-sampling errors come from problems in coverage, collection, recording, processing, or analysis.
Error Type | Description | Investment Example |
|---|---|---|
Sampling error | Random difference between a sample statistic and population parameter | A random sample mean differs from the true market mean |
Coverage error | Part of the population is missing from the sampling frame | A database excludes delisted companies |
Measurement error | A value is recorded or measured incorrectly | Returns are entered using the wrong currency |
Processing error | Data are coded, merged, or calculated incorrectly | Duplicate observations remain after datasets are combined |
Nonresponse error | Selected units do not provide information | Certain investor groups ignore a survey |
Model error | The analysis omits or misstates an important relationship | A risk model leaves out a relevant market factor |
Observing the full population removes ordinary sampling error, but measurement, processing, and model errors can still affect the conclusion.
How Does Sample Size Affect Sampling Error and Bias?
For a properly selected random sample, standard error decreases as sample size increases:
A larger sample therefore improves precision by reducing random sampling variability.
The square-root relationship means that sample size must increase substantially to produce a large reduction in standard error. For example, quadrupling the sample size cuts standard error in half.
Selection-related bias follows a different pattern. Suppose an online poll receives 100,000 voluntary responses but excludes people without reliable internet access. The large number of responses may create a small standard error within that group, while undercoverage and self-selection continue to affect the estimate.
The result may be highly precise for the population that participated and inaccurate for the broader population the analyst intended to study.
Precision vs Accuracy
Precision describes how closely repeated estimates cluster together. Accuracy describes how close an estimate is to the true population value.
Result | Precision | Accuracy |
|---|---|---|
Small random error and little bias | High | High |
Small random error but substantial bias | High | Low |
Large random error but little systematic bias | Low | Estimates may average to the correct value |
Large random error and substantial bias | Low | Low |
A small standard error supports precision. A representative sampling process and reliable data are also needed for accuracy.
Worked Investment Example
A population of investment funds has an actual mean annual return of 7.0%. An analyst selects a valid random sample and calculates a sample mean of 8.2%.
The realized sampling error is:
The estimate is 1.2 percentage points above the population mean. Another properly selected sample might produce a lower value, since random sampling error can move in either direction.
Now assume the analyst includes only funds that remained active throughout the measurement period. The calculated mean may also contain survivorship bias because failed or closed funds are missing.
Adding more surviving funds can reduce standard error within the survivor sample. The broader estimate may remain overstated because the missing funds are still excluded.
This example shows how an estimate can become more precise without becoming more accurate.
How Can Analysts Reduce Sampling Problems?
Analysts can improve the quality of sample-based estimates by:
Using a complete and relevant sampling frame.
Applying probability sampling when practical.
Using stratified sampling when important subgroups need representation.
Following up with nonrespondents.
Checking for failed, closed, merged, or delisted entities.
Reviewing datasets for missing values, duplicates, and recording errors.
Documenting selection and data-processing rules.
Reporting limitations alongside the estimate.
These steps reduce avoidable problems and make the remaining uncertainty easier to interpret.
Common Exam Traps
Common mistakes include:
Treating standard error as the realized error in one sample.
Assuming a large sample must be representative.
Saying random sampling eliminates all sampling error.
Treating small standard error as proof of accuracy.
Confusing survivorship bias with ordinary random variation.
Assuming sample size corrects undercoverage or self-selection.
Forgetting that sampling error can be positive or negative.
Assuming resampling can repair an unrepresentative original dataset.
Overlooking non-sampling errors when the full population is observed.
Practice Question
An analyst estimates average hedge fund returns using only funds that remain active at the end of the measurement period. The sample contains several thousand funds and has a small standard error.
Which statement is most accurate?
The large sample eliminates survivorship bias.
The estimate may be precise but systematically overstated.
The small standard error confirms that the sample represents all hedge funds.
Correct Answer: B
Excluding funds that closed or failed creates survivorship bias when those funds had systematically different returns. The large sample may produce a small standard error among the surviving funds, but it does not restore the missing observations.
The estimate can therefore be precise while remaining systematically too high.
Option A assumes sample size removes systematic bias.
Option C treats standard error as evidence of representativeness. Standard error measures sampling precision.
Continue Your CFA Level I Prep With KeyPoint
Use structured lessons, practice questions, mock exams, and progress tracking to focus on the time you have left
FAQs About Sampling Error and Bias
What Is Sampling Error in Simple Terms?
Sampling error is the difference between a value calculated from one sample and the corresponding value for the full population.
It occurs because different valid samples can contain different observations and produce different results.
What Is the Difference Between Sampling Error and Sampling Bias?
Sampling error reflects random variation between valid samples. Sampling bias comes from a selection process that systematically overrepresents or excludes certain observations.
Sampling error can move an estimate above or below the population value. Sampling bias tends to push estimates consistently in a particular direction.
Does Increasing Sample Size Reduce Sampling Error?
A larger properly selected random sample generally lowers standard error and improves precision.
The reduction follows a square-root relationship. Quadrupling the sample size typically cuts the standard error in half.
Can Increasing Sample Size Remove Sampling Bias?
Increasing the size of a sample drawn through the same biased process does not restore excluded or underrepresented population members.
Reducing sampling bias requires improving the sampling frame, selection method, response process, or data quality.
What Is the Difference Between Sampling Error and Non-Sampling Error?
Sampling error occurs because only part of the population is observed. Non-sampling error comes from problems such as missing coverage, incorrect measurements, nonresponse, processing mistakes, or unsuitable models.
Non-sampling errors can affect both samples and full-population datasets.