Updated for the 2026-2027 CFA® Level I curriculum.
Hypothesis testing gives analysts a structured way to evaluate a claim about a population using sample evidence. The process starts with two competing hypotheses and ends with a decision based on the strength of the observed data.
For CFA Level I, you should understand how the hypotheses, test statistic, significance level, p-value, and decision rule fit together. You should also be able to interpret Type I and Type II errors, test power, and the practical meaning of the final result.
Quick Answer
A hypothesis test compares sample evidence with a claim stated in the null hypothesis. The analyst selects a significance level, calculates a test statistic, and evaluates the result using a critical value or p-value. The final decision is to reject or fail to reject the null hypothesis. A failure to reject means the evidence was insufficient at the chosen significance level.
Key Takeaways About Hypothesis Testing
The null hypothesis is the claim evaluated directly.
The alternative hypothesis represents the competing claim.
Equality usually appears in the null hypothesis.
The test statistic measures the sample result’s distance from the null value in standard-error units.
The significance level sets the acceptable probability of a Type I error.
A p-value below the significance level supports rejection of the null hypothesis.
Failing to reject the null leaves the claim unresolved rather than confirming it.
Test power measures the probability of correctly rejecting a false null hypothesis.
Statistical significance should be considered alongside economic significance.
What You Need to Know for CFA Level I
For CFA Level I, focus on:
Stating null and alternative hypotheses.
Identifying right-tailed, left-tailed, and two-tailed tests.
Explaining the level of significance.
Calculating or interpreting a test statistic.
Using critical values and p-values to make a decision.
Distinguishing Type I and Type II errors.
Explaining the power of a test.
Interpreting the result in the context of an investment problem.
Separating statistical significance from economic significance.
What Is Hypothesis Testing?
Hypothesis testing is a formal procedure for evaluating a population claim using information from a sample.
Investment questions that may be tested include:
Is a portfolio manager’s mean excess return greater than zero?
Does a population correlation coefficient differ from zero?
Are the mean returns of two strategies equal?
Is a regression coefficient statistically significant?
Are two categorical variables independent?
The procedure manages decision risk when the analyst has incomplete information. Every conclusion remains subject to sampling variation and the assumptions of the chosen test.
What Are the Null and Alternative Hypotheses?
The null hypothesis, written as , is the claim evaluated directly. It usually includes an equality.
Examples include:
The alternative hypothesis, written as or , describes the competing claim. The form of the alternative determines whether the test is right-tailed, left-tailed, or two-tailed.
Two-Tailed Alternative
A two-tailed test is appropriate when results above or below the null value both matter.
For example, an analyst may want to know whether a strategy’s mean return differs from 5%, regardless of direction.
Right-Tailed Alternative
A right-tailed test is appropriate when only values above the null value support the research claim.
For example, an analyst may test whether a manager’s mean excess return is greater than zero.
Left-Tailed Alternative
A left-tailed test is appropriate when only values below the null value support the research claim.
For example, an analyst may test whether a portfolio’s average loss falls below a specified threshold.
The direction should come from the research question. It should be selected before the sample results are reviewed.
What Is the Procedure for Hypothesis Testing?
A complete hypothesis test usually follows seven steps:
State the null and alternative hypotheses.
Choose the appropriate statistical test.
Select the significance level.
Collect the sample and calculate the test statistic.
Determine the critical value or p-value.
Reject or fail to reject the null hypothesis.
Interpret the result in the context of the investment question.
The correct test depends on the parameter being evaluated, the number and structure of the samples, the available population information, and the assumptions supported by the data.
How Does a Test Statistic Work?
A test statistic compares the observed sample result with the value stated in the null hypothesis.
Its general structure is:
Where:
= estimate calculated from the observed sample
= parameter value stated in the null hypothesis
= estimated sampling variability of the statistic
A large positive or negative test statistic means the sample result lies farther from the null value relative to its standard error.
Common test statistics include:
z-statistics
t-statistics
chi-square statistics
F-statistics
Each statistic follows a particular probability distribution. That distribution determines the appropriate critical value and p-value.
What Is the Level of Significance?
The level of significance, written as , is the chosen probability of making a Type I error.
A Type I error occurs when the analyst rejects a true null hypothesis.
Common significance levels are:
A lower value of requires stronger evidence before the null hypothesis can be rejected.
For a 5% two-tailed z-test:
2.5% of the rejection area lies in the left tail.
2.5% lies in the right tail.
The approximate critical values are 1.96 and +1.96.
The rejection rule is:
For a 5% right-tailed z-test, the approximate critical value is:
The rejection rule is:
The same 5% significance level produces different critical values because the rejection area is distributed differently across the tails.
How Does the P-Value Decision Rule Work?
The p-value measures how unusual the observed result would be if the null hypothesis and the test assumptions were correct.
It can also be interpreted as the smallest significance level at which the observed result would lead to rejection.
The decision rule is:
When the p-value is greater than or equal to the significance level:
A smaller p-value indicates stronger evidence against the null hypothesis.
The p-value should be interpreted carefully. It is a probability calculated under the null hypothesis. It does not represent:
The probability that the null hypothesis is true.
The probability that the alternative hypothesis is true.
The economic importance of the result.
The probability that the result occurred “by chance” in a general sense.
Type I Error, Type II Error, and Power
A hypothesis test can produce two correct decisions and two possible errors.
Actual Situation | Fail to Reject | Reject |
|---|---|---|
is true | Correct decision | Type I error |
is false | Type II error | Correct decision |
Type I Error
A Type I error occurs when the analyst rejects a true null hypothesis.
Its probability is:
Type II Error
A Type II error occurs when the analyst fails to reject a false null hypothesis.
Its probability is:
Power of a Test
Power is the probability of correctly rejecting a false null hypothesis.
Where:
= probability of a Type I error
= probability of a Type II error
= power of the test
Holding other factors constant, test power generally increases when:
The sample size increases.
The true parameter lies farther from the null value.
The variability of the data decreases.
The significance level increases.
Lowering makes rejection more difficult. With other factors held constant, this can increase and reduce power.
Worked Example: Testing a Population Mean
An analyst wants to test whether the mean monthly return of an investment strategy differs from 0.5%.
The hypotheses are:
This is a two-tailed test because returns above or below 0.5% would support the alternative.
Assume:
Population standard deviation = 1.6%
Sample size = 64
Sample mean = 0.8%
Significance level = 5%
Step 1: Calculate the Standard Error
Substituting the values:
Step 2: Calculate the Test Statistic
Substituting the values:
Where:
= observed sample mean
= population mean stated in the null hypothesis
= standard error of the sample mean
= calculated z-statistic
Step 3: Apply the Decision Rule
For a 5% two-tailed z-test, the critical values are:
The calculated statistic lies between the critical values:
The analyst therefore fails to reject the null hypothesis.
Step 4: Interpret the Result
At the 5% significance level, the sample does not provide enough evidence to conclude that the population mean monthly return differs from 0.5%.
The result leaves the null hypothesis unresolved. The observed difference may reflect sampling variation under the assumptions of the test.
Statistical Significance vs Economic Significance
Statistical significance indicates that the observed result is unlikely under the null hypothesis at the chosen significance level.
Economic significance asks whether the effect is large enough to matter in an investment decision.
A statistically significant return difference may have little practical value after considering:
Transaction costs.
Management fees.
Taxes.
Risk exposure.
Liquidity.
Implementation constraints.
An economically meaningful estimate may also fail to reach statistical significance when the sample is small or the data are highly variable.
Investment analysis should consider both the strength of the statistical evidence and the practical size of the effect.
Common Exam Traps
Common mistakes include:
Placing the equality in the alternative hypothesis.
Choosing the direction of the test after reviewing the sample result.
Using two-tailed critical values for a one-tailed test.
Saying “accept the null hypothesis” instead of “fail to reject the null hypothesis.”
Rejecting the null when the p-value exceeds .
Interpreting as the probability that the null hypothesis is true.
Treating the p-value as the probability that the null hypothesis is true.
Reversing Type I and Type II errors.
Confusing test power with the significance level.
Treating statistical significance as proof of economic value.
Practice Question
An analyst conducts a hypothesis test using a 5% significance level and obtains a p-value of 0.08. Which conclusion is most accurate?
Reject the null hypothesis because the p-value is positive.
Fail to reject the null hypothesis because the p-value exceeds 0.05.
Accept the null hypothesis as proven true.
Correct Answer: B
The p-value is:
The significance level is:
Because:
the sample evidence is insufficient to reject the null hypothesis at the 5% significance level.
Option A applies the wrong decision rule. A positive p-value does not determine the decision.
Option C overstates the result. Failing to reject the null means the evidence did not reach the chosen rejection threshold.
Continue Your CFA Level I Prep With KeyPoint
Use structured lessons, practice questions, mock exams, and progress tracking to focus on the time you have left
FAQs About Hypothesis Testing
What Is Hypothesis Testing in Simple Terms?
Hypothesis testing uses sample data to evaluate a claim about a population.
The analyst compares the observed result with the value stated in the null hypothesis and decides whether the evidence reaches a chosen rejection threshold.
What Does Fail to Reject the Null Hypothesis Mean?
Failing to reject the null means the sample evidence was not strong enough to reject it at the chosen significance level.
The conclusion leaves the null unresolved. A different sample, larger sample size, or lower variability may produce a different result.
What Is the Level of Significance in Hypothesis Testing?
The significance level is the chosen probability of rejecting a true null hypothesis.
It is written as , with common values of 10%, 5%, and 1%.
What Is the Power of a Test in Hypothesis Testing?
The power of a test is the probability of correctly rejecting a false null hypothesis.
It equals , where is the probability of a Type II error.
What Is the Difference Between a P-Value and a Significance Level?
The significance level is chosen before the test and sets the rejection threshold.
The p-value is calculated from the sample result. The analyst rejects the null hypothesis when the p-value is below the chosen significance level.