Updated for the 2026-2027 CFA® Level I curriculum.
A hypothesis test can produce a correct decision or one of two errors. You may reject a null hypothesis that is actually true, or fail to reject a null hypothesis that is actually false.
For CFA Level I, focus on connecting each outcome with , , and statistical power. These relationships help explain what the significance level controls and why reducing one type of error can make the other more likely.
Quick Answer
A Type I error occurs when you reject a true null hypothesis. A Type II error occurs when you fail to reject a false null hypothesis. The significance level controls the probability of a Type I error, while represents the probability of a Type II error. The power of a test is , which is the probability of correctly rejecting a false null hypothesis.
Key Takeaways About Type I and Type II Errors, Power, and Significance
A Type I error means rejecting a null hypothesis that is true.
A Type II error means failing to reject a null hypothesis that is false.
A Type I error is often described as a false positive.
A Type II error is often described as a false negative.
The significance level represents the tolerated probability of a Type I error.
represents the probability of a Type II error for a specified alternative or effect size.
Statistical power equals .
Lowering generally raises and reduces power when other factors remain constant.
Increasing the sample size generally raises power without changing the selected significance level.
The more costly error depends on the real consequences of the decision.
What You Need to Know for CFA Level I
For CFA Level I, you should be able to:
Identify Type I and Type II errors from a written scenario.
Connect with the probability of a Type I error.
Connect with the probability of a Type II error.
Calculate statistical power from .
Read all four outcomes in a hypothesis-testing decision matrix.
Explain how a change in affects and power.
Explain how sample size, variability, and effect size affect power.
Evaluate which error may be more costly in a given investment setting.
What Is a Type I Error?
A Type I error occurs when the test rejects a null hypothesis that is actually true. The analyst concludes that an effect exists even though it does not.
The probability of a Type I error is represented by , the significance level selected before the evidence is evaluated.
For example, suppose an analyst tests whether a trading signal produces positive excess returns.
Assume the signal has no genuine ability to generate excess returns. Sampling variation could still produce a test statistic large enough to fall inside the rejection region.
If the analyst rejects and recommends the signal, the decision is a Type I error. The firm may then commit capital to a strategy that has no real investment edge.
Type I Error in Plain Language
True state: The null hypothesis is true.
Test decision: Reject the null hypothesis.
Result: False positive.
Probability: .
What Is a Type II Error?
A Type II error occurs when the test fails to reject a null hypothesis that is actually false. A genuine effect exists, but the test does not detect it.
The probability of a Type II error is represented by .
Unlike , beta is not normally selected as a direct decision threshold. It depends on several features of the test, including:
The significance level
The sample size
The variability of the observations
The true effect size
The form and direction of the hypothesis test
Suppose a second trading signal genuinely produces positive excess returns. The sample may still be too small or volatile to generate a test statistic inside the rejection region.
If the analyst fails to reject and abandons the signal, the decision is a Type II error. The firm misses a genuine source of excess return.
Type II Error in Plain Language
True state: The null hypothesis is false.
Test decision: Fail to reject the null hypothesis.
Result: False negative.
Probability: .
Type I vs Type II Errors
The possible outcomes can be organized according to the true state of the null hypothesis and the decision produced by the test.
Test Decision | Is True | Is False |
|---|---|---|
Fail to reject | Correct decision, probability | Type II error, probability |
Reject | Type I error, probability | Correct decision, power |
Read the table by starting with the true state:
When is true, you either fail to reject it correctly or make a Type I error.
When is false, you either fail to reject it incorrectly or reject it correctly.
The correct rejection of a false null hypothesis is what statistical power measures.
What Is the Significance Level?
The significance level is the probability threshold used to control Type I error. It is selected before the sample evidence is evaluated and helps determine the rejection region.
Common significance levels include:
A lower significance level requires stronger evidence before the null hypothesis can be rejected. The critical value moves farther into the tail, which reduces the probability of rejecting a true null hypothesis.
For example, changing from 0.05 to 0.01 makes a Type I error less likely. It also makes rejection more difficult when the null hypothesis is false, which generally increases and lowers power when the sample size, variability, and effect size remain unchanged.
What Is the Power of a Test?
The power of a hypothesis test is the probability of correctly rejecting a false null hypothesis.
Where:
= probability of a Type II error
= probability of correctly rejecting a false null hypothesis
Suppose a test has .
The test has power of 0.72, or 72%, for the specified alternative or effect size. Under repeated samples generated under those same conditions, the test would correctly detect that effect about 72% of the time.
Power should always be discussed in relation to a particular effect. A large departure from the null hypothesis is easier to detect than a very small departure.
What Affects the Power of a Hypothesis Test?
Several factors influence statistical power.
Significance Level
Increasing enlarges the rejection region. This makes it easier to reject both true and false null hypotheses.
Holding the other factors constant:
Higher increases power.
Higher also increases the probability of a Type I error.
Lower reduces power and increases .
Sample Size
A larger sample generally reduces the standard error. The test can then distinguish a genuine effect from sampling variation more clearly.
Holding other factors constant:
A larger sample increases power.
A larger sample reduces .
The selected significance level does not need to change.
Variability
Lower variability makes the sample estimate more precise. A more precise estimate produces a smaller standard error and generally increases power.
Higher variability makes genuine differences harder to separate from ordinary sampling noise.
Effect Size
The effect size is the true difference between the population parameter and the value stated in the null hypothesis.
Larger effects are easier to detect. A strategy generating 3% of annual excess return will generally be easier to distinguish from zero than a strategy generating 0.15%.
How Do Alpha, Beta, and Power Trade Off?
When the sample size, variability, and true effect size remain constant, reducing generally increases .
A lower significance level moves the critical value farther into the tail. This reduces false rejections when is true. It also causes some genuine effects to fall outside the new rejection region.
The relationship can be summarized as follows:
Change | Type I Error Probability | Type II Error Probability | Power |
|---|---|---|---|
Lower | Decreases | Increases | Decreases |
Higher | Increases | Decreases | Increases |
Larger sample size | Unchanged if is fixed | Decreases | Increases |
Lower variability | Unchanged if is fixed | Decreases | Increases |
Larger true effect | Unchanged if is fixed | Decreases | Increases |
Lowering makes the test more conservative about Type I errors. It does not improve every part of the test at once.
Increasing the sample size offers a more useful improvement because it can raise power while leaving the chosen Type I error threshold unchanged.
Why Do the Costs of Type I and Type II Errors Matter?
The consequences of the two errors can differ significantly. The analyst should consider what happens after each possible decision.
Suppose a compliance team tests whether a portfolio manager has exceeded a firm risk limit.
A Type I error would incorrectly flag a manager who is within the limit.
A Type II error would fail to detect a manager who has genuinely exceeded the limit.
The Type I error could lead to an unnecessary investigation and damaged trust. The Type II error could expose the firm to losses, regulatory action, or reputational harm.
In this setting, the Type II error may carry the greater cost. The firm may therefore prefer a test with more power, even if that means accepting a somewhat higher Type I error threshold.
Now consider an asset manager testing whether a new strategy produces positive alpha.
A Type I error may lead the firm to fund a strategy with no genuine edge.
A Type II error may cause the firm to reject a potentially useful strategy.
If launching an ineffective strategy is especially costly, the firm may place more weight on controlling Type I error.
The statistical relationships remain the same. The economic consequences determine which error deserves greater attention.
Worked Example: Testing Whether a Strategy Adds Value
An analyst tests whether a systematic momentum strategy generates a positive mean monthly excess return.
Step 1: State the Hypotheses
The analyst selects:
Significance level:
Type II error probability for the specified effect:
Step 2: Calculate the Power
The test has 72% power to detect the specified positive excess return under the assumed conditions.
Step 3: Interpret the Possible Outcomes
Actual State | Test Decision | Interpretation |
|---|---|---|
Strategy adds no value | Fail to reject | Correct decision |
Strategy adds no value | Reject | Type I error |
Strategy adds value | Fail to reject | Type II error |
Strategy adds value | Reject | Correct decision, represented by power |
A Type I error would lead the analyst to recommend a strategy that has no genuine edge.
A Type II error would lead the analyst to reject a strategy that does generate positive excess returns.
Step 4: Consider Changes to the Test
Lowering from 0.05 to 0.01 would reduce the probability of funding an ineffective strategy. Holding the other factors constant, the stricter threshold would also increase and reduce power.
Increasing the sample size would generally improve the test’s ability to detect the specified effect without requiring a higher significance level.
Common Exam Traps
Reversing the error definitions. Type I means rejecting a true null hypothesis. Type II means failing to reject a false null hypothesis.
Confusing with . Alpha controls Type I error, while beta represents Type II error.
Calculating power as . Power equals .
Forgetting that beta depends on the alternative. The probability of a Type II error changes with the true effect size.
Assuming a lower significance level improves every outcome. Lower reduces Type I error but generally raises Type II error when other factors are fixed.
Treating failure to reject as proof that is true. The evidence may simply be too weak or the test may have low power.
Ignoring the true state of the null hypothesis. Whether an error occurred depends on both the test decision and the actual state.
Discussing power without holding other factors constant. Sample size, variability, significance level, and effect size interact.
Practice Question
An analyst conducts a hypothesis test with a significance level of 0.05. For the effect size the analyst wants to detect, the probability of a Type II error is 0.35.
Which of the following correctly states the test’s power and its interpretation?
0.60, the probability of failing to reject a false null hypothesis
0.65, the probability of correctly rejecting a false null hypothesis
0.95, the probability of correctly failing to reject a true null hypothesis
Correct Answer: B
Explanation
Power equals one minus the probability of a Type II error.
The test has 65% power to correctly reject the null hypothesis when it is false and the specified effect exists.
Option A gives an incorrect calculation and describes a Type II error rather than power.
Option C calculates . This is the probability of correctly failing to reject a true null hypothesis, not the probability of detecting a false one.
Continue Your CFA Level I Prep With KeyPoint
Use structured lessons, practice questions, mock exams, and progress tracking to focus on the time you have left
FAQs About Type I and Type II Errors, Power, and Significance
What Is the Difference Between a Type I and Type II Error?
A Type I error occurs when a true null hypothesis is rejected. A Type II error occurs when a false null hypothesis is not rejected.
Type I error is often described as a false positive, while Type II error is described as a false negative.
Is the Significance Level the Probability of a Type I Error?
The significance level sets the tolerated probability of rejecting a true null hypothesis under the test procedure.
A lower significance level makes rejection more difficult and reduces the probability of a Type I error.
How Do You Calculate the Power of a Test?
Power is calculated as:
A test with has power of 0.80, or 80%, for the specified alternative.
How Does Sample Size Affect Statistical Power?
Increasing the sample size generally reduces the standard error. This makes a genuine effect easier to distinguish from sampling variation and raises the power of the test.
The significance level can remain unchanged while power increases.
Which Is Worse, a Type I or Type II Error?
Neither error is always worse. The more serious error depends on the consequences of the decision.
Funding a strategy with no genuine edge may make a Type I error especially costly. Missing an actual compliance breach may make a Type II error more serious.