Updated for the 2026-2027 CFA® Level I curriculum.
Simple linear regression describes the average linear relationship between two variables. One variable is the outcome you want to explain, while the other is used to help explain or predict it.
For CFA Level I, you should be able to set up the model, identify the dependent and independent variables, interpret the intercept and slope, and calculate fitted values and residuals.
Quick Answer
Simple linear regression estimates the linear relationship between one dependent variable and one independent variable. The intercept gives the predicted value of the dependent variable when the independent variable equals zero. The slope gives the expected change in the dependent variable for a one-unit increase in the independent variable.
Key Takeaways About Simple Linear Regression
A simple linear regression contains one dependent variable and one independent variable.
The dependent variable, , is the outcome being explained or predicted.
The independent variable, , is used to explain changes in .
Population coefficients are written as and .
Sample estimates of those coefficients are written as and .
The intercept is the predicted value of when .
The slope is the expected change in for a one-unit increase in .
A fitted value comes from the estimated regression line.
A residual equals the observed value minus the fitted value.
Regression measures association and supports prediction. It does not establish causation on its own.
What You Need to Know for CFA Level I
For CFA Level I, focus on:
Identifying the dependent and independent variables from the research question.
Writing the population and estimated regression equations.
Distinguishing population parameters from sample estimates.
Interpreting the intercept and slope in the units of the problem.
Calculating a fitted value for a given value of .
Calculating and interpreting a residual.
Explaining how least squares determines the fitted regression line.
Distinguishing simple linear regression from correlation.
Recognizing why statistical association does not automatically mean causation.
The 2026 curriculum requires candidates to describe a simple linear regression model, explain how least squares estimates its coefficients, and interpret those coefficients.
What Is Simple Linear Regression?
Simple linear regression models the average value of a dependent variable as a linear function of one independent variable.
The word simple means the model has one independent variable. The word linear means the relationship is represented by a straight line and is linear in the model coefficients.
For example, an analyst might use simple linear regression to study:
How a portfolio’s return changes with the market return.
How company sales growth relates to industry sales growth.
How bond returns respond to changes in interest rates.
How a company’s valuation multiple relates to expected earnings growth.
The model converts the observed relationship into an equation. That equation can then be used to interpret the relationship and calculate predicted values.
Individual observations will usually sit above or below the fitted line. The line summarizes the average relationship rather than matching every point exactly.
What Are the Dependent and Independent Variables?
The dependent variable is the outcome the analyst wants to explain or predict. It is usually written as .
The independent variable is the variable used to explain changes in the dependent variable. It is usually written as .
Consider the question:
How does a portfolio’s return respond to changes in the market return?
In this case:
Portfolio return is the dependent variable, .
Market return is the independent variable, .
The research question determines the variable roles. Reversing them creates a different regression and answers a different question.
You may also see the following terms:
Standard Term | Alternative Terms |
|---|---|
Dependent variable | Explained variable, response variable, outcome variable |
Independent variable | Explanatory variable, predictor variable |
These labels refer to the roles the variables play in the model. They do not prove that the independent variable causes the dependent variable to change.
What Is the Simple Linear Regression Formula?
Population Regression Model
The population model describes the true relationship across the full population:
Where:
= dependent variable for observation
= independent variable for observation
= population intercept
= population slope coefficient
= population error term
The population coefficients and are unknown. An analyst uses sample data to estimate them.
Estimated Regression Equation
The fitted regression line is:
Where:
= fitted or predicted value of
= estimated intercept
= estimated slope coefficient
= observed value of the independent variable
The estimates and are calculated from the sample. They are used as estimates of the unknown population parameters and .
Residual Formula
The residual measures the difference between the observed and fitted values:
Where:
= residual for observation
= observed value
= fitted value
A positive residual means the observed value is above the fitted regression line. A negative residual means the observed value is below it.
What Is the Difference Between the Population and Sample Regression?
The population model describes the true relationship, while the sample regression estimates that relationship using observed data.
Component | Population Model | Estimated Sample Regression |
|---|---|---|
Equation | ||
Intercept | ||
Slope | ||
Unexplained component | Error term | Residual |
Observable? | Parameters and errors are unknown | Estimates and residuals are calculated from data |
The error term and residual are closely related, but they are not interchangeable.
The error term measures the distance from an observation to the unknown population regression line. The residual measures the distance from an observation to the estimated sample regression line.
How Do You Interpret the Regression Intercept?
The intercept, , is the fitted value of when .
Suppose the estimated regression is:
The intercept is 1.20. When , the predicted value of is:
The intercept should always be interpreted in the units of the dependent variable.
Its practical usefulness depends on the context. When is realistic and falls within or near the observed data range, the intercept may have a useful economic interpretation.
When zero is far outside the observed range, the intercept mainly serves as the mathematical point where the regression line crosses the vertical axis. Using it as a practical forecast would require extrapolation beyond the available evidence.
How Do You Interpret the Slope of a Regression Line?
The slope, , is the expected change in for a one-unit increase in .
For the equation:
The slope is 1.35.
If both variables are measured in percentage points, the interpretation is:
A one-percentage-point increase in is associated with an expected 1.35-percentage-point increase in .
The slope carries units:
Always include those units in your interpretation.
Positive Slope
A positive slope indicates that the fitted value of increases as increases.
Negative Slope
A negative slope indicates that the fitted value of decreases as increases.
Zero Slope
A slope of zero indicates that the model shows no linear change in the fitted value of as changes.
An estimated slope close to zero suggests a weak estimated linear relationship, but statistical testing is needed to determine whether the population slope differs significantly from zero.
How Does Least Squares Fit the Regression Line?
Ordinary least squares chooses the intercept and slope that minimize the sum of squared residuals.
The least-squares objective is:
Because:
The same objective can be written as:
Squaring the residuals prevents positive and negative residuals from cancelling each other. It also gives greater weight to observations that sit farther from the fitted line.
The detailed coefficient calculations belong in the Least Squares Criterion study note. At this stage, remember that least squares chooses the line with the smallest possible total squared prediction error within the sample.
What Are Fitted Values, Errors, and Residuals?
Fitted Value
A fitted value is the value predicted by the estimated regression line for a particular .
Error Term
The error term is the difference between the observed value and the value implied by the unknown population regression line.
Because the true population line is unknown, the error term cannot be directly observed.
Residual
The residual is the difference between the observed value and the fitted value from the sample regression.
Residuals are observable and form the basis of many regression diagnostics. Analysts inspect them to look for curvature, changing variance, dependence, and unusual observations.
Regression vs Correlation: What Is the Difference?
Correlation and regression both examine linear relationships, but they answer different questions.
Area | Correlation | Simple Linear Regression |
|---|---|---|
Main purpose | Measure the strength and direction of linear association | Estimate how changes as changes |
Variable roles | Symmetric | Dependent and independent roles are assigned |
Units | Unitless | The slope has units of per unit of |
Main output | Correlation coefficient | Fitted regression equation |
Prediction | Does not create a prediction equation | Can calculate fitted values |
Swapping variables | Correlation remains the same | Regression coefficients change |
Correlation between and is the same as correlation between and .
A regression of on , however, differs from a regression of on . The analyst chooses a direction based on the question being studied.
Worked Example: Company and Industry Sales Growth
An analyst studies the relationship between a specialty retailer’s quarterly sales growth and the quarterly sales growth of its industry.
Both variables are measured in percentage points.
The estimated equation is:
Where:
= predicted company sales growth
= industry sales growth
Interpreting the Intercept
The intercept is 1.20.
When industry sales growth equals zero, the model predicts company sales growth of 1.20 percentage points.
This interpretation is reasonable when observations with industry growth near zero are included in the sample.
Interpreting the Slope
The slope is 1.35.
A one-percentage-point increase in industry sales growth is associated with an expected 1.35-percentage-point increase in company sales growth.
The slope describes how the company’s predicted growth changes with the industry. It does not mean that industry growth necessarily causes the company’s growth.
Calculating a Fitted Value
Suppose industry sales growth is 4.0%.
The predicted company sales growth is 6.60%.
Calculating a Residual
Suppose the company’s actual sales growth is 7.10%.
The residual is positive 0.50 percentage points. The company’s actual growth was 0.50 percentage points above the model’s prediction.
A single residual describes one observation. Analysts look for patterns across many residuals before drawing conclusions about the model.
Interpreting the Relationship Carefully
The regression shows a positive association between company and industry sales growth.
It does not prove that industry growth causes the company’s sales to change. Both variables may respond to consumer spending, economic conditions, seasonality, pricing, or other influences that are not included in the regression.
Common Exam Traps
Reversing the dependent and independent variables. The research question determines which variable is and which is .
Confusing population and sample notation. Use and for population parameters and and for sample estimates.
Interpreting the slope without units. A slope should be stated in units of per unit of .
Using the intercept as a practical forecast without checking the data range. The interpretation may have little economic meaning when is outside the observed sample.
Forgetting the intercept when calculating a fitted value. Use both and .
Subtracting in the wrong direction when calculating a residual. A residual is observed minus fitted.
Treating the residual as the error term. The residual comes from the estimated line. The error comes from the unknown population line.
Assuming a positive slope proves causation. Regression establishes an estimated association, not a causal relationship.
Treating regression and correlation as identical. Regression assigns variable roles and produces a prediction equation.
Practice Question
An analyst estimates a simple linear regression of a portfolio’s monthly return on the monthly return of a benchmark index. Both returns are measured in percent.
The estimated regression is:
If the benchmark return is 6.0%, the portfolio’s predicted return is closest to:
4.8%
7.3%
15.8%
Solution
Correct Answer: B
Substitute the benchmark return into the fitted equation:
The predicted portfolio return is 7.3%.
Option A. 4.8% includes the slope contribution but leaves out the intercept.
Option C. 15.8% incorrectly treats the intercept as the coefficient multiplied by .
Continue Your CFA Level I Prep With KeyPoint
Use structured lessons, practice questions, mock exams, and progress tracking to focus on the time you have left
FAQs About Simple Linear Regression
What Is Simple Linear Regression?
Simple linear regression estimates the average linear relationship between one dependent variable and one independent variable.
It produces an intercept and a slope that can be used to interpret the relationship and calculate fitted values within a relevant range of the data.
What Is the Simple Linear Regression Formula?
The population model is:
The estimated regression equation is:
The population coefficients are unknown, while the estimated coefficients are calculated from sample data.
What Does the Slope of a Regression Line Mean?
The slope is the expected change in the dependent variable for a one-unit increase in the independent variable.
Its interpretation must include the measurement units. For example, a slope may represent percentage points of portfolio return per percentage point of market return.
What Is the Difference Between a Residual and an Error?
An error is the difference between an observed value and the value from the unknown population regression line.
A residual is the difference between an observed value and the fitted value from the estimated sample regression. Residuals can be calculated, while population errors cannot be directly observed.
What Is the Difference Between Regression and Correlation?
Correlation measures the strength and direction of a linear association without assigning dependent and independent roles.
Regression assigns those roles, estimates the expected change in one variable as the other changes, and produces an equation that can be used for prediction.
Can Simple Linear Regression Prove That X Causes Y?
No. Simple linear regression estimates an association between the variables.
A causal conclusion requires additional economic reasoning and a research design that addresses omitted variables, reverse causality, and other possible explanations.