Web Analytics
QUANTITATIVE METHODS

Chi-Square Tests of Independence Using Contingency Tables

By KeyPoint Learning 9-minute read
CFA CFA Level I

Updated for the 2026-2027 CFA® Level I curriculum.

When two variables are categorical, their relationship appears as a pattern of counts rather than as a correlation coefficient. A chi-square test of independence examines that pattern and asks whether the two classifications appear related.

For CFA Level I, you should be comfortable reading a contingency table, calculating expected frequencies, finding the chi-square statistic and degrees of freedom, and interpreting the final decision.

Quick Answer

A chi-square test of independence evaluates whether two categorical variables are associated. It compares the observed counts in a contingency table with the counts expected if the variables were independent. Larger differences between observed and expected frequencies produce a larger chi-square statistic and stronger evidence against the null hypothesis of independence.

Key Takeaways About Chi-Square Tests of Independence

  • A contingency table cross-classifies observations using two categorical variables.

  • The null hypothesis states that the variables are independent.

  • The alternative hypothesis states that the variables are associated.

  • Observed frequencies come directly from the sample.

  • Expected frequencies represent the counts predicted under independence.

  • Each expected frequency uses the relevant row total, column total, and grand total.

  • The chi-square statistic combines the difference between observed and expected frequencies across all cells.

  • Degrees of freedom equal .

  • The chi-square test of independence is right-tailed.

  • A significant result supports association, not causation.

What You Need to Know for CFA Level I

For CFA Level I, you should be able to:

  • Identify the row and column variables in a contingency table.

  • Distinguish observed frequencies from expected frequencies.

  • State the null and alternative hypotheses.

  • Calculate an expected frequency for a given cell.

  • Calculate an individual cell’s contribution to the test statistic.

  • Calculate the total chi-square statistic.

  • Determine the correct degrees of freedom.

  • Apply a critical-value or p-value decision rule.

  • Interpret the result without making a causal claim.

What Is a Contingency Table?

A contingency table is a grid of frequency counts that organizes observations according to two categorical variables.

The rows represent the categories of one variable, while the columns represent the categories of the other. Each interior cell shows how many observations belong to both categories at the same time.

For example, an investment firm might classify clients by:

  • Risk tolerance: conservative or aggressive

  • Product choice: bond fund, balanced fund, or equity fund

The table would then show how many conservative and aggressive clients selected each product.

The row totals and column totals are called marginal totals. The grand total is the full number of observations in the sample. These totals are used to calculate the expected frequency for each cell.

What Does a Chi-Square Test of Independence Measure?

The test measures how far the observed counts differ from the counts that would be expected if the two categorical variables were independent.

The hypotheses are:

Under , knowing the category of one variable provides no information about the category of the other.

Suppose risk tolerance and product choice are independent. In that case, the proportion of clients choosing each product should be broadly similar across the risk-tolerance categories, after accounting for the different row totals.

Large differences between the observed and expected patterns provide evidence that the variables may be associated.

The chi-square test of independence is a nonparametric procedure. It works with category counts and does not require an assumption that an underlying numerical variable follows a normal distribution.

How Do You Calculate Expected Frequencies?

The expected frequency is the number of observations that would fall in a cell if the two variables were independent.

Use the following formula:

Where:

  • = expected frequency in row and column

  • = total number of observations in row

  • = total number of observations in column

  • = total number of observations in the table

Suppose a cell belongs to a row with 110 observations and a column with 70 observations. The table contains 200 observations in total.

The expected frequency is 38.5.

Expected frequencies do not need to be whole numbers. Keep the available decimal places during the calculation because rounding too early can change the final chi-square statistic.

What Is the Chi-Square Test Statistic Formula?

The chi-square test statistic adds the contribution from every cell in the contingency table.

Where:

  • = calculated chi-square test statistic

  • = observed frequency in row and column

  • = expected frequency in row and column

  • = number of row categories

  • = number of column categories

Each cell contribution is:

The difference between the observed and expected frequency is squared, so both positive and negative deviations increase the statistic.

Dividing by the expected frequency adjusts the difference for the size of the cell. A difference of 10 observations is more notable when the expected frequency is 20 than when it is 500.

When observed and expected counts are close throughout the table, remains relatively small. Larger differences cause the statistic to increase.

How Do You Calculate Degrees of Freedom?

For a chi-square test of independence, degrees of freedom are:

Where:

  • = degrees of freedom

  • = number of row categories

  • = number of column categories

A table with two row categories and three column categories has:

Do not count the totals row or totals column as additional categories. They summarize the interior cells and do not increase the degrees of freedom.

How Do You Make the Test Decision?

The chi-square test of independence is right-tailed.

All cell contributions are zero or positive because their differences are squared. A larger value of therefore represents a greater departure from the independence assumption.

You can apply either of two equivalent rules.

Critical-Value Rule

Reject when:

The critical value depends on:

  • The significance level

  • The degrees of freedom

P-Value Rule

Reject when:

If you reject the null hypothesis, conclude that the two variables are associated at the stated significance level.

If you fail to reject it, conclude that the available evidence is insufficient to establish an association. Avoid saying that the variables have been proven independent.

What Conditions Should You Check?

The chi-square test requires observations that are independent and expected frequencies that are large enough for the chi-square approximation to be appropriate.

Independent Observations

Each observation should appear in one cell only, and one observation should not determine another observation’s classification.

For example, treating repeated monthly records from one client as if they came from unrelated clients could violate the independence condition.

Suitable Expected Frequencies

The reliability of the chi-square approximation depends on the expected cell frequencies. Very small expected counts can make the approximation unsuitable.

Check the expected frequencies rather than only the observed frequencies. A cell may have a low observed count while still having an acceptable expected count under the null hypothesis.

Frequency Counts

The table should contain counts, not percentages or averages. If a question provides percentages, convert them into counts using the sample size before applying the formulas.

Worked Example: Risk Tolerance and Product Choice

An analyst reviews 200 client accounts. Each client is classified according to stated risk tolerance and selected investment product.

Observed Frequencies

Risk Tolerance

Bond Fund

Balanced Fund

Equity Fund

Row Total

Conservative

48

44

18

110

Aggressive

22

31

37

90

Column Total

70

75

55

200

Step 1: State the Hypotheses

Step 2: Calculate the Expected Frequencies

For conservative clients choosing a bond fund:

Applying the same calculation to every cell gives:

Risk Tolerance

Bond Fund

Balanced Fund

Equity Fund

Conservative

38.50

41.25

30.25

Aggressive

31.50

33.75

24.75

Step 3: Calculate Each Cell Contribution

Cell

Conservative, bond

48

38.50

2.3442

Conservative, balanced

44

41.25

0.1833

Conservative, equity

18

30.25

4.9607

Aggressive, bond

22

31.50

2.8651

Aggressive, balanced

31

33.75

0.2241

Aggressive, equity

37

24.75

6.0631

Step 4: Calculate the Chi-Square Statistic

Step 5: Calculate Degrees of Freedom

Step 6: Apply the Decision Rule

At with 2 degrees of freedom, the critical value is approximately:

Because:

the calculated statistic falls in the rejection region.

Decision: Reject

Step 7: Interpret the Result

At the 5% significance level, the evidence supports an association between stated risk tolerance and product choice.

The largest cell contributions come from the equity-fund column. Conservative clients selected equity funds less often than independence would predict, while aggressive clients selected them more often.

The test identifies an association between the variables. It does not establish that stated risk tolerance directly caused the product selection.

Common Exam Traps

  • Using percentages instead of counts. Convert percentages into frequencies before applying the formula.

  • Dividing by the observed frequency. The denominator is always the expected frequency.

  • Forgetting to square the difference. Without the square, positive and negative deviations can cancel.

  • Rounding expected counts too early. Keep sufficient decimal places until the final statistic is calculated.

  • Including totals in the degrees-of-freedom calculation. Count only the actual row and column categories.

  • Using the observed frequency to judge the small-cell condition. Review the expected frequencies.

  • Treating the test as two-tailed. The rejection region lies entirely in the right tail.

  • Saying that failure to reject proves independence. It only indicates insufficient evidence of association.

  • Interpreting association as causation. The test cannot determine causal direction.

Practice Question

An analyst classifies 200 equity recommendations by analyst rating and subsequent 12-month performance.

Rating

Outperform

Underperform

Row Total

Buy

62

58

120

Hold

28

52

80

Column Total

90

110

200

Under the null hypothesis of independence, the expected frequency for the Buy and Outperform cell is closest to:

  1. 45.0

  2. 54.0

  3. 60.0

  • Correct Answer: B

Use the Buy row total, the Outperform column total, and the grand total.

The expected frequency is 54.0.

  • Option A divides the Outperform column equally between the two rating categories, even though the row totals are different.

  • Option C divides the Buy row equally between the two performance categories, even though the column totals are different.

Continue Your CFA Level I Prep With KeyPoint

Use structured lessons, practice questions, mock exams, and progress tracking to focus on the time you have left

FAQs About Chi-Square Tests of Independence

The null hypothesis states that the two categorical variables are independent.

The alternative hypothesis states that the variables are associated or not independent.

Multiply the relevant row total by the relevant column total, then divide by the grand total.

Calculate the expected frequency separately for every interior cell.

The formula is:

It adds the scaled difference between the observed and expected frequency across all cells.

Use:

Count only the row and column categories. Exclude the totals row and totals column.

Yes. Greater differences between observed and expected frequencies produce larger positive values of the chi-square statistic.

You reject the null hypothesis only when the statistic is sufficiently large.

No. A significant result supports an association between the two categorical variables.

It does not show that one variable causes the other or identify the direction of any causal relationship.

On This Page

Explore KeyPoint Learning

  • Video Lessons
  • Study Notes
  • Practice Quizzes
  • Mock Exams
  • Progress Tracking
Explore CFA Study Packages

Get CFA Insights in Your Inbox

Adding to Cart

Preparing your study package access...