Variance for Grouped Data and Frequency Distributions

Find the variance from a frequency table or grouped class intervals using midpoints: formulas for sample and population, with a worked example table.

Variance of Grouped Data and Frequency Tables

You have a frequency table and need the variance. The raw data is gone. The exact steps, the formulas, the mistakes that cost you points, and how to run the calculation on a TI-84 or in Excel are all included. The variance of grouped data is always an approximation. Accept that and move on.

The core idea is unchanged: variance is the average of the squared differences from the mean. The only difference is that you work with class midpoints and frequencies instead of individual values. OpenStax Introductory Statistics 2e (section 2.7, "Measures of the Spread of the Data") gives the standard approach for grouped frequency tables and is the source for the formulas used here.

Ungrouped Frequency Tables

Before grouped classes, understand the simpler case: a frequency table where each row lists a single value x and its frequency f. You have five data points with value 3, eight with value 7, and so on. The mean is (Σ f x) / (Σ f). The sum of squared deviations is Σ f (x − x̄)². Divide by N for population variance or by n−1 for sample variance. This is the foundation. The grouped version extends it by replacing each class with its midpoint.

Grouped Classes: Using Midpoints

How Midpoints Replace Raw Data

When data is binned into class intervals (1-10, 11-20, etc.), you no longer know the individual values. The standard workaround is to use the midpoint of each class interval as the representative value for every observation in that class. For the interval 10-19 with a lower boundary of 9.5 and an upper boundary of 19.5, the midpoint m is (9.5 + 19.5) / 2 = 14.5. OpenStax uses m for this variable. The assumption is that values are uniformly distributed within each class. This is never exactly true, so the variance from grouped data is an estimate.

Calculate the mean of the grouped data as x̄ = (Σ f m) / n, where n is the total number of observations (the sum of all frequencies). From there, the squared deviations are (m − x̄)², weighted by f.

Grouped Data Variance Formula

One-Pass Formula for Sample and Population

OpenStax presents the grouped frequency table formula in a computationally efficient form that avoids calculating each deviation separately. For a sample:

s² = [Σ(f m²) − ( (Σ(f m))² / n )] / (n − 1)

For a population:

σ² = [Σ(f m²) − ( (Σ(f m))² / N )] / N

The variable m is the midpoint of each class interval, f is the frequency of that class, and n or N is the total number of data values (sum of frequencies). The numerator Σ(f m²) − ( (Σ(f m))² / n ) is identical to the sum of squared deviations (SS) but calculated in one pass.

Bessel's correction (n−1) appears in the sample formula to produce an unbiased estimator of the population variance. Using n instead would systematically underestimate the true spread. This is the single most common error in variance calculations, and it matters for every exam and every real dataset.

Worked Example: Variance From Frequency Distribution

Textbook Data: Book Prices in Six Class Intervals

Use the textbook data from OpenStax Introductory Statistics 2e, Table 2.28. The dataset has n = 50 book prices grouped into six class intervals. The following table works through the calculation step by step.

Worked Example: Grouped Frequency Table of Textbook Data
Class IntervalFrequency (f)Midpoint (m)f × mf × m²
1–1055.527.5151.25
11–20815.5124.01922.0
21–301225.5306.07803.0
31–401035.5355.012602.5
41–50745.5318.514491.75
51–60855.5444.024642.0
Totaln = 50—Σ(f m) = 1575Σ(f m²) = 61612.5

Σ(f m) = 1575, so the mean is 1575 / 50 = 31.5. The sample variance s² = [61612.5 − (1575² / 50)] / (50 − 1) = [61612.5 − 49612.5] / 49 = 12000 / 49. The sample standard deviation s ≈ 15.65. OpenStax reports s² = 812.26 and s = 28.5 for a different dataset; the values here are calculated from the table above and verify the method.

Check the population version: σ² = [61612.5 − (1575² / 50)] / 50 = 12000 / 50. The population standard deviation σ ≈ 15.49. The difference between the sample and population figures is entirely due to Bessel's correction (n−1 vs. N).

Accuracy: Why Grouped Variance Is an Estimate

Grouping Error and Sheppard's Correction

The midpoint assumption introduces error. Within a class from 21 to 30, some books cost 22 and some cost 29; using 25.5 for every book in that class loses that variation. The true variance of the original raw data would be slightly different from the grouped estimate. This is called the grouping error. For data with roughly uniform distributions within each class and many narrow intervals, the error is small. For wide or irregular classes, it can be meaningful.

Sheppard's correction is a formula that adjusts variance for the grouping effect, specifically for data grouped into equal-width intervals. It subtracts (width² / 12) from the calculated variance. For the example above with a class width of 10, the correction would subtract 100 / 12 from the grouped variance. Most introductory courses and exams do not require Sheppard's correction, but you should know it exists and that grouped variance is an approximation, not a precise measure.

Doing It in Excel

Manual Expansion and VAR.S

Excel has no dedicated function for grouped data variance. Build the calculation manually using the formulas above or by creating a column of midpoints repeated by their frequencies and applying VAR.S or VAR.P to that expanded list. For the textbook example, expand 5 copies of 5.5, 8 copies of 15.5, and so on, producing a 50-row dataset. Enter =VAR.S(A1:A50) for sample variance. The result matches the grouped formula exactly. Using VAR.P instead with the same expanded data gives the population variance.

The failure case: using VAR.P on a sample expands the grouping error with the bias error, underestimating the true population variance twice.

Doing It on a TI-84

1-Var Stats and the Sx vs. σx Trap

Enter the midpoints in one list (L1) and the frequencies in another (L2). Run 1-Var Stats L1, L2. The output shows x̄ (the grouped mean), n (the total number of observations), Sx (the sample standard deviation), and σx (the population standard deviation). Square Sx to get sample variance; square σx to get population variance. The TI-84 Plus CE Guidebook confirms these outputs.

The failure case: reporting σx when the problem asks for sample variance. The TI-84 labels both outputs on the same screen, and students routinely grab the wrong one. Check the problem: if the dataset is a sample, use Sx. If it is the entire population, use σx.

Common Questions

What is the formula for variance of grouped data for a sample?

s² = [Σ(f m²) − ( (Σ(f m))² / n )] / (n − 1), where m is the class midpoint, f is the class frequency, and n is the total number of observations.

What is the formula for variance of grouped data for a population?

σ² = [Σ(f m²) − ( (Σ(f m))² / N )] / N, where N is the total number of observations.

How do I find the midpoint of a class interval?

Add the lower and upper class boundaries and divide by two. If boundaries are not given, use the lower and upper limits of the interval.

Why is the grouped variance only an estimate?

Because all values in a class are represented by the midpoint, losing the variation within each class. This grouping error is smaller with narrow, equal-width intervals.

What is Sheppard's correction?

For equal-width intervals, subtract (class width² / 12) from the calculated variance to approximate the raw data variance. It is not always required but explains the gap between grouped and ungrouped variance.

Which Excel function should I use for a sample's grouped variance?

VAR.S. Expand the midpoints by their frequencies into a single column and apply VAR.S to that column.

On the TI-84, do I use Sx or σx for sample variance?

Sx is the sample standard deviation. Square it to get sample variance. σx is the population standard deviation. The TI-84 labels both, so check your data context.