Sample Variance vs Population Variance

When to divide by n and when by n - 1: how to tell a sample from a population, what Bessel's correction does, and how much difference it makes in practice.

Sample Variance vs Population Variance

You are looking at a column of numbers in Excel and you need the spread. Two buttons sit next to each other: VAR.S and VAR.P. Pick the wrong one and your variance estimate is biased low. That is not a rounding error; it is a structural mistake that compounds every time you take the square root for standard deviation. The distinction between sample vs population variance comes down to one denominator: n-1 versus n. Here is how to decide, why the minus one exists, and what happens when you get it wrong.

Variance measures how far a set of numbers spread out from their average, in squared units. Population variance (σ²) divides the sum of squared deviations by the total number of values N. Sample variance (s²) divides by n-1. The honest version: it is the average of the squared differences from the mean, and the single biggest mistake newcomers make is thinking the square root (standard deviation) is the variance, or that variance itself is in the original units. It is not; it is in squared units, which is why standard deviation exists.

Population vs Sample: Deciding Which You Have

Identify Your Data Type First

Before you compute anything, decide whether your data is the whole population or just a sample. A population includes every member of the group you care about. A sample is a subset used to estimate the population parameter. The decision is not mathematical; it is logical.

You have a population if you have measured every single item. Example: the exam scores of all 30 students in one class. Example: every transaction in a store on a specific date. Use population variance (σ²) with denominator N.

You have a sample if your data is a subset drawn from a larger group. Example: 200 voters surveyed to estimate the opinion of a city. Example: 12 monthly returns of a fund used to estimate its true volatility. Use sample variance (s²) with denominator n-1.

Use This Checklist

  • Is your data the complete set? Use N.
  • Is your data a subset meant to represent a larger group? Use n-1.
  • Are you using a built-in calculator or Excel function? Check the label: VAR.P for population, VAR.S for sample.
  • Are you working from a textbook problem that says "sample" or "population"? Follow the problem statement.
  • When in doubt, assume sample. Most real-world data is a sample, not a population.

OpenStax Introductory Statistics 2e section 2.7 states that the population variance formula uses denominator N and the sample variance formula uses denominator (n-1).

The Two Formulas Side by Side

Population variance formula:

σ² = Σ(x - μ)² / N

where μ is the population mean, x is each value, and N is the population size.

Sample variance formula:

s² = Σ(x - x̄)² / (n - 1)

where x̄ is the sample mean, and n is the sample size.

The numerator is the same: the sum of squared deviations from the mean. The denominator is the only difference. That difference is called Bessel's correction.

Variance Formulas and Notation
AttributePopulation VarianceSample Variance
Symbolσ²s²
DenominatorNn - 1
FormulaΣ(x - μ)² / NΣ(x - x̄)² / (n - 1)
Excel FunctionVAR.P (or VARP)VAR.S (or VAR)
TI-84 Outputσx²Sx² (square Sx)
Used whenData is the whole populationData is a sample

Why n - 1: Degrees of Freedom and Intuition

The Lost Degree of Freedom

The sample mean x̄ is itself calculated from the data. That means the deviations (x - x̄) are not all free to vary; once you know n-1 of them, the last deviation is forced. You lose one degree of freedom for estimating the mean. Dividing by n-1 corrects for that lost degree of freedom and makes s² an unbiased estimator of σ².

Degrees of freedom is the number of independent pieces of information available to estimate a parameter. For sample variance, that number is n-1. The intuition: if you have two numbers and you know their mean, the two numbers are not both free. The second deviation is determined by the first. Dividing by n-1 instead of n inflates the variance estimate by a factor of n/(n-1), which exactly compensates for the bias introduced by using the sample mean in place of the true population mean.

Bessel's Correction Explained

Bessel's correction is the name for using n-1 instead of n. It is named after Friedrich Bessel, who introduced the correction in the 19th century. Casella and Berger's Statistical Inference (2nd edition) provides the formal proof in Theorem 7.1, showing E[s²] = σ².

Simulation and Illustration: n vs n-1 Over Many Samples

Imagine a population with true variance σ² = 100. You repeatedly draw samples of size n=10. For each sample, you compute sample variance two ways: one using denominator n and one using n-1. Average the results over thousands of samples:

  • Denominator n: The average of the computed variances will be about 90. It is biased low.
  • Denominator n-1: The average of the computed variances will be about 100. It is unbiased.

The n-1 version hits the true value on average. The n version consistently underestimates. That is the practical consequence: using VAR.P on a sample produces a variance estimate that is, on average, too small. For n=10, the bias is 10%. For n=30, the bias is about 3.3%. For n=100, the bias is about 1%.

How Much N-1 Matters as Sample Size Grows
Sample Size (n)Bias if Using n Instead of n-1Correction Factor n/(n-1)
520.0%1.250
1010.0%1.111
205.0%1.053
303.3%1.034
502.0%1.020
1001.0%1.010
10000.1%1.001

Notation: σ² vs s², and Calculator and Excel Names

Symbols and What They Mean

σ² (sigma squared) is the symbol for population variance. s² is the symbol for sample variance. The Greek letter communicates that it is a population parameter; the Latin letter communicates that it is a sample statistic.

Excel and Calculator Commands

Excel: VAR.S and VAR use denominator n-1 for samples. VAR.P and VARP use N for populations. Microsoft Support states VAR.S computes sample variance using denominator (n-1) and VAR.P computes population variance using denominator N. VARA and VARPA include logical and text values (TRUE=1, FALSE=0, text=0). Google Docs Editors Help confirms the same for VAR, VARP, VAR.S, and VAR.P.

TI-84: The 1-Var Stats command outputs Sx (sample standard deviation) and σx (population standard deviation). To get variance, square the standard deviation: Sx² for sample variance, σx² for population variance. The TI-84 Plus CE Guidebook lists both outputs in the same screen. Confusing σx and Sx is the most common error students make on this calculator.

Common Questions

Why is it called Bessel's correction?

It is named after German mathematician Friedrich Bessel, who introduced the use of n-1 in the sample variance formula. He recognized that dividing by n systematically underestimates the true population variance and proposed the correction.

What happens if I use VAR.P on a sample in Excel?

You will underestimate the true population variance. The bias is larger for small samples. For n=10, your estimate will be about 10% too low on average. This compounds when you take the square root for standard deviation, producing a standard deviation that is also too low.

Can I use sample variance when I have the whole population?

No. If you have the entire population, use the population variance formula with denominator N. Using n-1 on a population overestimates the true spread. The only time you use n-1 is when your data is a sample and you want an unbiased estimate of the population variance.

Is sample variance always an unbiased estimator?

Yes, sample variance with denominator n-1 is an unbiased estimator of population variance. This is a proven mathematical result. Casella and Berger's <em>Statistical Inference</em> provides the proof: E[s²] = σ². The sample standard deviation s, however, is not an unbiased estimator of the population standard deviation σ, but that is a separate issue.

Does the n-1 correction matter for large sample sizes?

It matters less as n grows. At n=100, the bias from using n is only about 1%. At n=1000, it is about 0.1%. But even for large samples, you should use the correct formula for the data type you have. The distinction between sample and population is logical, not based on sample size.