What Is Variance in Statistics?
Variance measures how far data spreads from its mean, in squared units. What it means, how to judge high or low variance, and where it is used in practice.
What Is Variance in Statistics?
OpenStax Introductory Statistics 2e covers what is variance in Section 2.7, and the unbiasedness proof for the n−1 denominator is found in Casella & Berger's Statistical Inference, 2nd edition. The term "variance" itself was first published by R.A. Fisher in 1918, in "The Correlation between Relatives on the Supposition of Mendelian Inheritance."
Why Deviations Are Squared (And Why They Sum to Zero Otherwise)
The failure case: a student who skips the squaring step and averages the raw deviations will always get zero and conclude there is no spread. That error is so common that some textbooks teach it as a demonstration of why squaring is necessary.
Two Datasets, Same Mean, Different Variance
This example makes the core point: variance measures spread, not central tendency. Two datasets can have the same average but completely different risk, consistency, or predictability. Variance is what tells you apart.
Is My Variance High or Low? Compare on the SD Scale, Relative to the Data
When Variance Is Zero
A variance of zero means every value in the dataset is identical. This is rare outside of controlled measurements or coded data. If you see a variance of zero in a real-world dataset, check for data entry error or a constant column.
Can Variance Be Negative?
No. Variance is the average of squared deviations, and a squared number is never negative. The smallest possible variance is zero. The question "can variance be negative" is one of the most common searches on the topic, and the answer is always no for any real-valued dataset. If your calculation produces a negative variance, you have made a computational error, likely a sign mistake, a wrong sign in the covariance term in a portfolio calculation, or an error in the sum of squared deviations.
Where Variance Is Used
The failure case: using VAR.P in Excel on a sample instead of VAR.S. Microsoft's documentation specifies that VAR.P computes population variance (divides by n). On a sample, this underestimates true population variance. Google Docs Editors Help documents the same distinction for VAR, VARP, VAR.S, and VAR.P. On a TI-84, the 1-Var Stats command outputs both σx (population) and Sx (sample). Reporting σx when your data is a sample is the most common error on that platform.
| Measure | Formula | Units | Interpretation |
|---|---|---|---|
| Population variance (σ²) | Σ(xᵢ − μ)² / N | Squared original units | Spread in squared units; not directly interpretable |
| Sample variance (s²) | Σ(xᵢ − x̄)² / (n − 1) | Squared original units | Unbiased estimate of σ²; uses Bessel's correction |
| Standard deviation (σ or s) | √(variance) | Original units | Average distance from the mean; the interpretable version of variance |
| Coefficient of variation (CV) | σ / μ | Unitless ratio | Relative spread; valid only for ratio-scale data with positive mean |
Common Questions
Can variance be negative?
No. Variance is the average of squared deviations, and a squared number is never negative. The minimum is zero, which occurs only when all values are identical.
What does a variance of zero mean?
All data points are identical. This is extremely rare in natural datasets and usually indicates a measurement error or a constant column.
How do I interpret a variance value directly?
You cannot easily, because it is in squared units. Take the square root to get the standard deviation, which is in the original units. Then compare that to the mean or to domain benchmarks.
What is the difference between σ² and s²?
σ² is population variance (divide by N). s² is sample variance (divide by n−1). The n−1 denominator, called Bessel's correction, makes s² an unbiased estimator of σ². Use s² when you have a sample, σ² only when you have the entire population.
What does it mean if my calculated variance is larger than the mean?
It means the standard deviation is larger than the mean. This is possible and common. It does not indicate an error, it just tells you the data have high relative spread. The coefficient of variation would be greater than 1.