noesisnoesis
How measurement works

Standard error of measurement: how much your score could be off

Every score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.

The standard error of measurement converts an abstract reliability coefficient into something concrete: a plus-or-minus band around a score. It answers the question most results screens leave unanswered — how much could this number have differed if I had taken the test on a different afternoon?

The formula is simple. Multiply the standard deviation of the scale by the square root of one minus its reliability. For a scale with a standard deviation of 10 and reliability of 0.85, the standard error is about 3.9 points.

What that band actually covers

Roughly two thirds of the time, a person's true standing lies within one standard error of their observed score. To be about 95 percent confident, you need roughly two standard errors either side.

Take the example above. Someone scores 62. The 95 percent interval runs from about 54 to 70. That is a sixteen-point window on a scale where the typical gap between average and clearly-above-average might be ten points.

This is not a defect of that particular test. A reliability of 0.85 is respectable. It is the normal precision of psychological measurement, and it is the reason careful reports present bands rather than points.

The consequences people miss

Small differences between scales are usually meaningless. If your extraversion is 58 and your openness is 54, the honest reading is that they are indistinguishable. Profile interpretations that build a narrative from four-point gaps are reading noise.

Small changes over time are usually meaningless too. Retaking a test and moving five points is within error for most instruments. Real trait change happens over years and shows up as movement well outside the error band, not as a few points month to month.

Rank orderings amplify the problem. A difference of a few raw points can move someone ten percentile places in the dense middle of a distribution, because that is where most people sit. Percentile ranks look precise and are the least stable way to express a score.

Precision is not uniform across the scale. Classical test theory assumes one error value for every score, which is a simplification. Item response theory models error separately at each level of the trait, and it is almost always larger at the extremes — precisely where the most consequential interpretations get made. A very high or very low score is typically less precise than a middling one, not more.

Why honest reporting looks less impressive

An instrument that reports confidence bands looks vaguer than one that reports a single number to one decimal place. The vaguer one is telling the truth. The precise-looking one has the same underlying error and has chosen not to show it.

This is worth holding on to when comparing a scientifically documented test with a commercial product that assigns you to a category. The category boundary sits at a specific score, and someone one point below it and someone one point above are, statistically, the same person. The type label conceals that entirely — which is one of the more serious arguments against type-based instruments, and the reason continuous scores with stated error are the standard in research.

Put it to the test

Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.

Keep reading