Standard error of measurement: how much your score could be off
Every score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
The standard error of measurement converts an abstract reliability coefficient into something concrete: a plus-or-minus band around a score. It answers the question most results screens leave unanswered — how much could this number have differed if I had taken the test on a different afternoon?
The formula is simple. Multiply the standard deviation of the scale by the square root of one minus its reliability. For a scale with a standard deviation of 10 and reliability of 0.85, the standard error is about 3.9 points.
What that band actually covers
Roughly two thirds of the time, a person's true standing lies within one standard error of their observed score. To be about 95 percent confident, you need roughly two standard errors either side.
Take the example above. Someone scores 62. The 95 percent interval runs from about 54 to 70. That is a sixteen-point window on a scale where the typical gap between average and clearly-above-average might be ten points.
This is not a defect of that particular test. A reliability of 0.85 is respectable. It is the normal precision of psychological measurement, and it is the reason careful reports present bands rather than points.
The consequences people miss
Small differences between scales are usually meaningless. If your extraversion is 58 and your openness is 54, the honest reading is that they are indistinguishable. Profile interpretations that build a narrative from four-point gaps are reading noise.
Small changes over time are usually meaningless too. Retaking a test and moving five points is within error for most instruments. Real trait change happens over years and shows up as movement well outside the error band, not as a few points month to month.
Rank orderings amplify the problem. A difference of a few raw points can move someone ten percentile places in the dense middle of a distribution, because that is where most people sit. Percentile ranks look precise and are the least stable way to express a score.
Precision is not uniform across the scale. Classical test theory assumes one error value for every score, which is a simplification. Item response theory models error separately at each level of the trait, and it is almost always larger at the extremes — precisely where the most consequential interpretations get made. A very high or very low score is typically less precise than a middling one, not more.
Why honest reporting looks less impressive
An instrument that reports confidence bands looks vaguer than one that reports a single number to one decimal place. The vaguer one is telling the truth. The precise-looking one has the same underlying error and has chosen not to show it.
This is worth holding on to when comparing a scientifically documented test with a commercial product that assigns you to a category. The category boundary sits at a specific score, and someone one point below it and someone one point above are, statistically, the same person. The type label conceals that entirely — which is one of the more serious arguments against type-based instruments, and the reason continuous scores with stated error are the standard in research.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.