noesisnoesis
How measurement works

What is reliability in psychological testing?

Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.

Reliability is the degree to which a measurement is repeatable rather than noise. If you weighed yourself three times in a minute and got 71, 78 and 66 kilograms, the scale would be unreliable — and you would have no way to know your actual weight from it. Psychological measurement faces the same problem with far less precise equipment.

The critical thing to understand about reliability is what it does not claim. A bathroom scale that consistently reads eight kilograms heavy is perfectly reliable and completely wrong. Reliability is about consistency; whether the number means the right thing is a separate question, called validity. Reliability is necessary for validity and nowhere near sufficient for it.

The forms it takes

Internal consistency asks whether the items on a scale agree with each other. If ten questions are all supposed to tap the same underlying trait, people who answer high on one should tend to answer high on the others. This is usually reported as Cronbach's alpha or, better, McDonald's omega.

Test-retest reliability asks whether the same person scores similarly on two occasions. This is the form that matters most for traits, which by definition are supposed to be stable. It is also the form most often left unreported, because it requires tracking down the same participants weeks later.

Inter-rater reliability applies when humans score the instrument — clinical interviews, projective techniques, behavioural observation. It asks whether two trained raters looking at the same material reach the same conclusion.

Parallel forms reliability asks whether two different versions of the same test, built to be equivalent, produce the same scores. It matters wherever a test is administered more than once and item exposure is a concern.

Reading the numbers

Reliability coefficients run from 0 to 1. The conventions are rough and widely misused, but roughly: above 0.90 is required where a score affects an individual's life, 0.80 to 0.90 is solid for most applied use, 0.70 to 0.80 is acceptable for research and for group comparisons, and below 0.70 means the score contains more noise than most conclusions can survive.

Two cautions about these thresholds. First, they are conventions, not laws — the acceptable level depends entirely on what the score will be used for. Second, a very high alpha on a short scale usually indicates the items are near-paraphrases of each other, which buys consistency at the cost of covering the construct. A scale of ten questions that all ask "are you organised?" in slightly different words will have superb internal consistency and a narrow reading of conscientiousness.

What it means for your own score

Reliability translates directly into how much confidence a single number deserves. A test with a reliability of 0.85 carries measurement error large enough that a difference of a few points between two people, or between your score today and your score last month, is very likely to be noise rather than signal.

This is the practical consequence: a good instrument reports its reliability so you know how wide the uncertainty around your score is. The technical expression of that uncertainty is the standard error of measurement, which converts a reliability coefficient into a plus-or-minus band around the score.

An instrument that reports no reliability figure at all is not thereby more precise. It has simply declined to tell you how imprecise it is.

Put it to the test

Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.

Keep reading