What is reliability in psychological testing?
Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
Reliability is the degree to which a measurement is repeatable rather than noise. If you weighed yourself three times in a minute and got 71, 78 and 66 kilograms, the scale would be unreliable — and you would have no way to know your actual weight from it. Psychological measurement faces the same problem with far less precise equipment.
The critical thing to understand about reliability is what it does not claim. A bathroom scale that consistently reads eight kilograms heavy is perfectly reliable and completely wrong. Reliability is about consistency; whether the number means the right thing is a separate question, called validity. Reliability is necessary for validity and nowhere near sufficient for it.
The forms it takes
Internal consistency asks whether the items on a scale agree with each other. If ten questions are all supposed to tap the same underlying trait, people who answer high on one should tend to answer high on the others. This is usually reported as Cronbach's alpha or, better, McDonald's omega.
Test-retest reliability asks whether the same person scores similarly on two occasions. This is the form that matters most for traits, which by definition are supposed to be stable. It is also the form most often left unreported, because it requires tracking down the same participants weeks later.
Inter-rater reliability applies when humans score the instrument — clinical interviews, projective techniques, behavioural observation. It asks whether two trained raters looking at the same material reach the same conclusion.
Parallel forms reliability asks whether two different versions of the same test, built to be equivalent, produce the same scores. It matters wherever a test is administered more than once and item exposure is a concern.
Reading the numbers
Reliability coefficients run from 0 to 1. The conventions are rough and widely misused, but roughly: above 0.90 is required where a score affects an individual's life, 0.80 to 0.90 is solid for most applied use, 0.70 to 0.80 is acceptable for research and for group comparisons, and below 0.70 means the score contains more noise than most conclusions can survive.
Two cautions about these thresholds. First, they are conventions, not laws — the acceptable level depends entirely on what the score will be used for. Second, a very high alpha on a short scale usually indicates the items are near-paraphrases of each other, which buys consistency at the cost of covering the construct. A scale of ten questions that all ask "are you organised?" in slightly different words will have superb internal consistency and a narrow reading of conscientiousness.
What it means for your own score
Reliability translates directly into how much confidence a single number deserves. A test with a reliability of 0.85 carries measurement error large enough that a difference of a few points between two people, or between your score today and your score last month, is very likely to be noise rather than signal.
This is the practical consequence: a good instrument reports its reliability so you know how wide the uncertainty around your score is. The technical expression of that uncertainty is the standard error of measurement, which converts a reliability coefficient into a plus-or-minus band around the score.
An instrument that reports no reliability figure at all is not thereby more precise. It has simply declined to tell you how imprecise it is.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.