What is validity in psychological testing?
Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
Validity asks whether a test measures what it claims to measure, and whether the conclusions people draw from its scores are warranted. It is the central question in psychometrics and the one most often answered with a marketing claim rather than evidence.
The modern understanding, which has held since Messick's work in the 1980s, is that validity is not a fixed attribute of an instrument. It is a property of an interpretation of scores for a specific use. The same questionnaire can be well validated for describing a person's self-perception and completely unvalidated for deciding whether to hire them. Saying "this test is valid" without saying valid for what is not a claim; it is a slogan.
The kinds of evidence
Validity is built from accumulated evidence of several types. Older texts present these as separate species of validity; the current view treats them as different arguments supporting a single case.
Content evidence asks whether the items cover the construct properly. A depression scale that asks only about sleep and appetite covers the somatic symptoms and misses the mood and cognitive ones. Content coverage is usually judged by subject-matter experts, and it is the reason very short scales are always a compromise.
Internal structure evidence asks whether the items group the way the theory says they should. If a model claims four distinct facets, factor analysis should recover four factors — and crucially, in an independent sample, not just the one the scale was built on. A great many instruments recover their intended structure in the development sample and fail to in anyone else's.
Relations to other variables is where most of the work happens. Convergent evidence shows the scale correlates with things it should correlate with. Discriminant evidence shows it does not correlate too highly with things it should be distinct from — a chronically neglected requirement, and the one that sinks a lot of newly branded constructs. Criterion evidence shows the score relates to an outcome that matters, either measured at the same time or predicted in advance.
Consequential evidence asks what happens when the test is used. If an instrument systematically disadvantages a group for reasons unrelated to the construct, that is a validity problem, not merely an ethical one.
The question that exposes most claims
The sharpest test of a validity claim is incremental validity: does this instrument add anything over what a cheaper, older, established measure already tells you?
A remarkable number of branded constructs fail here. They correlate 0.7 or higher with an existing Big Five domain, predict the same outcomes at the same magnitude, and add nothing once the older measure is statistically controlled. The construct is new; the measurement is not. This is why the critiques sections of serious instrument documentation so often turn on redundancy rather than inaccuracy.
What to look for
When someone claims a test is validated, the useful questions are concrete. Validated against what criterion? In which population, of what size? Was the factor structure confirmed in an independent sample? Is there published evidence from researchers with no commercial interest in the answer? Does it add anything over an established alternative?
An instrument with honest answers to those questions is worth taking seriously, limitations included. An instrument whose validation consists of a testimonial page and a claim that millions of people have taken it has answered a different question entirely — popularity is not evidence, and the number of people who have taken a test says nothing about whether its scores mean anything.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.