What does it mean when a test is called validated?
The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
"Scientifically validated" is not a protected term. Nobody certifies it, no body audits it, and it appears on products whose entire evidence base is that a lot of people have taken them. It is worth knowing what it means when it is used properly.
The short version
A validated instrument has a published record showing what it measures, how precisely, in which populations, and what its scores can legitimately be used for. Validation is not a certificate. It is a body of evidence, usually assembled over years, usually including findings that constrain the instrument's use rather than flattering it.
The presence of documented limitations is one of the better signals. Genuine validation literature is full of them: this scale performs poorly in adolescents, that facet does not replicate, the cutoff is population-dependent. A product that reports only strengths has not published its validation; it has published its marketing.
Four questions that do the work
What is the source publication? A validated instrument has an original paper in a peer-reviewed journal, with authors, a year and usually a DOI. That paper reports how the items were developed and what the initial psychometric properties were. If nobody can name the paper, there is no validation to discuss.
Was the structure confirmed by someone else? The development team will recover the structure they designed. What matters is whether an independent group, with independent data, found the same thing. This is the single most informative question you can ask, and the one most often left unanswered.
Which population were the norms built on, and when? A score is a position in a reference distribution. An instrument that reports a percentile without stating what population it refers to has given you a number with no denominator.
What does it add? Incremental validity over an established measure. A great many branded constructs fail here, correlating so highly with an existing trait that they explain no additional variance in anything. This is a normal outcome of scientific scrutiny, and it is invisible unless someone runs the test.
Where translation complicates it
An instrument validated in English is not thereby validated in German. Proper cross-cultural adaptation requires forward and back translation, expert review, cognitive interviewing, and then a fresh psychometric study in the new language testing whether the factor structure holds and whether individual items function equivalently across groups.
Instruments that skip this are common. A questionnaire whose original validation is impeccable can behave quite differently in a language where a key item carries a different connotation, and there is no way to know without the study. When a test is offered in ten languages, the useful question is whether it was validated in ten languages or validated in one and translated into nine.
What validation does not buy
It does not make a score certain. A validated instrument still carries measurement error, still describes a population-level regularity rather than an individual destiny, and still measures only what its items reach.
What validation buys is knowable limits. You can find out how much error, in which population, for which purpose. That is a smaller claim than "this test reveals who you really are", and it is the only one that holds up.
If you want to see what the documentation looks like when it exists, every test in the NOESIS catalog names the instrument it implements, its authors and year, its scoring method, and the evidence behind it — including where that evidence is contested.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.