Convergent and discriminant validity
Two mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
Convergent and discriminant validity are the two halves of a single requirement. A measure must correlate with things it theoretically should, and must not correlate too highly with things it is supposed to be distinct from. Meeting one without the other is not partial success; it is usually a sign that the construct is not what it says it is.
Convergent validity
Convergent evidence shows a measure lines up with other indicators of the same thing. A new anxiety scale should correlate substantially with established anxiety scales, with clinician ratings, and — ideally — with something that is not another questionnaire, such as physiological measures or observed avoidance behaviour.
Correlations here are usually expected above 0.50, often above 0.70 when the comparison measure is a well-established instrument for the same construct. Below that, either the new scale or the comparison is measuring something else.
There is a trap. A correlation of 0.95 with an existing scale is not triumph — it means the new instrument is a repackaging, and the honest question becomes what it adds. Convergent validity is a floor, not a target to maximise.
Discriminant validity
Discriminant evidence shows the measure stays separate from constructs it claims to differ from. This is where the interesting failures live.
The pattern recurs across the literature. Grit correlates around 0.84 with conscientiousness, close enough that meta-analysis found it adds almost nothing to academic prediction once conscientiousness is controlled. Self-report emotional intelligence overlaps heavily with emotional stability, extraversion and conscientiousness combined. The shared variance across the Dark Triad largely disappears once Honesty-Humility from the HEXACO model is taken into account. In each case the construct was presented as new territory and turned out to be a renamed region of an existing map.
Discriminant validity is also where method effects surface. If every self-reported construct in a study correlates 0.4 with every other, the common factor is probably the respondent's response style rather than any shared psychology.
The test that settles it
The decisive procedure is incremental validity: enter the established measure into a regression first, then the new one, and see whether the new one explains any additional variance in the outcome. If it does not, the construct may still be theoretically interesting, but the instrument is not measuring anything the field did not already have.
This is a demanding standard and a great many published scales have never been subjected to it. When they are, the result is frequently the redundancy finding above. That is not a scandal — it is normal scientific pruning — but it does mean that a construct's popularity is uninformative about whether it survives the test.
What it changes for a reader
When a test reports several scores, the useful question is whether those scores are actually distinct. Four dimensions that correlate 0.8 with each other are one dimension with four labels, and a profile built from them will look differentiated while carrying almost no independent information.
Instruments that report their inter-scale correlations let you check this yourself. Instruments that present a four-quadrant type without ever showing how the quadrants relate have made the check impossible, which is usually the point.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.