What it means when a test predicts something
A correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
Personality research reports that conscientiousness predicts job performance. Reading that sentence, most people picture something far stronger than what the data show. The gap between the statistical claim and the everyday meaning of "predicts" is where most misuse of psychological testing begins.
The numbers involved
Effect sizes in personality research are usually reported as correlations. Meta-analytic values for well-established relationships tend to land between 0.20 and 0.35. Conscientiousness and job performance sits around 0.20 to 0.25. Job satisfaction and job performance is around 0.30. General mental ability, the strongest single predictor in the field, reaches roughly 0.50 for complex work.
Square a correlation and you get the proportion of variance explained. A correlation of 0.25 accounts for about six percent of the differences between people. Ninety-four percent is something else: ability, opportunity, training, management, health, luck.
Why 0.25 is nonetheless a real finding
At population scale, small correlations produce substantial aggregate effects. If you are selecting a thousand people, a predictor at 0.25 improves your average outcome measurably and repeatedly. Insurers and epidemiologists build entire practices on effects of this size.
The catch is that the aggregate benefit and the individual prediction are different quantities. A correlation of 0.25 means that among people who score high, more will perform well than among people who score low — and also that a great many high scorers will perform poorly and a great many low scorers will excel. Both statements come from the same number.
The individual case
This is the step that gets skipped. A correlation describes the relationship between two distributions. It does not license a statement about any particular person.
If conscientiousness correlates 0.25 with performance, then knowing someone's conscientiousness score narrows the plausible range of their performance only slightly. The prediction interval for one individual remains wide enough to include almost the entire outcome range. Being told you are at the 80th percentile on conscientiousness tells you very little about how you will do in a specific job.
This is why using personality scores as a primary selection filter is poorly supported even when the underlying research is sound. The research supports the claim that the trait matters. It does not support the claim that a score tells you what an applicant will do.
Two further cautions
Predicting and causing are different. Almost all of this literature is correlational. Conscientiousness predicting health outcomes is consistent with conscientiousness causing better health behaviour, with a third factor causing both, and with reverse causation. Longitudinal designs narrow the possibilities; they rarely close them.
Self-report criteria inflate the numbers. When both the predictor and the outcome are self-reported by the same person on the same afternoon, common-method variance pushes the correlation up. Studies using objective outcomes or informant ratings usually report smaller effects, and those are the ones to weight.
What a score is actually good for
Read as description rather than prophecy, a test score is genuinely useful. It tells you where you stand relative to a reference population on a trait that has documented consequences, and it gives you vocabulary for a pattern you may have noticed without being able to name.
That is worth having. It is a different thing from a forecast, and the distinction is one an honest report should make for you rather than leave you to infer.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.