What is psychometrics?
The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
Psychometrics is the branch of psychology concerned with measurement itself. Not with what extraversion means, but with whether a given set of questions actually measures extraversion, how much error is in the resulting number, and what that number is entitled to predict.
The distinction matters more than it sounds. Anyone can write twenty questions about how outgoing someone is and add up the answers. Psychometrics is the work that comes after: showing that the twenty questions hang together, that they measure one thing rather than three, that the same person scores roughly the same next month, that the score relates to how that person actually behaves, and that the whole apparatus works the same way for a 22-year-old in Lisbon as for a 55-year-old in Munich.
Why a discipline was needed
Physical measurement has an advantage psychology does not: you can put a metre rule next to the thing you are measuring. There is no metre rule for conscientiousness. The trait is not directly observable — it is inferred from behaviour, and questionnaire responses are one more layer of inference on top of that.
This is what psychometricians call a latent variable: something real enough to have consequences, but only ever reachable through indirect indicators. The whole methodological apparatus exists because of that indirectness. Every technique you will encounter — factor analysis, reliability coefficients, item response theory, norm samples — is an attempt to say something defensible about a quantity nobody can observe.
The three questions
Strip away the vocabulary and psychometrics asks three things of any instrument.
Is it consistent? If the measurement is mostly noise, nothing else matters. Consistency is checked across items (do the questions agree with each other?), across time (does the same person score the same next month?), and across raters where relevant. This is reliability.
Is it measuring what it claims? A questionnaire can be perfectly consistent and still measure the wrong thing. A scale that claims to measure leadership potential but correlates 0.8 with plain self-esteem is measuring self-esteem. This is validity, and it is not a property an instrument either has or lacks — it is an argument built from accumulated evidence.
Compared to whom? A raw score of 34 means nothing on its own. It becomes interpretable only against a reference group: 34 is high relative to what population, measured when, and where? This is the work of norms, and it is the step most consumer tests skip entirely.
What good practice looks like
A well-built instrument has a paper trail. Someone published how the items were written and which ones were discarded. Someone reported the reliability coefficients, in named samples, with sample sizes. Someone tested whether the factor structure held up in a second, independent sample rather than the one used to build it. Someone checked whether the items behave the same across languages and age groups. Someone else — ideally with no stake in the instrument — tried to replicate the findings and reported what happened.
None of this makes a test infallible. It makes its limits knowable, which is a different and more useful property. An instrument whose measurement error is documented tells you how much to trust a small difference between two scores. An instrument with no published error estimate is not more accurate; it is just quieter about being inaccurate.
Where it leaves the reader
Psychometrics will not tell you that a test is true. It will tell you how far a score can be pushed before it stops meaning anything — which questions it can answer, which populations it was calibrated on, and how large a difference has to be before it is worth noticing.
That is a narrower promise than most personality products make. It is also the only promise that survives contact with the evidence.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.