How a psychometric test is actually built
From construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
A psychometric instrument that reaches publication has usually been through five or six years of work and thrown away most of what was written for it. The sequence is worth knowing, because every stage of it corresponds to a way a test can be inadequate.
1. Define the construct
Before any item is written, the construct has to be specified precisely enough to say what falls outside it. This means a definition, a statement of dimensional structure — one dimension or several, and if several, how they relate — and an explicit account of which neighbouring constructs it is meant to be distinct from.
Skipping this stage is the origin of most redundant instruments. A scale built from an intuitive sense of a concept usually ends up measuring an established trait under a new name.
2. Generate an item pool
Items are drafted well in excess of what will survive — typically three to five times the final count. They come from theory, from existing instruments, from interviews with people who have the experience being described, and from subject-matter experts.
At this stage the pool is reviewed for content coverage (does it span the whole construct or cluster on one facet?), for reading level, for double-barrelled wording, and for items that are really asking about something else.
3. Pilot and cut
The pool goes to a first sample. The analysis is mostly destructive.
Items with almost no variance go — a question everyone answers the same way distinguishes nobody. Items that correlate poorly with the total score go. Items that cross-load onto more than one factor go. Items that behave differently across demographic groups for reasons unrelated to the trait — differential item functioning — go, and this check is one of the most commonly skipped.
What survives is typically a third of what was written.
4. Confirm the structure in a fresh sample
This is the step that separates serious instruments from the rest. Factor analysis on the development sample will recover the structure that was designed in — it has to, since the items were selected to produce it. The real test is confirmatory factor analysis in an independent sample.
A great many published scales recover their intended structure only in the data used to build them. When someone else collects new data, the factors do not appear. That failure is common enough that "was the structure confirmed in an independent sample?" is one of the most efficient questions to ask about any instrument.
5. Accumulate validity evidence
Correlations with established measures of the same construct. Correlations with constructs it should be distinct from, which had better be lower. Relationships with outcomes that matter. Incremental validity over what already exists. Test-retest stability over an interval appropriate to the construct.
This stage does not end. It continues for as long as the instrument is in use, and it is where instruments are most often found wanting years after publication.
6. Build norms
A reference sample large enough and representative enough for the intended population, stratified where the trait varies by age or sex, documented as to year and country. Then repeated periodically, because norms drift.
7. Translate, if it will be used elsewhere
Translation is not a linguistic exercise. Proper adaptation means forward translation, independent back translation, expert review, cognitive interviews with target-language speakers, and then a fresh psychometric study in the new language to test whether the structure holds and whether items function equivalently. An instrument translated without that study is an untested instrument in that language, whatever its credentials in the original.
What this implies about the quick quiz
A test written in an afternoon has completed step two. It may well produce a plausible-looking result, a percentile and a paragraph of description. The steps it skipped are the ones that would have established whether any of that means anything.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.