noesisnoesis
How measurement works

Why most internet personality tests are not reliable

Not because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.

A free personality test online and a validated psychometric instrument produce output that looks identical: a score, a percentile, a paragraph describing you. The difference is entirely in work you cannot see from the results screen.

This is worth being precise about, because the usual framing — that internet tests are scams — is mostly wrong and not very useful. Most are made by people with no intent to deceive. They are unreliable for structural reasons.

What typically gets skipped

Item analysis. In a properly built scale, most drafted items are discarded after piloting: items nobody varies on, items that correlate poorly with the total, items that load on more than one factor, items that behave differently across age or language groups. A test written and published without a pilot has kept everything, including the items that do not work.

Independent structural confirmation. Factor analysis on the data you built the scale from will recover the structure you designed. Confirming it in a fresh sample is a separate study, and it is the step where a large share of scales fail.

Norms. The percentile has to come from somewhere. If the site does not state its reference population, the number is derived either from previous visitors — a self-selected group of unknown composition — or from nothing in particular.

Test-retest data. Establishing that people score the same next month requires finding them again. Almost nobody does this, which means the stability that the word "trait" implies has not been demonstrated.

The design incentives run the wrong way

A test built for engagement is optimised for a different outcome than a test built for accuracy, and the two pull apart in predictable ways.

Flattering results perform better. A result that tells you something unwelcome gets shared less, so results drift towards descriptions everyone recognises in themselves. This is the Barnum effect, demonstrated in the 1940s and reliably reproduced since: people rate generic personality descriptions as highly accurate about themselves specifically.

Types perform better than dimensions. "You are an Architect" is more shareable than "you are at the 68th percentile on openness with a wide confidence interval". Types also conceal measurement error, since someone one point either side of a category boundary receives entirely different labels.

Short performs better than adequate. Every additional question loses respondents, so tests get cut to a length that maximises completion rather than one that covers the construct.

None of these are lies. They are optimisation for a metric other than accuracy.

How to tell in about thirty seconds

Look for a named source instrument with authors and a year. Look for a statement of what population the norms come from. Look for any acknowledgement of limitations or measurement error. Look for whether the result is a continuous score or a category. Look for whether the same page also promises to identify your ideal career, partner or life purpose from twenty questions.

A test that names its instrument, cites its source, states its norms and says what it cannot tell you has done the work. One that does none of that may still be a pleasant way to spend five minutes — but the number it produces is decoration.

Every instrument in the NOESIS catalog states which published test it implements, who wrote it, in what year, how it is scored, and what its results do not establish. That last part is the one worth comparing against.

Put it to the test

Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.

Keep reading

Why most internet personality tests are not reliable | NOESIS