Screening is not diagnosis
A screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.
A depression screening questionnaire, an anxiety screener, an alcohol-use screener — these are among the best-validated instruments in existence, and they are routinely misread. A score above the cutoff does not mean you have the condition. It was never designed to.
The design trade-off
Every screening instrument sits on a trade-off between two errors: missing people who have the condition, and flagging people who do not. Screening tools are deliberately tuned towards the second. Missing a case of depression in primary care is costly; a false positive costs a conversation with a clinician.
The consequence follows arithmetically. A screener with 88 percent sensitivity and 85 percent specificity sounds excellent. Apply it in a population where 5 percent actually have the condition, and of every 100 people who screen positive, roughly 74 do not have it. The instrument is working exactly as specified. Most positives are still false.
This is base-rate arithmetic, and it is the reason universal screening in low-prevalence settings is contested even for well-validated instruments.
What a cutoff is
A cutoff is a decision threshold chosen for a purpose in a population, not a natural boundary. The conventional PHQ-9 cutoff of 10 was derived in populations with relatively high prevalence; applied elsewhere it behaves differently. Individual-participant meta-analysis has shown that published accuracy estimates for that cutoff were systematically inflated by selective reporting of optimised thresholds.
For some instruments there is deliberately no universal cutoff at all. The World Health Organization's SRQ-20 is used across dozens of countries, and its optimal threshold varies enough by country, language and sex that publishing one global number would do more harm than leaving it to local validation.
Diagnosis is a different procedure
A diagnosis requires a clinician establishing the presence and duration of symptoms, their functional impact, and the exclusion of alternative explanations — medical conditions, substances, other disorders that present similarly. It considers history and context. A questionnaire does none of this. It counts self-reported symptom frequency over a window.
Two people can produce identical scores from entirely different clinical pictures, since the total compresses several heterogeneous symptoms into one number. The score is a signal to look further; the looking is the diagnosis.
What the score is legitimately for
A prompt. A high score is a reason to talk to a professional. That is the intended use and it is a genuinely valuable one.
Tracking change. This is the underused function. Repeated measurement over the course of treatment shows whether something is working, and it is far more informative than any single score. Measurement-based care is built on this.
Population description. In research and public health, these instruments estimate symptom burden across groups, which is what they are best at.
Reading your own result
If you score above a threshold on a screener, the correct interpretation is: this warrants a conversation with someone qualified. Not: I have this condition.
If you score below, the correct interpretation is: this particular screen did not flag anything. Not: I am fine. Screeners miss cases, which is the other half of the trade-off, and a low score in someone who feels unwell is not reassurance.
Where NOESIS publishes screening instruments, the reports state the published thresholds, name the source, and say plainly that the result is not a diagnosis. That framing is not a legal formality; it is what the instruments themselves say about their own use.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- What is Cronbach's alpha — and what it is notThe most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.