What is Cronbach's alpha — and what it is not
The most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.
Cronbach's alpha is a coefficient of internal consistency: the degree to which the items on a scale correlate with each other. Published in 1951, it is the single most reported psychometric statistic, and its ubiquity has outrun its usefulness.
Alpha runs from 0 to 1 and is driven by two things: the average correlation among items, and the number of items. That second dependency is the source of most of the misunderstanding.
What a high alpha does not mean
It does not mean the scale measures one thing. This is the most common error. Alpha can be high for a scale with two or three distinct factors, provided the items correlate enough overall. Unidimensionality has to be established by factor analysis; alpha cannot establish it and was never meant to.
It does not mean the scale is valid. Alpha says nothing about what the items measure. Twenty questions about shoe size would have excellent internal consistency and zero validity as a measure of anything psychological.
It does not mean the scale is good. Because alpha rises with item count, a long scale of mediocre items will out-score a short scale of excellent ones. A forty-item scale with an average inter-item correlation of 0.2 lands above 0.90.
A very high alpha is often a warning. Above roughly 0.95, the items are usually near-paraphrases. That buys consistency at the price of construct coverage: the scale measures one narrow facet very precisely and misses the rest of the construct. This is called attenuation paradox, and it is why more consistency is not monotonically better.
What it assumes
Alpha is a lower bound on reliability under a specific assumption — tau-equivalence, meaning every item contributes equally to the underlying trait. Real scales almost never satisfy this. Items differ in how strongly they load on the factor, and when they do, alpha underestimates reliability.
McDonald's omega drops the equal-contribution assumption and estimates reliability from the actual factor loadings. Methodologists have recommended it over alpha for two decades. Reporting practice has been slow to follow, largely because alpha is a single line in every statistics package and omega is not.
Reading it in practice
Rough conventions: 0.70 and above is acceptable for research comparing groups, 0.80 and above for applied use, 0.90 and above where a score affects an individual decision. Treat these as orientation, not thresholds — the required level depends on the stakes of the decision the score informs.
The more informative figures are usually elsewhere in the same paper. The average inter-item correlation (0.15 to 0.50 is a healthy range) tells you whether high alpha came from item quality or item quantity. The item count tells you the same thing from the other direction. Test-retest reliability tells you something alpha structurally cannot: whether the score is stable in the same person over time.
A scale reporting alpha alone has reported the easiest of the reliability figures to obtain, and the least demanding one to pass.
Put it to the test
Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.
Keep reading
- What is psychometrics?The discipline that works out whether a psychological measurement is any good — and the reason two tests asking similar questions can differ enormously in what their results are worth.
- What is reliability in psychological testing?Reliability is consistency of measurement, not correctness. A test can be highly reliable and still measure the wrong thing — which is why it is the first question asked and never the last.
- What is validity in psychological testing?Validity is not a property a test has. It is an argument about whether a particular interpretation of a particular score, for a particular purpose, is defensible — and it has to be rebuilt for every new use.
- What is construct validity?The hardest question in measurement: whether the thing you are measuring exists as you have defined it, and whether your instrument reaches it rather than something adjacent.
- Convergent and discriminant validityTwo mirror-image requirements: a measure must agree with what it should agree with, and must stay distinct from what it claims not to be. Most weak instruments fail the second one.
- Standard error of measurement: how much your score could be offEvery score carries a margin of error, and for most psychological tests it is wider than people assume. This is the number that tells you when a difference is real and when it is noise.
- Norms and percentiles: what your score is compared againstA raw score is uninterpretable on its own. It becomes meaningful only against a reference group — and which group was used is one of the most consequential and least advertised facts about any test.
- How a psychometric test is actually builtFrom construct definition to published norms, the sequence takes years and discards most of what goes into it. Knowing the steps makes it obvious which ones a quick online quiz skipped.
- What does it mean when a test is called validated?The word is unregulated and used freely by products that have done none of the work. Here is what it means when it means something, and the four questions that separate the two cases.
- Why most internet personality tests are not reliableNot because they are dishonest, but because the steps that make a measurement trustworthy are invisible, expensive and easy to skip — and skipping them changes nothing about how the result looks.
- What self-report can and cannot measureQuestionnaires ask you to be the observer of yourself, which works better for some things than others. Knowing where the method is strong and where it fails changes how you read any result.
- What it means when a test predicts somethingA correlation of 0.30 is a strong finding in personality research and a weak basis for a decision about one person. Both statements are true, and the gap between them causes most misuse of test results.
- Screening is not diagnosisA screening questionnaire is built to catch as many possible cases as it can, accepting a high false-positive rate as the price. Reading a screening score as a diagnosis inverts what it was designed to do.