noesisnoesis
How measurement works

What is Cronbach's alpha — and what it is not

The most reported statistic in psychological testing, and the most misread. Alpha does not tell you a scale is unidimensional, and a high value is often a symptom rather than a virtue.

Cronbach's alpha is a coefficient of internal consistency: the degree to which the items on a scale correlate with each other. Published in 1951, it is the single most reported psychometric statistic, and its ubiquity has outrun its usefulness.

Alpha runs from 0 to 1 and is driven by two things: the average correlation among items, and the number of items. That second dependency is the source of most of the misunderstanding.

What a high alpha does not mean

It does not mean the scale measures one thing. This is the most common error. Alpha can be high for a scale with two or three distinct factors, provided the items correlate enough overall. Unidimensionality has to be established by factor analysis; alpha cannot establish it and was never meant to.

It does not mean the scale is valid. Alpha says nothing about what the items measure. Twenty questions about shoe size would have excellent internal consistency and zero validity as a measure of anything psychological.

It does not mean the scale is good. Because alpha rises with item count, a long scale of mediocre items will out-score a short scale of excellent ones. A forty-item scale with an average inter-item correlation of 0.2 lands above 0.90.

A very high alpha is often a warning. Above roughly 0.95, the items are usually near-paraphrases. That buys consistency at the price of construct coverage: the scale measures one narrow facet very precisely and misses the rest of the construct. This is called attenuation paradox, and it is why more consistency is not monotonically better.

What it assumes

Alpha is a lower bound on reliability under a specific assumption — tau-equivalence, meaning every item contributes equally to the underlying trait. Real scales almost never satisfy this. Items differ in how strongly they load on the factor, and when they do, alpha underestimates reliability.

McDonald's omega drops the equal-contribution assumption and estimates reliability from the actual factor loadings. Methodologists have recommended it over alpha for two decades. Reporting practice has been slow to follow, largely because alpha is a single line in every statistics package and omega is not.

Reading it in practice

Rough conventions: 0.70 and above is acceptable for research comparing groups, 0.80 and above for applied use, 0.90 and above where a score affects an individual decision. Treat these as orientation, not thresholds — the required level depends on the stakes of the decision the score informs.

The more informative figures are usually elsewhere in the same paper. The average inter-item correlation (0.15 to 0.50 is a healthy range) tells you whether high alpha came from item quality or item quantity. The item count tells you the same thing from the other direction. Test-retest reliability tells you something alpha structurally cannot: whether the score is stable in the same person over time.

A scale reporting alpha alone has reported the easiest of the reliability figures to obtain, and the least demanding one to pass.

Put it to the test

Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.

Keep reading