noesisnoesis
Frameworks compared

Is the MBTI scientific?

The honest answer is more specific than yes or no: the questionnaire is reasonably well constructed, and the type theory it reports results in is not supported. Those two facts are usually collapsed into one argument.

The Myers-Briggs Type Indicator is the most widely taken personality instrument in the world and the one academic personality psychology takes least seriously. Both facts are worth explaining, because the popular versions of the argument on each side are wrong.

What the criticisms actually establish

The dichotomies are not dichotomous. This is the strongest objection and it is not really disputable. The MBTI reports you as an E or an I, a T or an F. If those were real categories, scores on each dimension would cluster into two humps with a gap between them. They do not. The distributions are unimodal — most people are near the middle, and the type boundary cuts through the densest part of the distribution.

The practical consequence is severe. Someone one point either side of the boundary receives a different letter and a different four-letter type, despite being statistically indistinguishable. This is not a subtle statistical quibble; it is why the same person can legitimately come out as two different types.

Test-retest classification is poor. Studies retesting after several weeks find that a substantial share of respondents — commonly reported around a third to half — receive at least one different letter, and therefore a different type. For an instrument whose whole output is a stable type, that is a serious problem. Note the mechanism: it follows directly from the dichotomisation. The underlying continuous scores are reasonably stable; it is the act of cutting them at a boundary that manufactures the instability.

The fourth dimension is weakly supported. Judging-Perceiving was Isabel Myers's addition to Jung, not part of the original theory, and it has the least empirical grounding of the four.

The theory is pre-empirical. Jung's type theory came from clinical observation and introspection in the 1920s. It was not derived from data and was not built to be tested. That is not a criticism of Jung, who was doing something else; it is a description of what kind of claim it is.

Predictive validity is weak where it is most claimed. The MBTI is heavily marketed for team composition and career guidance. Evidence that type predicts job performance or team effectiveness is thin, and the instrument's own publisher advises against using it for selection.

What the criticisms do not establish

The questionnaire itself is not badly built. The item-level psychometrics are respectable — internal consistency for the four scales is generally acceptable, and the items were developed with more care than the popular criticism implies.

More interestingly, the underlying continuous dimensions map recognisably onto established traits. E-I corresponds closely to extraversion. T-F corresponds substantially to agreeableness. J-P corresponds substantially to conscientiousness. S-N corresponds to openness. Three of the four are recovering real variance that the Five-Factor Model also captures.

So the measurement is not the problem. The problem is what is done with it afterwards: continuous scores are cut into categories, the categories are given names, and the names are treated as kinds of person.

Why it remains popular

It is pleasant. Every one of the sixteen types is described positively, none is a bad type to be, and the descriptions are written at a level of generality where most people recognise themselves — the Barnum effect, working exactly as it was documented to.

It also gives teams a shared, non-threatening vocabulary. Consultants report that this is genuinely useful in facilitation, and there is no reason to doubt them. It is a legitimate use that has nothing to do with whether the types are real.

The reasonable conclusion

Taking the MBTI is not harmful and the experience of reading your type description is not worthless. Using type as a basis for hiring, promotion, team assignment or career direction is not supported by evidence, and the publisher agrees.

If what you want from a personality instrument is a measurement you can defend, the Five-Factor Model is where the evidence sits: continuous scores rather than types, forty years of independent replication, published norms, and documented measurement error. NOESIS publishes Big Five instruments in the catalog with their sources and limitations stated, which is the comparison worth making for yourself.

Put it to the test

Reading about measurement is one thing. Seeing your own score reported with its source, its norm sample and its limits is another. Free to take, no signup required.

Keep reading