Item Response Theory
Lord, F. M. · 1980
Also known as: IRT · Rasch
Item Response Theory is a family of mathematical models that describe the relationship between latent traits (abilities, attitudes) and responses to test items, enabling sample-free item calibration and test-free person measurement. Unlike Classical Test Theory, IRT models item-level characteristics (difficulty, discrimination, guessing) and person ability on a common scale, providing measurement precision estimates at each ability level. IRT underpins modern computerized adaptive testing, large-scale educational assessments (PISA, GRE, GMAT), and clinical patient-reported outcome measures.
Methodological (not trait-content) framework modeling the probability of a response as a function of latent trait level and item parameters (difficulty, discrimination, guessing). Backs adaptive/CAT administration rather than any single instrument's content.
- Source:
- Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems.
Historical Context
IRT's foundations were laid independently by Georg Rasch in Denmark (1960) and Frederic Lord and Allan Birnbaum in the United States (1960s), with Lord and Novick's 1968 'Statistical Theories of Mental Test Scores' providing the mathematical groundwork. Lord's 1980 'Applications of Item Response Theory to Practical Testing Problems' consolidated the field by demonstrating practical applications in test construction, equating, and bias detection. The theory gained widespread adoption in the 1980s-1990s with increasing computational power, making parameter estimation feasible for large-scale testing programs.
Constructs
No constructs documented yet.
Instruments
No instruments built on this theory yet.
Tests built on this theory
No tests built on this theory yet.
Practical Applications
IRT is the foundation of computerized adaptive testing (CAT), used in high-stakes assessments such as the GRE, GMAT, NCLEX nursing licensure exam, and the CAT-ASVAB military battery, enabling shorter yet more precise tests by selecting items tailored to each examinee's ability level. In clinical settings, IRT powers the PROMIS (Patient-Reported Outcomes Measurement Information System) initiative, enabling efficient measurement of health outcomes. IRT methods also support test equating across forms, detection of differential item functioning (DIF) for fairness analysis, and item bank development for standardized testing programs worldwide.
How It's Measured
IRT is itself a measurement methodology rather than a substantive psychological theory. Common models include the one-parameter (1PL/Rasch), two-parameter (2PL), and three-parameter (3PL) logistic models for dichotomous items, and the Graded Response Model and Partial Credit Model for polytomous items. Parameter estimation uses maximum likelihood or Bayesian methods.
Critiques & Limitations
A key criticism is the strong parametric assumptions required by IRT models (such as unidimensionality and local independence), which may not hold for complex psychological constructs, leading to model misfit. The Rasch model community and the broader IRT community disagree fundamentally on whether the model should be fit to data (IRT approach) or data should meet model requirements for fundamental measurement (Rasch approach). IRT requires relatively large sample sizes for stable parameter estimation (typically 200-1000+ examinees depending on the model), limiting its applicability for small-scale assessments.
Key Publications
- Lord, F. M.. (1980). Applications of Item Response Theory to Practical Testing Problems.
- Lord, F. M. & Novick, M. R.. (1968). Statistical Theories of Mental Test Scores.
- Rasch, G.. (1960). Probabilistic Models for Some Intelligence and Attainment Tests.
- Hambleton, R. K., Swaminathan, H. & Rogers, H. J.. (1991). Fundamentals of Item Response Theory.
- Embretson, S. E. & Reise, S. P.. (2000). Item Response Theory for Psychologists. 10.4324/9781410605269