Item Response Theory
Lord, F. M. · 1980
也被称为: IRT · Rasch · Rasch Measurement Theory · Latent Trait Theory
Item Response Theory is a family of mathematical models that describe the relationship between latent traits (abilities, attitudes) and responses to test items, enabling sample-free item calibration and test-free person measurement. Unlike Classical Test Theory, IRT models item-level characteristics (difficulty, discrimination, guessing) and person ability on a common scale, providing measurement precision estimates at each ability level. IRT underpins modern computerized adaptive testing, large-scale educational assessments (PISA, GRE, GMAT), and clinical patient-reported outcome measures.
Methodological (not trait-content) framework modeling the probability of a response as a function of latent trait level and item parameters (difficulty, discrimination, guessing). Backs adaptive/CAT administration rather than any single instrument's content.
- 来源:
- Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems.
证据概览
- 实证支持
- 强证据支持
- 重复验证
- 已充分重复验证
- 跨文化性
- 具普遍性
- 綜合分析(统合分析)
- 已收录 10 项
这些评等总结了关於该理論已发表的文獻,而非以此理論建构的任何单一测验的品质。它们是根据下方所列来源所作的編輯判斷,並會隨著证据累积而變化。
历史背景
IRT's foundations were laid independently by Georg Rasch in Denmark (1960) and Frederic Lord and Allan Birnbaum in the United States (1960s), with Lord and Novick's 1968 'Statistical Theories of Mental Test Scores' providing the mathematical groundwork. Lord's 1980 'Applications of Item Response Theory to Practical Testing Problems' consolidated the field by demonstrating practical applications in test construction, equating, and bias detection. The theory gained widespread adoption in the 1980s-1990s with increasing computational power, making parameter estimation feasible for large-scale testing programs.
构念
尚无构念文档记录。
工具
尚无基于此理论的工具。
基于此理论的测试
尚无基于此理论的测试。
实际应用
IRT is the foundation of computerized adaptive testing (CAT), used in high-stakes assessments such as the GRE, GMAT, NCLEX nursing licensure exam, and the CAT-ASVAB military battery, enabling shorter yet more precise tests by selecting items tailored to each examinee's ability level. In clinical settings, IRT powers the PROMIS (Patient-Reported Outcomes Measurement Information System) initiative, enabling efficient measurement of health outcomes. IRT methods also support test equating across forms, detection of differential item functioning (DIF) for fairness analysis, and item bank development for standardized testing programs worldwide.
如何测量
IRT is itself a measurement methodology rather than a substantive psychological theory. Common models include the one-parameter (1PL/Rasch), two-parameter (2PL), and three-parameter (3PL) logistic models for dichotomous items, and the Graded Response Model and Partial Credit Model for polytomous items. Parameter estimation uses maximum likelihood or Bayesian methods.
批评与局限
A key criticism is the strong parametric assumptions required by IRT models (such as unidimensionality and local independence), which may not hold for complex psychological constructs, leading to model misfit. The Rasch model community and the broader IRT community disagree fundamentally on whether the model should be fit to data (IRT approach) or data should meet model requirements for fundamental measurement (Rasch approach). IRT requires relatively large sample sizes for stable parameter estimation (typically 200-1000+ examinees depending on the model), limiting its applicability for small-scale assessments.
重要文献
- Lord, F. M.. (1980). Applications of Item Response Theory to Practical Testing Problems.
- Lord, F. M. & Novick, M. R.. (1968). Statistical Theories of Mental Test Scores.
- Rasch, G.. (1960). Probabilistic Models for Some Intelligence and Attainment Tests.
- Hambleton, R. K., Swaminathan, H. & Rogers, H. J.. (1991). Fundamentals of Item Response Theory.
- Embretson, S. E. & Reise, S. P.. (2000). Item Response Theory for Psychologists. 10.4324/9781410605269