Item Response Theory
Lord, F. M. · 1980
別名: IRT · Rasch · Rasch Measurement Theory · Latent Trait Theory
Item Response Theory is a family of mathematical models that describe the relationship between latent traits (abilities, attitudes) and responses to test items, enabling sample-free item calibration and test-free person measurement. Unlike Classical Test Theory, IRT models item-level characteristics (difficulty, discrimination, guessing) and person ability on a common scale, providing measurement precision estimates at each ability level. IRT underpins modern computerized adaptive testing, large-scale educational assessments (PISA, GRE, GMAT), and clinical patient-reported outcome measures.
Methodological (not trait-content) framework modeling the probability of a response as a function of latent trait level and item parameters (difficulty, discrimination, guessing). Backs adaptive/CAT administration rather than any single instrument's content.
- 出典:
- Lord, F. M. (1980). Applications of Item Response Theory to Practical Testing Problems.
エビデンス概要
- 実証的サポート
- 強力なエビデンス
- 再現性
- 十分に再現
- 異文化
- 普遍的
- メタ分析
- 10件索引済み
これらの評価は、理論に関する公開文献を要約したもので、それに基づいて構築された個々のテストの品質ではありません。以下に記載された出典に基づく編集判断であり、エビデンスが蓄積されるにつれて変化します。
歴史的背景
IRT's foundations were laid independently by Georg Rasch in Denmark (1960) and Frederic Lord and Allan Birnbaum in the United States (1960s), with Lord and Novick's 1968 'Statistical Theories of Mental Test Scores' providing the mathematical groundwork. Lord's 1980 'Applications of Item Response Theory to Practical Testing Problems' consolidated the field by demonstrating practical applications in test construction, equating, and bias detection. The theory gained widespread adoption in the 1980s-1990s with increasing computational power, making parameter estimation feasible for large-scale testing programs.
構成概念
まだ構成概念が文書化されていません。
機器
まだこの理論に基づく機器がありません。
この理論に基づくテスト
まだこの理論に基づくテストがありません。
実践的応用
IRT is the foundation of computerized adaptive testing (CAT), used in high-stakes assessments such as the GRE, GMAT, NCLEX nursing licensure exam, and the CAT-ASVAB military battery, enabling shorter yet more precise tests by selecting items tailored to each examinee's ability level. In clinical settings, IRT powers the PROMIS (Patient-Reported Outcomes Measurement Information System) initiative, enabling efficient measurement of health outcomes. IRT methods also support test equating across forms, detection of differential item functioning (DIF) for fairness analysis, and item bank development for standardized testing programs worldwide.
測定方法
IRT is itself a measurement methodology rather than a substantive psychological theory. Common models include the one-parameter (1PL/Rasch), two-parameter (2PL), and three-parameter (3PL) logistic models for dichotomous items, and the Graded Response Model and Partial Credit Model for polytomous items. Parameter estimation uses maximum likelihood or Bayesian methods.
批評と制限事項
A key criticism is the strong parametric assumptions required by IRT models (such as unidimensionality and local independence), which may not hold for complex psychological constructs, leading to model misfit. The Rasch model community and the broader IRT community disagree fundamentally on whether the model should be fit to data (IRT approach) or data should meet model requirements for fundamental measurement (Rasch approach). IRT requires relatively large sample sizes for stable parameter estimation (typically 200-1000+ examinees depending on the model), limiting its applicability for small-scale assessments.
主要文献
- Lord, F. M.. (1980). Applications of Item Response Theory to Practical Testing Problems.
- Lord, F. M. & Novick, M. R.. (1968). Statistical Theories of Mental Test Scores.
- Rasch, G.. (1960). Probabilistic Models for Some Intelligence and Attainment Tests.
- Hambleton, R. K., Swaminathan, H. & Rogers, H. J.. (1991). Fundamentals of Item Response Theory.
- Embretson, S. E. & Reise, S. P.. (2000). Item Response Theory for Psychologists. 10.4324/9781410605269