How Item Signal combines decades of psychometric research with modern AI to validate test items — transparently and reproducibly.
Our quality analysis is a coded implementation of Haladyna, Downing & Rodriguez (2002), AERA/APA/NCME Standards (2014), and ITC/ATP Guidelines (2022). 57 deterministic rules run first — no LLM, no hallucination. Each rule maps to a specific citation and produces a finding with severity, span location, and suggested fix.
For subjective assessments — Bloom's taxonomy classification, nuanced DIF analysis, rewrite suggestions — we use Claude (Anthropic) as a second-pass reviewer. The AI operates within a structured system prompt requiring JSON output, confidence intervals, and evidence-based reasoning. The AI never overrides a deterministic rule.
Text-only item difficulty prediction was validated as a BEA 2024 NAACL shared task using USMLE items (Tack et al., 2024). We position our predictions as pre-field triage signals with 80% confidence intervals — not as substitutes for field-test calibration. Current v0 baseline: RMSE ≤ 0.32 on held-out data.
For batch engagements, we generate a 15–25 page branded PDF report covering: executive summary, methodology citations, per-item findings with prediction cards, aggregate statistics, and an appendix of cited works.
| Orthodox — Decades of Research | Frontier — 2024–2025 Research |
|---|---|
| Haladyna item-writing rules | Text-only difficulty prediction |
| Item Response Theory (IRT) | AI contamination detection |
| Classical Test Theory (CTT) | LLM-assisted rewrite suggestions |
| Differential Item Functioning (DIF) | |
| Bloom's Taxonomy classification | |
| Standard setting methods |