Methodology

How Item Signal combines decades of psychometric research with modern AI to validate test items — transparently and reproducibly.

1. The Rule Engine

Our quality analysis is a coded implementation of Haladyna, Downing & Rodriguez (2002), AERA/APA/NCME Standards (2014), and ITC/ATP Guidelines (2022). 57 deterministic rules run first — no LLM, no hallucination. Each rule maps to a specific citation and produces a finding with severity, span location, and suggested fix.

2. AI-Assisted Analysis

For subjective assessments — Bloom's taxonomy classification, nuanced DIF analysis, rewrite suggestions — we use Claude (Anthropic) as a second-pass reviewer. The AI operates within a structured system prompt requiring JSON output, confidence intervals, and evidence-based reasoning. The AI never overrides a deterministic rule.

3. Difficulty Prediction

Text-only item difficulty prediction was validated as a BEA 2024 NAACL shared task using USMLE items (Tack et al., 2024). We position our predictions as pre-field triage signals with 80% confidence intervals — not as substitutes for field-test calibration. Current v0 baseline: RMSE ≤ 0.32 on held-out data.

4. The Evidence Packet

For batch engagements, we generate a 15–25 page branded PDF report covering: executive summary, methodology citations, per-item findings with prediction cards, aggregate statistics, and an appendix of cited works.

5. What's Orthodox vs. Frontier

Orthodox — Decades of Research Frontier — 2024–2025 Research
Haladyna item-writing rules Text-only difficulty prediction
Item Response Theory (IRT) AI contamination detection
Classical Test Theory (CTT) LLM-assisted rewrite suggestions
Differential Item Functioning (DIF)
Bloom's Taxonomy classification
Standard setting methods

6. Cited Works