Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple-choice item-writing guidelines for classroom assessment. Applied Measurement in Education, 15(3), 309–333.
31 rules across 5 categories
Content Concerns — 4 Rules
1
Test important learning outcomes, not trivial content
Items targeting trivial facts waste testing time and reduce construct validity. Every item should measure a meaningful, defensible learning objective.
2
Base each item on a specific, defensible content source
Items without a traceable content source cannot survive content validity challenges in legal or accreditation review.
3
Present a single, clearly formulated problem
Double-barreled items confound measurement. If one part is known and the other is not, the response is uninterpretable.
4
Items must be independent — no inter-item cues
When one item gives away or depends on the answer to another, it introduces construct-irrelevant variance and violates local independence.
Formatting Concerns — 3 Rules
5
Format items vertically, not horizontally
Horizontal layouts increase cognitive load and reading errors, especially under timed conditions. Vertical formatting improves scanability.
6
Arrange options in logical order
Numerical, alphabetical, or chronological ordering eliminates positional bias and makes the item easier to process without affecting difficulty.
7
Avoid “All of the above” and “None of the above”
These options create logical flaws: partial knowledge can eliminate AOTA, and NOTA measures recognition of incorrectness rather than knowledge of the correct answer.
Style Concerns — 5 Rules
8
Keep reading level appropriate for the target population
Unnecessarily complex language measures reading ability rather than content knowledge, introducing construct-irrelevant difficulty.
9
Use positive phrasing — avoid NOT/EXCEPT unless essential
Negative stems are frequently misread under test pressure. When negatives are necessary, bold and capitalize them (e.g., NOT, EXCEPT).
10
Avoid absolute terms (always, never, all, none)
Test-wise examinees recognize absolutes as typically false, making the item easier for strategic guessers without content knowledge.
11
Ensure grammatical consistency between stem and all options
When only the key is grammatically correct with the stem, examinees can identify it without content knowledge — a construct-irrelevant cue.
12
Minimize wording in options — put shared information in the stem
Redundant text in options increases reading time and cognitive load without adding measurement value. Move common phrasing to the stem.
Writing the Stem — 3 Rules
13
The stem must contain the central idea
A well-focused stem lets the examinee formulate an answer before reading options, supporting construct-relevant cognitive processing.
14
Avoid irrelevant information and window dressing
Extraneous details increase reading time, penalize slower readers, and can introduce cultural or contextual bias without improving measurement.
15
Use a direct question or completion format, not a vague prompt
Vague stems like “Regarding X, which is true?” fail to frame a specific problem and allow poorly constructed distractors to go unnoticed.
Writing the Choices — 16 Rules
16
All options must be plausible to uninformed test-takers
Implausible distractors are never selected, effectively reducing the item to fewer options and inflating guessing probability.
17
There must be exactly one correct or best answer
Multiple defensible answers make scoring arbitrary and invite legal challenges. Subject-matter expert agreement on the key must be unambiguous.
18
Options should be homogeneous in content
Heterogeneous options let examinees eliminate outliers without content knowledge, reducing discrimination and introducing construct-irrelevant variance.
19
Options should be similar in length — avoid longest-is-correct cue
Item writers tend to qualify the correct answer more carefully, making it longer. Test-wise examinees exploit this pattern.
20
Avoid overlapping or all-inclusive options
When one option subsumes another, logical deduction (not content knowledge) can identify the answer, undermining construct validity.
21
Avoid clang associations (same word in stem and key)
Repeating a distinctive word from the stem in only the correct option provides a surface-level cue unrelated to content mastery.
22
Avoid specific determiners (always, never) in options
Options with absolutes are almost always distractors. Test-savvy examinees eliminate them immediately, reducing effective option count.
23
Avoid giving away the answer through grammatical structure
Article agreement (a/an), verb conjugation, or pronoun matching between stem and only one option signals the key without requiring knowledge.
24
Keep options mutually exclusive
Overlapping options create logical contradictions in scoring and confuse examinees who correctly identify the overlap.
25
Phrase options positively — avoid negatives in options
Negative options combined with negative stems create double negatives, dramatically increasing cognitive complexity without improving measurement.
26
Avoid trick questions
Trick items measure carefulness and test-taking skill rather than content knowledge, reducing construct validity and examinee trust in the assessment.
27
Avoid humor and cultural references
Humor can distract, offend, or confuse examinees from different cultural backgrounds, introducing bias and reducing the seriousness of the assessment.
28
Don’t use compound options (“A and B” or “B and C”)
Compound options allow partial knowledge to cue the answer: if a student knows one part is wrong, the compound is eliminated — or confirmed — without full knowledge.
29
Place the answer key randomly across positions
Non-random key placement (e.g., favoring “C”) creates positional bias and rewards guessing strategies over content knowledge.
30
Use “Which of the following” format when appropriate
This construction clearly signals that the answer is among the listed options and frames the task as selection, which aligns with construct-relevant processing.
31
Ensure item matches the test blueprint/specification
Items that drift from the blueprint compromise content validity and can create gaps or overrepresentation in the construct being measured.
Framework 2 of 7
Bloom’s Revised Taxonomy
Anderson, L. W., & Krathwohl, D. R. (Eds.). (2001). A taxonomy for learning, teaching, and assessing: A revision of Bloom’s taxonomy of educational objectives. Longman.
Assessment check: Verify the item genuinely requires retrieval, not recognition cued by stem wording. Flag items that test recall of trivial or isolated facts.
2
Understand
Construct meaning from instructional messages: interpret, exemplify, classify, summarize, infer, compare, explain.
Assessment check: Confirm the item requires paraphrasing or restating — not verbatim recognition. Distractors should represent common misunderstandings.
3
Apply
Carry out or use a procedure in a given situation: execute, implement, solve, demonstrate, use.
Assessment check: Ensure the item presents a novel scenario requiring procedure application, not a rehearsed example from training materials.
4
Analyze
Break material into constituent parts and determine relationships: differentiate, organize, attribute, deconstruct.
Assessment check: Verify the item requires examinees to distinguish relevant from irrelevant information or identify cause-effect relationships.
5
Evaluate
Make judgments based on criteria and standards: check, critique, judge, justify, assess, defend.
Assessment check: Confirm the item requires criterion-based judgment, not mere preference. The evaluation criteria should be inferable from training content.
6
Create
Put elements together to form a coherent whole or produce an original product: generate, plan, produce, design, construct.
Assessment check: Flag if claimed as Create-level but actually tests lower-order skills. True creation is difficult to assess via multiple-choice format.
Quality Dimensions Checked Per Level — 6 Checks
A
Cognitive level classification accuracy
Verify the item actually measures the claimed Bloom's level, not a lower-order proxy. Misclassification undermines blueprint alignment.
B
Stem-to-level alignment
The stem verb and structure must elicit the targeted cognitive process. “Which is true?” rarely rises above Remember, regardless of content complexity.
C
Distractor cognitive plausibility
Distractors should represent errors at the targeted level. A Remember-level distractor on an Analyze-level item provides no diagnostic information.
D
Scenario authenticity for higher-order items
Apply, Analyze, Evaluate, and Create items require realistic contexts. Abstract or decontextualized stems reduce these to lower-order recall.
E
Knowledge dimension alignment
Cross-reference the cognitive process with the knowledge type (factual, conceptual, procedural, metacognitive) to ensure the item maps to the correct taxonomy cell.
F
Level-appropriate difficulty calibration
Higher Bloom's levels should generally be more difficult. A “Create”-level item that is trivially easy likely misclassifies the cognitive demand.
Framework 3 of 7
Item Response Theory — Three-Parameter Logistic (3PL) Model
Lord, F. M. (1980). Applications of item response theory to practical testing problems. Erlbaum. • Baker, F. B., & Kim, S.-H. (2004). Item response theory: Parameter estimation techniques (2nd ed.). Marcel Dekker.
18 checks — 3 parameters × 6 quality criteria
a
Discrimination
How well the item differentiates between examinees of different ability levels. The slope of the item characteristic curve at the inflection point.
Acceptable: 0.80 – 2.50
b
Difficulty
The ability level at which an examinee has a 50% probability of answering correctly (adjusted for guessing). Measured on the theta scale.
Acceptable: −3.00 – +3.00
c
Pseudo-Guessing
The lower asymptote — the probability that very low-ability examinees answer correctly by guessing. Should approximate 1/k for k options.
Acceptable: 0.00 – 0.35
Quality Criteria Checked Per Parameter — 6 Checks Each
1
Parameter within acceptable range
Out-of-range parameters indicate items that are too easy, too hard, non-discriminating, or susceptible to excessive guessing.
2
Standard error of estimate
Large standard errors indicate unstable parameter estimates, often due to insufficient sample size or model-data misfit.
3
Model-data fit (chi-square residual)
Significant misfit between observed and expected response patterns suggests the 3PL model does not adequately describe item behavior.
4
Information function contribution
Items should contribute meaningful information at the target ability range. Items with low information are measurement deadweight.
5
Distractor response curve analysis
Each distractor should show decreasing selection probability as ability increases. Non-functional distractors indicate implausible options.
6
Parameter stability across samples
IRT parameters should be sample-invariant. Instability across subgroups suggests DIF or multidimensionality concerns.
Framework 4 of 7
AERA/APA/NCME Standards for Educational and Psychological Testing
American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). Standards for educational and psychological testing. AERA.
38 checks across 3 chapters
Chapter 1: Validity — 14 Standards
1.0
Clear articulation of proposed score interpretations
Validity is not a property of the test but of the score interpretation. Each use case requires its own validity argument.
1.1
Evidence of content validity through domain representation
Items must proportionally represent the content domain as defined by the test blueprint.
1.2
Evidence of response processes
The cognitive processes examinees actually use should match those the item intends to measure.
1.3
Internal structure evidence
Item intercorrelations and factor structure should be consistent with the construct being measured.
1.4
Relations to external variables
Scores should correlate as expected with external criteria, other tests of the same construct, and tests of different constructs.
1.5
Consequences of testing
Evidence should address both intended and unintended consequences of score use, including impact on subgroups.
1.6
Construct definition clarity
The construct must be defined precisely enough to guide item development and evaluate whether items are on-construct.
1.7
Construct-irrelevant variance minimization
Items should not measure abilities or traits outside the intended construct (e.g., reading ability on a math test).
1.8
Construct underrepresentation avoidance
The test must not systematically omit important facets of the construct, which would narrow the inference.
1.9
Alignment between item format and construct
The chosen item format (MC, constructed-response, performance) must be capable of eliciting the targeted knowledge or skill.
1.10
Expert judgment of item-objective congruence
Subject matter experts must agree that each item measures its intended objective at the intended cognitive level.
1.11
Cross-population validity generalization
Validity evidence should extend to all populations for which score interpretations are intended.
1.12
Documentation of validity evidence
All validity evidence must be documented and made available for review by qualified professionals.
1.13
Ongoing validity evidence collection
Validity is not established once — it requires ongoing data collection as the construct, population, and use context evolve.
Chapter 2: Reliability/Precision — 12 Standards
2.0
Appropriate reliability coefficient reported
Internal consistency (alpha, omega), test-retest, or alternate-form reliability must be reported for the intended use context.
2.1
Standard error of measurement reported
SEM quantifies score precision and is essential for constructing confidence intervals around individual scores.
2.2
Conditional SEM at cut scores
Reliability matters most at decision points. SEM should be evaluated specifically at pass/fail boundaries.
2.3
Item-total correlation adequacy
Each item should positively correlate with the total score. Negative or near-zero correlations indicate miskeyed or off-construct items.
Network drops during testing must not result in lost responses or forced restarts, which disadvantage affected examinees.
16
Timer accuracy verified across browsers
Timer discrepancies across browsers or devices can give some examinees more or less time than intended.
17
Navigation controls are clear and consistent
Confusing navigation increases construct-irrelevant variance and test anxiety, particularly for less tech-savvy examinees.
18
Review and flag functionality available
Examinees should be able to mark items for review and revisit them, mimicking the strategic test-taking available in paper formats.
19
Progress indicators displayed
Examinees need to manage their time effectively; hiding progress information introduces construct-irrelevant anxiety and poor time allocation.
Data Quality & Scoring — 5 Checks
20
Automated scoring accuracy verified
Automated scoring must exactly match the answer key; even rare misscoring undermines the entire assessment’s credibility.
21
Response data completeness validated
Missing or corrupted response records lead to inaccurate scoring and invalid psychometric analyses.
22
Item metadata captured accurately
Bloom's level, content domain, key, and item parameters must be stored with each item for psychometric analysis and bank management.
23
Score scale consistency maintained
Scores from different forms or administrations must be equated to a common scale for fair cross-administration comparisons.
24
Audit trail for score calculations
A complete, reproducible record of how each score was calculated is essential for defending score validity in appeals or litigation.
Framework 6 of 7
Differential Item Functioning (DIF) Analysis
Holland, P. W., & Wainer, H. (Eds.). (1993). Differential item functioning. Erlbaum. • Zumbo, B. D. (1999). A handbook on the theory and methods of differential item functioning. National Defense Headquarters.
Detects items that function differently for male and female examinees after controlling for overall ability level.
🌐
Cultural DIF
Identifies items where cultural background knowledge provides an advantage or disadvantage unrelated to the construct.
📍
Geographic DIF
Flags items requiring region-specific knowledge (laws, climate, customs) that is not part of the intended construct.
💬
Language DIF
Detects items where linguistic complexity or idiomatic expressions disadvantage non-native speakers beyond intended difficulty.
💰
Socioeconomic DIF
Identifies items that assume specific socioeconomic experiences (travel, technology access, lifestyle) irrelevant to the construct.
♿
Disability DIF
Ensures items are accessible and function equivalently for examinees with disabilities when accommodations are provided.
Analytic Methods Applied — 6 Methods Per Dimension
M1
Mantel-Haenszel (MH) chi-square statistic
The most widely used DIF detection method. Compares odds ratios of correct response across groups at each ability level. Classifies DIF as A (negligible), B (moderate), or C (large).
M2
IRT-based likelihood ratio test
Compares item parameters estimated separately for reference and focal groups. More powerful than MH for detecting non-uniform DIF.
M3
Logistic regression DIF detection
Models the probability of correct response as a function of ability, group membership, and their interaction. Detects both uniform and non-uniform DIF simultaneously.
M4
Standardized mean difference (SMD) effect size
Provides an interpretable effect size metric for DIF magnitude, complementing statistical significance with practical significance.
M5
Expert sensitivity review panel
Statistical DIF methods alone cannot identify the source of bias. Expert panels review flagged items for content-based explanations of differential functioning.
Examines whether specific distractors attract different groups at differential rates, which can reveal culturally loaded wrong answers even when the key shows no DIF.
Framework 7 of 7
Legal & Regulatory Compliance
Multiple regulatory frameworks governing the development, administration, and defensibility of credentialing and certification assessments.