Item Analysis: Improve Questions After Every Exam

Learn how item analysis helps teachers interpret question difficulty, discrimination and distractor performance, then improve a reusable question bank.

Item analysis in education turns student response data into a practical quality check for every question. Instead of judging an item only by how it reads, teachers can see how difficult it was, whether it separated stronger and weaker overall performance, and whether its wrong answer options revealed useful misconceptions.

This guide explains the core measures in plain language, works through a classroom example, and gives assessment teams a repeatable review process. The aim is not to let statistics make decisions for you. It is to combine evidence with subject expertise so your question bank improves after every assessment.

What does item analysis measure?

Item analysis examines how learners responded to individual questions and how each question behaved within the assessment. The University of Washington identifies two common measures: item difficulty and item discrimination. For multiple-choice questions, a distractor analysis adds another useful layer.

Measure Question it answers Possible follow-up
Difficulty index What proportion answered correctly? Check whether the level matched the learning objective.
Discrimination Did stronger overall performers tend to answer correctly? Review the key, wording, alignment and scoring.
Distractor performance Which wrong options were selected, and by whom? Improve weak options or investigate misconceptions.

These measures describe performance in a particular administration. They do not prove that a question is valid or invalid. Results can change with the learners, instruction, stakes, curriculum coverage and the rest of the assessment. Treat every flag as a prompt for review, not an automatic order to delete the item.

How to calculate difficulty and discrimination

Difficulty index

For a one-mark question with one correct answer, the difficulty index is the proportion of learners who answered correctly:

Difficulty index = number correct ÷ number who attempted the item

If 36 of 50 learners answered correctly, the index is 36 ÷ 50 = 0.72, or 72%. A higher value means the item was easier for that group. The name can feel backward, which is why some assessment teams call it the facility value instead.

There is no single ideal value for every purpose. A diagnostic quiz may deliberately include accessible questions that confirm essential knowledge. A selection test may need items that provide more separation. Start with the intended learning objective and the test blueprint, then ask whether the observed difficulty makes sense.

Discrimination index

Discrimination describes whether success on one item is associated with stronger performance on the rest of the assessment. Testing software often reports a point-biserial correlation. A positive value usually means higher-scoring learners were more likely to answer the item correctly. A value near zero suggests little relationship. A negative value deserves prompt investigation.

For a simple classroom calculation, divide learners into an upper and lower group of equal size:

Discrimination index = proportion correct in upper group − proportion correct in lower group

Suppose 18 of 20 learners in the upper group and 10 of 20 in the lower group answered correctly. The index is 0.90 − 0.50 = 0.40. That is a useful positive result because the item distinguished between the groups in the expected direction.

Do not apply one threshold blindly. The University of Washington’s ScorePak guidance labels values above .30 as good, .10 to .30 as fair and below .10 as poor, but it explicitly describes those classifications as arbitrary. A broad assessment that samples several topics and thinking skills may produce different patterns from a narrow, highly consistent test.

Read the measures together

A single number rarely explains why a question performed as it did. Read difficulty, discrimination, distractor choices and curriculum alignment together.

Observed pattern What it might mean What to inspect
Very easy, positive discrimination An essential concept was widely mastered. Keep if blueprint coverage requires it.
Very hard, positive discrimination A demanding but functioning question. Confirm that content and cognitive demand were taught.
Moderate difficulty, weak discrimination Ambiguity, guessing or mixed skills may be involved. Stem, key, options, curriculum tag and scoring.
Negative discrimination A wrong key, misleading wording or a plausible alternative answer may exist. Review immediately with a subject expert.
One distractor never selected The option may be implausible or visibly different. Replace it with a misconception-based option.

Consider a question answered correctly by 62% of learners with a discrimination value of −0.12. The difficulty alone looks reasonable, but the negative relationship is a warning. Check whether the keyed answer is correct, whether two options could be defended, and whether high-performing learners interpreted a nuance differently.

The reverse also matters. An item answered correctly by 95% may have weak discrimination simply because almost everyone mastered an essential prerequisite. Removing it could damage content coverage. The decision belongs to the assessment purpose, not the statistic alone.

Use distractor analysis to improve MCQs

Distractors are not filler. Each wrong option should be clearly incorrect while remaining plausible to a learner with a recognizable misconception. Cambridge Assessment notes that undesirable difficulty can come from ambiguity, grammatical clues or poorly designed options rather than the knowledge being assessed.

For every multiple-choice question, record how many learners chose each option. Then ask:

  • Did almost nobody select a distractor? It may be too obviously wrong.
  • Did many strong performers choose the same wrong option? It may be defensible, or the stem may be unclear.
  • Did one distractor attract learners with the same misconception? Keep it and use the result to guide teaching.
  • Did choices appear almost random? The content may not have been taught, the wording may be confusing, or learners may have guessed.

The ETS guidelines for valid and fair test development recommend examining item difficulty, discrimination, correlations and distractor performance. ETS also notes that distractor analysis can reveal a second plausible answer or an option that is unusually attractive to higher-ability test takers.

Use response patterns to revise the item, then review the new version before reuse. A changed distractor creates a new measurement condition, so retain a revision note rather than silently overwriting the history.

A repeatable item review workflow

  1. Preserve the administration record. Keep the paper version, answer key, learner group, date and any scoring changes.
  2. Calculate the core measures. Record difficulty, discrimination and the response count for every option.
  3. Flag questions for review. Prioritize negative discrimination, challenged keys, unexpected difficulty and weak distractors.
  4. Review with context. Compare the item with the learning objective, curriculum coverage, teaching sequence and blueprint.
  5. Choose an action. Keep, revise, reclassify or retire the item. Never delete history that is needed to understand past papers.
  6. Document the reason. Add a short note such as “key corrected after expert review” or “difficulty retained because this is an essential prerequisite.”
  7. Verify before reuse. Return revised items to Draft or Review status until a second person checks them.

This routine works best when your question bank stores the item, answer, tags, review status and usage history together. The guidance in How to Build a Reusable Question Bank explains how to establish that structure before you begin collecting statistics.

Examiar can help you organize reviewed questions by curriculum, topic, type and difficulty, then reuse those pools in balanced exam patterns. Item statistics can be maintained as part of your review evidence while your team keeps one controlled source for approved content. Explore the Examiar assessment workflow or start with a free account.

Frequently asked questions

How many student responses are needed for item analysis?

Larger groups generally produce more stable estimates, but there is no universal minimum that makes every statistic dependable. A classroom teacher can still use small-group results as clues, provided each decision is checked against the item content and repeated evidence from future uses.

Should every difficult question be revised?

No. A difficult question may appropriately assess advanced content or higher-order thinking. Revise it when the difficulty is caused by ambiguity, an incorrect key, poor alignment or unnecessary reading load rather than the intended skill.

What should I do with a negatively discriminating item?

Check the answer key first. Then inspect alternative answers, wording, scoring, curriculum alignment and whether the item measured a different skill from the rest of the assessment. Do not use the item again until the cause has been reviewed.

Is item analysis the same as test validity?

No. Item analysis describes response patterns and internal relationships. Validity requires evidence that score interpretations and uses are appropriate for the assessment’s purpose. Statistics support that judgment, but they do not replace it.

Turn the framework into a workflow

Build your question bank and first exam.

Explore Examiar with your own curriculum and questions. No credit card required.