Diagnostic Questions: How Wrong Answers Reveal Mental Models
A right answer tells you a learner got there. A well-designed wrong answer tells you exactly where they got lost — and which mental model led them there.
Every teacher has met the moment: a learner hands in a wrong answer and the mark book records a zero, but the why stays invisible. A score punishes the mistake without ever naming it. Yet decades of research show that a wrong answer, when the question is built correctly, is one of the richest signals in education — it can reveal the exact mental model a learner is using. That is the idea behind the diagnostic question, and it quietly underpins much of modern learning science.
What is a diagnostic question?
A diagnostic question looks like an ordinary multiple-choice item, but it is engineered in reverse. Instead of padding the question with throwaway wrong options, every distractor is written to match a specific, well-documented misconception, so the option a learner selects becomes evidence in its own right. A conventional test asks only whether the answer is right. A diagnostic question asks which wrong idea a learner is holding when they miss it. The approach was formalised by Philip Sadler in 1998, who showed that instruments whose distractors mirror real learner conceptions behave very differently from conventional tests.1 The shift sounds small, but it changes what a test can do: a single well-built question stops being a pass-or-fail gate and becomes a probe into the structure of a learner's understanding. Each option is a hypothesis about how a mind might go wrong — and the answer chosen tells you which hypothesis fits.
Why do wrong answers carry more information than right ones?
A correct answer is almost silent. It tells you the learner reached the right place, but not how, and not whether they could do it again. A wrong answer, when the options are designed, is far richer — because there are many distinct ways to be wrong, and each one points to a different broken idea. Two learners can both miss the same question and need completely different help. One has a shaky procedure. The other has a confident but mistaken rule. If the distractors are written so that each maps to a known misconception, the specific option a learner picks names the fault rather than merely flagging it. This is why diagnostic assessment treats the wrong answer as the signal, not the noise. A right answer closes the conversation. A well-chosen wrong answer opens it, turning a single mistake into a precise, actionable description of where a learner's thinking diverged.
How did the Force Concept Inventory prove the idea?
The most famous demonstration came from physics. In 1992, David Hestenes, Malcolm Wells and Gregg Swackhamer published the Force Concept Inventory, a multiple-choice test of basic Newtonian mechanics whose distractors deliberately embodied the commonsense beliefs students bring to a physics class.2 Because each wrong option captured a real intuition — that motion implies a force, or that heavier objects fall faster — the pattern of answers exposed exactly which everyday ideas were getting in the way. The result was uncomfortable and influential. Conventional instruction shifted these beliefs far less than anyone assumed, and the effect held across teachers and teaching styles. The inventory became the gold-standard concept test in the physical sciences precisely because its wrong answers were the point, not an afterthought. It proved, at scale, that a carefully designed distractor could measure understanding that a right-or-wrong score completely missed.
What did Sadler discover about distractors?
Sadler's 1998 study went further and examined what happens to the numbers when distractors are tied to misconceptions.1 He found that such instruments have a fundamentally different psychometric profile from conventional tests. Most strikingly, the probability of choosing a particular wrong answer did not always fall as a learner's overall ability rose. Some misconceptions persist, and even strengthen, before they finally give way. That pattern only makes sense if misconceptions are treated as developmental states a learner passes through, rather than as random errors. The practical consequence is large: a wrong answer is not just a measure of what a learner lacks, but a clue to where they currently are on the path to understanding. Reading the spread of wrong answers across a class tells a teacher which mistaken models are common and which are fading — turning assessment from a ranking exercise into a map of how understanding is actually developing.
How does this scale to millions of answers?
What began with single tests now runs at the scale of millions of responses. The NeurIPS 2020 Education Challenge, published by Wang and colleagues in 2021, was built on more than 20 million real answers to diagnostic mathematics questions from the Eedi platform, and drew nearly 400 teams competing to predict learner behaviour.3 Crucially, one track asked models not just whether a learner would be right, but which option they would choose — making the misconception itself the prediction target. That framing only works because the questions were diagnostic to begin with: the wrong answers carry the information the models learn from. The challenge made explicit what the earlier studies implied — good question design and good prediction are two halves of the same loop. A model is only as diagnostic as the items feeding it, which is why the quality of the question bank matters as much as the sophistication of the method.
Why does this matter in the classroom?
All of this changes what a teacher can do on a Monday morning. A conventional mark says a learner scored six out of ten and leaves the teacher to guess what went wrong. A diagnostic result says something a teacher can act on directly. This learner is adding numerators and denominators separately, or believes a heavier object falls faster. The wrong answer has become a named diagnosis instead of a verdict. That is the whole promise of distractor-driven assessment — it makes the reason for a struggle visible early enough to do something about it, and it points to the specific idea to reteach rather than the whole topic. The goal was never to rank learners more precisely. It was to hand the teacher a clearer picture of each mind in the room, so that limited classroom time goes exactly where the evidence says it is needed most.
That principle — every result should explain itself — is the one CogniTrace is built around.
Key takeaways
- A diagnostic question writes each wrong option to match a known misconception, so the answer chosen reveals the mental model behind it.
- Wrong answers carry more information than right ones: there are many ways to be wrong, and each points to a different broken idea.
- The Force Concept Inventory (1992) and Sadler (1998) showed distractors can measure understanding that a score misses entirely.
- At scale — 20M+ answers in the NeurIPS 2020 / Eedi challenge — predicting which wrong answer a learner picks became the goal itself.
References
- Sadler, P. M. (1998). Psychometric models of student conceptions in science: Reconciling qualitative studies and distractor-driven assessment instruments. Journal of Research in Science Teaching, 35(3), 265–296.
- Hestenes, D., Wells, M., & Swackhamer, G. (1992). Force Concept Inventory. The Physics Teacher, 30(3), 141–158. doi:10.1119/1.2343497
- Wang, Z., Lamb, A., Saveliev, E., Cameron, P., Zaykov, Y., Hernández-Lobato, J. M., Turner, R. E., Baraniuk, R. G., Barton, C., Peyton Jones, S., Woodhead, S., & Zhang, C. (2021). Results and Insights from Diagnostic Questions: The NeurIPS 2020 Education Challenge. Proceedings of the NeurIPS 2020 Competition and Demonstration Track, PMLR, 133, 191–205. arXiv:2104.04034