Asking how good your subjunctive is gets you an opinion. Asking you to repair a broken sentence gets you evidence.The check-up only does the second one.
Every question is a real item you have to answer. Retrieval practice is one of the better-evidenced ways to learn, so the diagnostic pays back part of the time it costs you. We looked at a self-rating shortcut and dropped it.
Three options, and what that does not fix
A selected-answer item here offers three credible options. Across eighty years of item research a fourth option is usually the one nobody picks, and writing it costs attention that a better third option deserves. So we write three and reject filler.
Three credible options still leave guessing on the table. Pick one at random and you are right about a third of the time. That is the reason the next section exists.
I'm not sure is a real answer
A lucky guess and a known answer look identical in the data, and an unlucky one is worse than useless: the option you picked can stay with you as though it were correct. Feedback after the test reduces that effect, which is why every item's reasoning is shown once you finish.I'm not sure takes the guess out of the record. The item counts as unanswered evidence for that check-up's score, and your mastery for that rule stays exactly where it was. Nothing about your profile gets worse because you said you did not know.
Choose an answer when you can reason it out. Use I'm not sure when you would only be guessing. We will show the explanation after the check-up.
The CEFR precedent
DIALANG was the first major CEFR-based diagnostic system: an online suite across fourteen European languages, built for diagnosis and feedback rather than certification. It settled two things we copied. Grammar can be diagnosed separately from vocabulary, and a result is worth reporting against a public scale.
One letter is not enough
A CEFR band tells you very little about what to study on Tuesday. The English Grammar Profile showed the alternative by deriving corpus-based grammar can-do statements keyed to levels. Our rule catalogue plays that role across the languages we support, and your result comes back rule by rule.
Free and Rich answer different questions
A free check-up is a recognition snapshot. It is quick, it costs nothing, and it tells you which rules are worth looking at. Because recognition evidence cannot show that you can produce a form, a free run recommends a level change for you to confirm rather than applying one.A Rich check-up adds written production, follow-up items chosen from how you answered, and anchor items from the levels either side of yours. Those anchors are what lets a Rich run apply a level change instead of proposing it.
What one check-up can tell you
Your result is an estimate from a bounded set of items. It describes the material that run actually covered, and nothing beyond it. Rules you answered once carry low confidence, and the profile marks them as thin instead of rounding up.
Re-take the check-up after a few courses. The change between two sittings is the measurement we trust most, because it is far harder to fake than a single sitting.
Getting the most out of it
Do the check-up in one sitting. A split session measures your evening as much as your grammar.
Use I'm not sure the moment you catch yourself picking at random. It protects the rule you have not learned yet.
Read the explanations at the end. They are written for the answer you actually gave.
Re-take it after two or three courses, not after one chapter.
Treat a thin-confidence row as a question, not a verdict.
Under the hood
Items are drawn per rule from the cheat-sheet catalogue for your language, with the count per rule bounded so that a check-up stays short.Each answer updates a per-rule estimate with a confidence value attached. Rules that never came up are marked unmeasured rather than assumed correct.An unsure response is stored as its own outcome. It enters the run's bounded score as unanswered and is excluded from the mastery update, so it can neither raise nor lower a rule.
Rodriguez (2005) Three options are optimal for multiple-choice items: a meta-analysis of 80 years of research. Educational Measurement: Issues and Practice.
Butler & Roediger (2008) Feedback enhances the positive effects and reduces the negative effects of multiple-choice testing. Memory & Cognition.
Alderson & Huhta (2005) The development of a suite of computer-based diagnostic tests based on the Common European Framework. Language Testing.
O'Keeffe & Mark (2017) The English Grammar Profile of learner competence. International Journal of Corpus Linguistics.