Study & Learning
Generating Quiz Questions That Predict Your Exam Score
By Jim Vernon, Editor, AI Intelligence International · Published 8 March 2026 · Reviewed against our editorial standards · About the author
Scoring ninety per cent on self-generated questions and sixty on the exam is a specific, common and avoidable failure. It happens when the practice questions test recognition and the exam tests application.
This article covers how to specify question difficulty and format so your practice score becomes a usable prediction rather than a comfort.
Key takeaways
- Match the cognitive level, not the topic: Two questions on the same topic can be trivially different in difficulty.
- Format matters more than people assume: Multiple choice and short answer measure different things.
- Calibrating difficulty: If you are scoring above about eighty-five per cent on practice questions, they are too easy to be informative.
- Avoiding source leakage: Questions generated from a passage often reuse its phrasing, which lets you answer by matching words rather than by understanding.
Match the cognitive level, not the topic
Two questions on the same topic can be trivially different in difficulty. Define this term is recall; given this scenario, which principle applies and why is application; here are two conflicting accounts, evaluate them is analysis.
Find out which level your exam operates at by reading past papers, then specify that level explicitly when generating. Unprompted, models default to recall because recall questions are easiest to produce cleanly.
A useful instruction: generate questions at the application level, each presenting a novel scenario not described in the source text, requiring the reader to select and justify a principle.
Format matters more than people assume
Multiple choice and short answer measure different things. If your exam is written answers, practising multiple choice will inflate your confidence, because recognising a correct option is far easier than producing one.
Where the exam is multiple choice, quality of distractors is everything. Ask for plausible wrong answers that reflect common misconceptions, not obviously incorrect filler.
For written exams, generate the question and the mark scheme separately, then attempt the answer before reading the scheme. Reading them together destroys the test.
Calibrating difficulty
If you are scoring above about eighty-five per cent on practice questions, they are too easy to be informative. The productive range for learning sits closer to sixty to eighty.
Ask for a stated difficulty distribution and check it empirically. Models are inconsistent at self-rating difficulty, so your own results are the calibration.
Regenerate with harder constraints rather than adding volume. Fifteen hard questions beat sixty easy ones on every dimension that matters.
Avoiding source leakage
Questions generated from a passage often reuse its phrasing, which lets you answer by matching words rather than by understanding. This is the single biggest reason practice scores overpredict.
Instruct the model to paraphrase all terminology and to construct scenarios not present in the source. Then spot-check: if a question's answer is obvious from a skim of the text, discard it.
Better still, generate questions from one chapter and attempt them a week later, when the surface phrasing has faded.
Grading yourself usefully
Mark against the official mark scheme where one exists, not against a generated answer. Real schemes reward specific things, and learning what they reward is a substantial part of exam performance.
Record the reason for each loss: did not know, knew but misread, knew but ran out of time, knew but explained poorly. The distribution tells you what to fix, and it is usually not what you assumed.
Repeat missed questions after a week, not immediately. Immediate repetition tests short-term memory and produces a flattering result.
Building toward full papers
Individual questions are for learning; full papers under timed conditions are for prediction. Neither substitutes for the other, and most people do far too little of the second.
Sit at least two complete papers under exam conditions before the real thing. Time pressure and question-order effects are skills in themselves.
Use generated questions to fill the gaps past papers leave — new topics, recently changed syllabus areas, and anything you keep failing.
Worked example: calibrating a set
First generation from a chapter produced thirty questions. Practice score: ninety-two per cent. That is a signal of a bad question set, not a good student.
Inspection showed twenty-two of thirty reused distinctive chapter vocabulary in the stem. Regenerated with a paraphrase constraint and novel scenarios; new score, sixty-eight per cent.
The error log across the harder set showed most losses were in one sub-topic and one question format, both of which had been invisible under the easy set.
Two targeted sessions later, the same set scored eighty-one per cent, and the eventual exam score landed within four points of that figure.
Make the questions harder than recognition
Generated multiple-choice questions default to recognition, where the correct answer is obvious beside three implausible distractors. Recognition feels like knowledge and predicts exam performance poorly.
Ask instead for short-answer questions, for 'explain why this is wrong' items built from common misconceptions, and for problems that require applying a concept to an unfamiliar situation. Specify that distractors must be plausible errors a student actually makes.
Always check generated answers against your source material before trusting them. A confidently wrong answer key learned through repeated self-testing is worse than no quiz at all.
Frequently asked questions
How many practice questions should I generate per topic?
Fifteen to twenty hard ones, not a hundred easy ones. Beyond that you are testing stamina rather than knowledge, and past papers do that better.
Can AI write good multiple-choice distractors?
Yes, if you ask for distractors based on specific common misconceptions. Left unspecified, it tends to produce one plausible option and three obvious rejects.
Should I generate questions before or after studying a topic?
Both. A pre-test primes attention and shows what to look for; a post-test measures what stuck. The pre-test score is expected to be poor and is not a problem.
Why do practice scores overpredict exam scores?
Almost always because of phrasing leakage and cognitive level mismatch. Fix those two and the gap narrows to a few points.
How often should I retest the same material?
Expanding intervals — a day, three days, a week, a fortnight — and always retest the items you got wrong, not the ones you found easy.