Click to upload or drag and drop
PDF, DOCX, PPTX, TXT, JPG, JPEG, PNG, HEIC, ODP, ODT, BMP, or TIFF
up to 20MB
Uploading...
To check whether an AI generated quiz covers your document, map each question back to the section it came from and compare that distribution against how much of the material each section represents. Most generated sets pass a read through and fail this test, because uneven coverage looks exactly like good coverage until you count. The audit takes about five minutes and it is the difference between a quiz that measures the reading and one that measures the first twenty pages of it.
This matters more as the source gets longer. On a six page handout, coverage takes care of itself. On a 180 page training manual or a full textbook chapter, models draw unevenly from what they retrieved, and nothing in the output signals that anything was skipped.
| Step | What you do | What it catches |
|---|---|---|
| 1. List the sections | Write out every major section or chapter heading in the source | Gives you the denominator, which nobody has by default |
| 2. Tag every question | Mark each question with the section it tests | Items that cannot be traced to any section at all |
| 3. Count per section | Tally questions by section | Zero coverage sections, the most common failure |
| 4. Compare to weight | Set the tally against how much class time or page count each section had | Sections that are technically covered but underweighted |
| 5. Fill the gaps | Generate additional items for the empty and thin sections only | Rebuilding a quiz that was mostly fine |
If you generated the quiz in a chat assistant, step 2 has a shortcut. Ask it directly: list every major section of the attached document, then state how many of the questions you wrote came from each. The gaps appear immediately. Treat the answer as a starting point rather than gospel, and spot check two or three of its claims against the actual questions.
Late sections are the ones that go missing. A chapter's final section often carries the synthesis, the exceptions, or the procedure that matters most in practice, and it is routinely the least represented in a generated set. Check the last twenty percent of any long source first.
Prose generates questions easily. A table of thresholds, a decision tree, or a diagram often generates none, even though that is frequently the content people most need to have learned. If your document's key information lives in a table, say so explicitly and ask for questions built on it.
Rules with exceptions get flattened. The source says a step applies except in three situations, and the generated question tests the general rule while ignoring the exceptions entirely. Those exceptions are usually the reason the section exists, particularly in procedures and compliance material.
A quiz can touch every section and still be badly balanced. If you spent three weeks on section four and one afternoon on section seven, a set with two questions from each is not a fair test of the course, even though the coverage table looks clean.
The fix is a rough blueprint before you generate: decide the percentage of items each section should carry, based on instruction time or on how much the material actually matters, then generate against those targets. Twenty questions with a stated split of 6, 5, 4, 3, 2 across five sections beats twenty questions distributed by whatever the model happened to retrieve. Doing this deliberately is the substance of building a unit test from your lesson plans.
Take a 90 page onboarding manual with nine sections and a request for 30 questions. A clean result would be roughly three or four per section, adjusted for section length. A typical unaudited result looks more like nine questions from the first two sections, nothing at all from the compliance appendix, and four questions that are variations on the same policy statement worded differently.
The duplicates are worth flagging separately. Near duplicate items inflate the question count without adding measurement, and they are easy to miss when you read a quiz top to bottom rather than grouped by topic. Sorting the questions by section, which the audit makes you do anyway, surfaces them immediately.
The audit above exists because chat assistants retrieve passages rather than working through a document systematically. A generator built for this reads across the whole file in one pass and draws items from throughout it, which does not remove your review but does change what you are reviewing for: question quality rather than whether half the material was silently skipped. That is the practical difference described on the ChatGPT quiz generator and Gemini quiz generator pages.
For source material you reuse every term, the better long term move is to build the items once into a question bank organized by section, then pull each test from the bank against your weighting targets. Coverage becomes a property of the bank rather than something you re verify with every new quiz.
In a corporate setting there is a second reason to care. If the quiz is the record that someone understood a policy or a procedure, a gap in coverage is a gap in your evidence. Teams that need to show who was trained on what usually pair the assessment with an LMS that records completions and scores, and a quiz that skipped the compliance appendix will not survive the first time an auditor asks which sections it tested.
Tag every question with the section it came from and count by section. Any section with zero questions was missed. Reading the quiz straight through will not reveal this, because the questions that exist look fine and there is nothing on the page representing what is absent.
Work from weighting rather than a fixed number. Decide the share of the test each section deserves based on instruction time or importance, then convert that share into item counts. For a 20 item test across five uneven sections, something like 6, 5, 4, 3, 2 is more defensible than four each.
Because prose converts to questions far more readily than structured content does. A table of thresholds contains dense, testable material but no sentence to reshape into a stem. Ask explicitly for questions based on the tables and figures, and check that the numbers in those items match the source.
Not reliably. Asking for 50 instead of 25 often produces more items drawn from the same sections, plus more near duplicates. Targeted requests for the specific sections that came back empty work better than raising the total.
About five minutes for a 25 question set, less once you have done it a few times. It is faster than the alternative, which is discovering during the exam review that a third of the material was never tested and the scores do not mean what you assumed.
From the same family of tools