Chapter 03 · Section I · 17 min read
Drafting assessment items aligned to learning objectives
The single instruction that decides whether your AI-drafted quiz tests recall or thinking — and the verification gate that keeps a wrong question from reaching a child.
The Class 10 SEE pre-board is in eleven days. You are setting the history paper. You sit down at the staff-room laptop, open the chatbot, type “give me twenty questions on the Rana regime for Class 10,” and the model produces twenty fluent-looking questions in about forty seconds. You print them. The first parent who looks at the printed paper closely notices that one of the dates is wrong, two questions ask the same thing in slightly different words, and almost every question is a fact-recall item even though the curriculum chapter is supposed to develop comparison and evaluation. This is the failure mode this section exists to prevent.
The verb is doing the work
Most teachers who first try AI for assessment make the same mistake. They name the topic — “the Rana regime”, “linear equations in two variables”, “the water cycle” — and they name the grade level, and they stop there. The model, given a topic and a grade and nothing else, defaults to the easiest cognitive level it can find: fact recall. Who, what, when, where. The questions look reasonable. They are also pedagogically wrong, because your unit was probably not asking the children to recall facts. It was asking them to explain, to compare, to evaluate, to apply.
The fix is small and it changes everything. Start the prompt with the verb from your learning objective, not the topic. “Write five questions that ask students to compare the administrative reforms of Chandra Shamsher with those of Juddha Shamsher” is a different prompt from “write five questions on the Rana regime.” The first one forces the model to construct items at the comparison level. The second one returns five variations of “in which year did X happen.” Same topic. Different cognitive demand. Different paper.
This is why every NEB unit objective in the curriculum starts with a verb. The verb is not decoration. It is the contract between you and the child about what the lesson was for, and it is the single most important word to feed the model.
Bloom as a prompt scaffold
If you want to be more deliberate — and on a board-style paper you should — Bloom’s taxonomy gives you a six-level vocabulary the model already understands. Remember, understand, apply, analyse, evaluate, create. Stating the level explicitly produces dramatically better-targeted items than topic prompts alone.
A serviceable pattern looks like this:
“Write four questions on the causes of the 1950 revolution against the Rana regime, for Class 10 students following the NEB curriculum. Two questions at the ‘understand’ level — students explain a cause in their own words. Two questions at the ‘analyse’ level — students identify which of several stated factors was most decisive and justify their choice. Short-answer format, 3 marks each. Use only facts from the NEB Class 10 history textbook, Chapter on the Rana period. If you are uncertain about any factual claim, mark it
[VERIFY]rather than guessing.”
Three things make that prompt work. The Bloom levels are named explicitly, so the model cannot quietly slide back to recall. The format and mark scheme are fixed, so you do not get a mix of MCQs and essays you did not ask for. And the [VERIFY] instruction is the gate that keeps invented “facts” out of your paper — more on that below.
NEB exam-style alignment
The children sitting in front of you will face specific question shapes in SEE and +2 board exams. Generic AI-drafted questions do not prepare them for those shapes. You need to name the shape.
Short-answer items on SEE history are typically 3 to 5 marks, two to four sentences expected, with a marking scheme that rewards a stated cause + a stated consequence + a brief justification. Long-answer items are 6 to 10 marks, structured paragraphs expected, marking scheme rewards thesis + evidence + counter-evidence + conclusion. MCQs on +2 papers expect one clearly correct answer, three plausible distractors (not silly ones), and no “all of the above.” If you do not tell the model these conventions, it will invent its own — usually closer to a Western standardised-test style than to NEB practice, and the children will be thrown by the shift on the actual paper.
A useful add-on to any item-generation prompt: “Format the questions in NEB SEE pre-board style: short-answer items at 3 marks each, with the marking scheme stating one mark for the cause, one for the consequence, one for the justification. Distractors on MCQs should be plausible misconceptions a Class 10 student might actually hold, not absurd options.”
Worked example: a balanced Class 10 history paper
Here is a full prompt for the Rana-regime paper that started this section.
“Draft a 20-question Class 10 history paper on the Rana regime in Nepal (1846-1951), aligned to the NEB curriculum. Mix: six questions at the remember/understand level (3 marks each, short-answer), ten at the apply/analyse level (5 marks each, short-answer), four at the evaluate/create level (8 marks each, long-answer). Topics to cover: administrative structure, social reforms and their limits, the 1934 earthquake response, foreign relations with British India, the rise of opposition movements, and the events of 1950-51. Use only facts that appear in the NEB Class 10 history textbook chapter on the Rana period. If you are uncertain about a date, a name, or a specific event, mark that item with
[VERIFY: claim]rather than asserting it. Also write a one-line marking scheme for each item. Provide the same paper in Nepali underneath in Devanagari.”
This is a long prompt. It is long because each clause buys you something specific. The Bloom mix bought you a paper that is not all recall. The mark distribution bought you SEE alignment. The “use only facts from the textbook” buys you a paper that the children’s preparation actually covers. The [VERIFY] instruction is the single most important line.
The verification gate
Here is the part the demo videos skip. The draft will come back fluent and confident. It will also, on a twenty-question paper, contain on average one to three quiet factual errors: a date off by a year, a wrong attribution of a reform to the wrong Shamsher, a “fact” the model has fabricated because it sounded reasonable. The [VERIFY] instruction reduces the rate sharply, but it does not eliminate it. The model does not always know when it is guessing.
So: before any AI-drafted assessment item reaches a student, open the textbook and check. Take a representative sample first — five items out of twenty, chosen to span topics and Bloom levels — and check each fact against the NEB chapter. If you find zero errors in the sample, check the rest at a faster pace; if you find one error, check every remaining item carefully; if you find two errors, throw the paper out and redraft with a tighter prompt. This is not bureaucratic caution. A wrong fact on a pre-board paper, marked correct in the answer key, teaches the wrong fact to forty children, and you will see it again — wrong — on their SEE answer sheets four months later.
A short note on differentiation
Once you have the discipline, AI lets you do something that was previously impractical: produce three versions of the same assessment at three difficulty levels for a mixed-ability classroom. Same Bloom levels, same topics, same total marks — but easier vocabulary and shorter prompts for the foundation tier, and stretch questions that demand multi-step analysis for the top tier. Forty minutes of prompting replaces what used to be a Sunday afternoon of rewriting. The verification gate still applies, three times over.
Check your understanding
Quick check
—You sit down to draft a Class 10 history quiz on the Rana regime using AI. Of the inputs below, which one most determines whether the quiz tests fact-recall or higher-order thinking?
Quick check
—Why is the instruction “if you are uncertain about a date, name, or event, mark the item with [VERIFY] rather than asserting it” worth including in every assessment-drafting prompt?
What comes next
Drafting the items is half the work. The other half — the half that eats the weekend after the children sit the paper — is grading forty-five short-answer responses against a defensible rubric without losing your evenings or your judgement. The next section is about rubric-based grading with AI in the loop: where it saves the marking time honestly, where it drifts dangerously, and the teacher-review gate that keeps a borderline score from quietly becoming a wrong one.