ailiteracynepal 🇳🇵
Text size

Chapter 03 · Section II · 17 min read

Rubric-based grading with AI in the loop

The defensible grading workflow that saves a Nepali teacher a weekend on internal assessment — and the borderline-review rule that keeps a child’s grade from drifting on the model’s whim.

It is Friday evening. Forty-five Class 9 history papers on the Rana regime are sitting in a stack on the dining table. Each paper has six short-answer questions and one long-answer question; each paper, marked properly with feedback, takes you about twelve minutes; the arithmetic is unforgiving and the parent-teacher meeting is on Monday. The honest question is whether AI can give you back a Saturday without giving away your professional judgement. The honest answer is: yes, but only if you follow a specific workflow, and the workflow is stricter than most teachers initially want it to be.

The rubric + anchor approach

The single failure mode in AI-assisted grading is the model “scoring on vibes.” Give it a student answer and ask “out of 10?”, and you will get a number that looks reasonable on the first paper, drifts on the tenth, and is wildly inconsistent by the fortieth. Models do not have a stable internal scale. They invent one each time and forget it between answers.

The fix is to give the model a scale it cannot invent away from. Provide the rubric AND one example answer at each band. Not just “9 means excellent, 5 means adequate.” A concrete student-style answer that, in your judgement, deserves a 9. Another that deserves a 7. Another at 5. Another at 3. These are called anchor responses, and they do most of the work.

A prompt that produces defensible draft grades looks like this:

“You are helping me grade a Class 9 short-answer question on the causes of the 1950 revolution against the Rana regime. The question is worth 10 marks. The rubric is: 4 marks for naming three valid causes, 3 marks for explaining each cause briefly, 3 marks for one sentence linking the causes to the events of 1950-51. Here are four anchor answers from previous student work that I have already graded: [paste anchor at 9, anchor at 7, anchor at 5, anchor at 3]. For each new student answer I paste below, give me: (1) a draft score out of 10, (2) the specific rubric clauses you relied on, (3) the closest anchor it most resembles, and (4) one flag if the answer is within 1 mark of a band boundary.”

The anchors do four things at once. They define the bands operationally instead of abstractly. They calibrate the model to your standard rather than its internal one. They give you a paper trail — “the model said this answer most resembles the anchor I personally graded at 7” — that is defensible if a parent questions a score. And they make the model’s scoring stable across the full set of forty-five papers, because the anchors do not drift between answers the way the model’s vibe does.

The teacher review gate

Here is the non-negotiable part. You read every score before it is recorded. Not just the borderline ones. Every score. But not all of them with the same intensity.

A practical rhythm: the model produces draft scores and rubric notes for all forty-five papers. You then run two passes.

Pass one — every borderline score, in full. Any draft score the model flagged as within one mark of a band boundary, plus every long-answer score (because long answers are where models drift most), plus every paper where the model’s confidence sounded uncertain in its notes. You read the student’s full answer, re-check the rubric, and decide. This is the pass where you catch the unfairness.

Pass two — a random sample of non-borderline scores. Pick roughly one in five of the remaining papers, chosen at random rather than by which child wrote them. Read each one against the model’s draft. If you agree with all of them, the system is working that day and you can accept the rest. If you disagree with even one, your random sample just became “read all of them” — because the model is drifting and you need to recalibrate.

This workflow is not as fast as “let the model grade everything.” It is dramatically faster than marking forty-five papers from scratch. Twelve minutes per paper becomes about three minutes per paper on average — the borderline cases take longer, the clear cases take less — and the weekend comes back.

Why not auto-grade

It is tempting, especially under time pressure, to skip the review and just record the model’s scores. Resist this completely, for three concrete reasons.

Models drift on creative and argumentative answers. Short-answer factual items the model handles reasonably; long-answer items where a student takes an unexpected angle, or argues a position the model did not anticipate, are where the score becomes unreliable. The same answer, scored twice in two different sessions, can come back with a two-mark spread. Your professional consistency cannot ride on that.

Auto-grading high-stakes work has documented unfairness. International studies on automated essay scoring have shown systematic biases against non-standard dialects, against unusual but legitimate arguments, and against students whose writing voice does not match the model’s training distribution. In a Nepali classroom — where some students are writing in their second or third language — these biases land directly on the children least likely to push back.

The pedagogical knowledge of the student is missing. You know that Anish has been struggling with paragraph structure all term but finally produced a clear thesis this week; the model sees an answer and applies the rubric. Your knowledge of where the child has come from is part of fair assessment, and it lives in your head, not in the prompt.

Where AI-assisted grading helps most

Be honest about where this saves the most time. The headline use is large internal assessments: unit tests, end-of-term papers, project reports, the marking stack that used to define your weekend. Forty-five Class 9 history papers, sixty Class 8 maths quizzes, twenty-two +2 economics essays. The rubric is stable, the question set is fixed, and the volume is large enough that the calibration overhead — building the anchors, writing the prompt — pays for itself.

Where it helps least: one-off pieces of work where building anchors costs more than just marking the thing, and high-stakes external papers (SEE, +2 boards) where you are not the final marker anyway and where defensibility requires the entire process to be human.

A concrete Nepali scenario

Sushila madam teaches Class 9 history at a community school in Lalitpur. Forty-five papers on the Rana regime, due Monday. Pre-AI: she sets aside Saturday morning and most of Sunday afternoon. Post-AI workflow: Friday evening she picks four papers from a previous term to use as anchors, spends thirty minutes confirming her own grades on them, and writes the prompt. Saturday morning she runs the forty-five papers through the model in batches of five, scanning the rubric notes as they come back. By eleven, she has draft scores and a flagged list of nine borderline papers. She reads the nine in detail and changes three. She picks seven more papers at random, reads them, agrees with all seven. She records the grades.

What changed: the weekend. What did not change: the grade on every paper still has Sushila madam’s name on it, the borderline cases were decided by her judgement, and if a parent asks about the score, she can name the rubric clauses and the anchor the decision rested on. That is the deal. Anything less and the productivity gain is borrowed from the children.

Check your understanding

Quick check

Which workflow describes the safe use of AI in grading 45 short-answer papers for an internal assessment?

What comes next

Defensible scores are necessary but not sufficient. A child who receives a fair grade and a vague paragraph of feedback walks away knowing what they got, not what to do next. The next section is about the kind of personalised feedback that actually changes how a student writes the following week — and the structured prompt that stops AI from defaulting to lengthy, generic praise that no child has ever re-read.