Chapter 04 · Section I · 17 min read
Interview question generation and structured scoring
The single highest-leverage AI workflow for a Nepali recruiter is not screening — it is generating a structured interview kit and the rubric every interviewer must read before the candidate walks in.
By the time a candidate sits down across the table from you, most of the fairness of the hiring decision has already been determined — not by what you ask, but by whether the next candidate is asked the same thing in the same way against the same yardstick. A free-form interview where each panelist follows their own instinct is the easiest interview to run and the easiest to lose a discrimination claim against. The good news is that the one piece of preparation that fixes this — a structured question bank with a written rubric — is exactly the piece of preparation today’s AI produces well in under five minutes. This section is the paste-this-prompt walkthrough.
Why structured beats unstructured, and not by a little
Two decades of meta-analyses on interview validity converge on a finding that still surprises most managers: structured interviews — same questions, same order, same written rubric across every candidate — predict on-the-job performance roughly twice as well as unstructured ones. The unstructured “let us just have a conversation and see how they come across” interview is, in the literature, barely better than a coin flip at picking the candidate who will perform. It is also the format most exposed to first-impression bias, affinity bias, and the quiet drift toward whoever reminds the panel of themselves.
The reason structure works is not mysterious. When every candidate answers the same question, the panel is comparing answers. When every candidate answers a different question, the panel is comparing impressions. Impressions favour the candidate who is most like the panel — which, in Nepali offices that skew Kathmandu-male-Khas-Arya at senior levels, systematically disadvantages strong candidates from elsewhere. Structure is not a bureaucratic chore. It is the cheapest fairness intervention available to an HR team, and current AI lowers the cost of building it almost to zero.
The three inputs the model needs
A good interview-kit prompt has exactly three inputs, and a prompt that skips any of them produces generic output. Input one: the job description, pasted in full — not summarised, not paraphrased. The model needs to see the actual words, including the must-haves, the nice-to-haves, and the seniority signals. Input two: the seniority and the salary band — a question bank for an engineer with two years of experience and one for an engineering lead with eight are not the same kit, and the model will guess if you do not tell it. Input three: the three to five competencies you have decided this role must test for. This last input is the one most HR teams skip, and it is the one that does the most work. “Communication, technical depth, ownership, conflict handling” is a different interview from “speed, scrappiness, customer obsession, business sense” — and you, not the model, decide which set fits the role.
What to ask the model to produce
A serviceable prompt asks for four distinct artefacts in one pass, not one blob of “interview questions”. Each artefact does a different job in the interview.
“Here is the JD for a backend engineer at a Kathmandu fintech. Seniority is mid-level, three to five years. The four competencies I want to test are: backend system design, code quality and debugging discipline, ownership of incidents in production, and clear written communication with non-engineers. Produce: (1) eight behavioural questions in the STAR format, two per competency, each with two follow-up probes; (2) four technical or role-specific questions appropriate for a mid-level backend engineer, with what a strong answer would mention; (3) two scenario questions of the form ‘you are on call at 2am and X happens, walk me through your next thirty minutes’; (4) a one-page scoring rubric covering all four competencies, with a 1-to-5 scale and a one-line description of what each level looks like in an answer. Tone: practical, not academic.”
What comes back is a first-cut kit that a senior recruiter can polish in fifteen minutes. Behavioural questions ask the candidate to describe specific past situations, which is harder to fake than hypotheticals. Technical questions test whether the CV claim is real — not by quizzing on trivia, but by asking the candidate to reason out loud through something a person in that role would actually face. Scenario questions test judgement under the pressure of an unfamiliar situation; they reveal how the candidate thinks, not what they have memorised. The rubric is what turns those answers into comparable scores.
Three Nepali kits, sketched
A backend engineer at a Kathmandu fintech. Competencies: system design at the scale of a payments gateway, code-quality discipline, on-call ownership, written clarity for non-engineers. Behavioural probes target real incidents — “tell me about a time you shipped a change that caused a production issue; walk me through the next hour.” Technical questions cover idempotency, retry semantics, and what happens when the bank’s API times out mid-transaction. Scenario question: “It is 11pm on the day before Dashain, transactions are failing for one bank, the CEO is on the phone, what is your next thirty minutes?”
A programme officer at an INGO in Kathmandu. Competencies: programme design against a logframe, stakeholder navigation across donor / government / community, monitoring discipline, written communication suitable for a donor report. Behavioural probes target navigating disagreement between a donor and a local partner; scenario questions test what the candidate does when the field team reports a data anomaly two days before a quarterly report is due. The rubric explicitly rewards specificity over jargon — every INGO interview drowns in jargon, and the rubric is what lets you score past it.
A relationship manager at a commercial bank in Birgunj. Competencies: portfolio management of SME clients, credit judgement at the file-prep stage, conflict handling with a stressed borrower, integrity under pressure. Behavioural probes target moments when a long-standing client asked for an exception the policy did not allow. Scenario questions test what the candidate does when the branch manager pressures them to push through a file they have private doubts about. The rubric for integrity is the most important and the hardest to write — get the AI to draft it, then sit with the panel and rewrite the levels until everyone agrees on what a “3” actually looks like.
The interviewer-calibration habit
Here is the part of the workflow that the AI cannot do for you, and that almost every team skips. Before any interview begins, every interviewer on the panel reads the same rubric and agrees on what each score level means. Not “we all received the rubric in a calendar invite.” Read it together, in the same room or the same call, for ten minutes, the day before.
The reason is simple: a rubric on paper is a different rubric in each interviewer’s head. One panelist’s “4 — strong evidence of ownership” is another’s “3 — adequate ownership.” Without a calibration conversation, the panel produces scores that look comparable on the form and are in fact noise. The AI can produce the rubric in a minute. It cannot make the panel agree on it. That agreement is the actual fairness mechanism, and it is a human conversation.
A few practical guardrails
Three small disciplines save a lot of trouble. Do not let the model invent the JD. If you ask for an interview kit without pasting the JD, the model will produce a generic kit aimed at a fictional company; you will then conduct an interview against a competency profile you never wrote down. Always paste the actual JD.
Do not ask the model for “the answer.” It is tempting to ask the model to produce both the question and the model answer the candidate “should” give. For technical questions where there is a right answer, this is fine. For behavioural and scenario questions, the model’s “model answer” is a confident guess at what a good candidate looks like — and the rubric is supposed to be your judgement of that, not the model’s. Use the rubric, score against it, and let strong candidates surprise you.
Treat the kit as a starting point, not a script. A good interviewer follows the candidate into the interesting bits and probes where the rubric tells them to probe harder. The kit ensures you ask the same opening questions and score against the same scale. It does not require you to read each question off a sheet like a checkout receipt. The structure is in the questions and the rubric; the conversation around them is yours.
Check your understanding
Quick check
—A Nepali recruiter has 45 minutes to prepare for tomorrow's panel interview of three backend engineering candidates. Which AI workflow produces the most defensible, fairest interview?
Quick check
—Which of the following best explains why structured interviews predict on-the-job performance roughly twice as well as unstructured ones?
What comes next
Generating the kit and calibrating the panel handle the half-hour before the interview. The next section moves to the ninety seconds and the ninety minutes after it — how to use AI to turn scattered interviewer notes into a structured summary against the rubric you just built, without inventing what was not said and without leaking the candidate’s identity to a public tool.