Chapter 03 · Section I · 17 min read
Screening resumes with AI — what is defensible
A defensible screening workflow uses a written rubric, surfaces a top-N for human review, samples the rejected pile, and never lets the model issue a rejection on its own — because a chatbot that cannot be cross-examined cannot defend a hiring decision.
You are an HR manager at a Lalitpur fintech. The CTO has signed off two junior engineer slots, you posted on Tuesday, and by Friday morning you have 214 CVs in your inbox. Some are PDFs, some are LinkedIn exports, some are scanned photos taken on a phone. You have a long-list to produce by Monday and three other open roles competing for the same hours. A vendor email lands offering an AI tool that will, the email promises, “intelligently screen all 214 CVs in under five minutes.” This section is about what you should and should not let that tool do — and what the workflow that will actually hold up looks like when somebody asks you, six months later, why a particular candidate did not get an interview.
Start with a written rubric — before the model touches anything
The single most important act in defensible AI-assisted screening happens before any tool is opened. Write down what you are screening for. Not in your head, not in the JD prose, but in a one-page rubric that a colleague who has never seen the role could apply to a stack of CVs and reach roughly the same shortlist you would.
The rubric has three parts. Must-haves — the small number of attributes a candidate cannot pass without. For the Lalitpur junior engineer role: a bachelor’s degree (any field) or demonstrated equivalent project work; ability to read and write English well enough for the codebase and the team’s standup; at least one piece of evidence of building software that exists outside coursework — a GitHub repo, a deployed project, a freelance gig. Three must-haves, written specifically enough that you can answer yes or no for each CV without arguing.
Evidence types — what counts as evidence for each must-have. “Building software” is vague; “a GitHub link, a Play Store app, a hackathon project with a working demo, or a paid freelance reference” is specific. Without this list, the model will make up its own evidence taxonomy on the fly, and so will three different reviewers.
Scoring bands — how the model and the human reviewers are to mark a CV against the rubric. A simple four-band scale (clearly above bar, plausibly above bar, plausibly below bar, clearly below bar) is enough; finer scales create false precision. The bands are not the decision — they are the input to the decision.
The rubric is the contract between you and the tool. When you prompt the model, you give it the rubric and ask it to apply your rubric, with specific evidence quotes from the CV for each judgement. You are not asking it to use its own judgement of “a good engineer.” You are asking it to be a fast, tireless reader who marks CVs against criteria you wrote. That distinction is the difference between AI as a force-multiplier for HR and AI as an unaccountable hiring manager.
Top-N plus sample-from-bottom — the human review pattern
Even with a clean rubric, the model will make mistakes. It will misread a CV from a Tribhuvan-affiliated rural campus whose format it has not seen often. It will over-weight an English-medium school in Kathmandu because that pattern is over-represented in its training data. It will miss the strongest junior in your pile because she described her project in two careful Nepali sentences instead of a paragraph of buzzwords. These errors are not malice; they are the predictable consequence of a model trained on the wrong distribution applied to a Nepali pool.
The workflow that catches these errors is top-N plus sample-from-bottom. You ask the model to surface the top thirty CVs against the rubric, with its evidence quotes for each judgement. Two human reviewers — ideally one HR and one technical — read all thirty in full and re-score them against the rubric. So far this is standard.
The unusual move is the second pile. Take the model’s rejected CVs and pull a random sample of twenty out of the bottom 180. The same two reviewers read those twenty against the rubric. The purpose is not to find every wrongly-rejected candidate — at this volume that is statistically impossible. The purpose is to catch the systematic error, the one where the model is dropping a whole category of candidate for a reason that is not in the rubric. If three of the twenty sampled rejections turn out to be clearly above-bar candidates, you have a signal that the model is broken for this pool and the top-N is suspect. You go back, fix the prompt or the rubric, and re-run.
Never auto-reject — especially at the entry level
The single most legally exposed pattern in AI-assisted hiring is the auto-reject. A threshold score, a hard filter, a “candidate did not meet our criteria” email generated and sent by software without a human eye on the CV first. It is fast, it is cheap, it is being marketed to you, and it is the thing that will end a Nepali firm’s reputation when the first wrongly-rejected candidate with a credible discrimination claim takes the screenshot to a journalist.
Auto-rejection is most dangerous at the entry-level / junior pool, because that is where CV noise is highest. A senior engineer with twelve years of experience has a CV the model can parse confidently — recognisable employers, recognisable titles, a clean LinkedIn. A junior fresh from a Pokhara campus may have a single-page CV with an email address, a degree, and three small projects described in two lines each. The signal-to-noise ratio is low, the model’s confidence interval is wide, and the cost of a wrong rejection — to a candidate who is at the start of her career, who may not have other offers — is high. The auto-reject pattern fails worst exactly where it is most tempting to apply.
The defensible alternative is the workflow above: the model surfaces and ranks, the human decides. Every rejected candidate has been seen, or at least sampled, by a human. The rejection emails go out from the HR account, with the rubric criterion that was missed if the candidate asks. This is slower than auto-rejection by perhaps an hour per cycle. That hour is the price of a defensible process, and it is small.
Document the tool, the rubric, the audit — in a workpaper
Borrow the discipline from accounting. For each hiring cycle that used AI, keep a single-page workpaper in the role folder with five things on it.
One. The tool used, by name and version — ChatGPT 4.7 web, paid tier, date — not just “we used AI.” Two. The prompt, in full. The model’s behaviour depends on the prompt; without the prompt, the audit cannot be reconstructed. Three. The rubric, attached. Four. The top-N list the model produced, and the human-revised final shortlist with notes where they diverged. Five. The sample-from-bottom results, with notes on any systematic errors caught and the fix applied.
This workpaper is twenty minutes of work per cycle. It is the artefact you produce if a rejected candidate complains, if the Department of Labour writes a letter, if your auditor asks how the HR function uses AI, if your insurer asks the same question. Without it you have memory and a chatbot history; with it you have a documented professional process. The difference matters most on the one day a year you need it to.
The 200-CV cycle, end to end
To make this concrete, return to the Lalitpur fintech. Friday, 214 CVs, two slots. You spend forty-five minutes writing the rubric with the engineering lead. You spend ten minutes writing the prompt — here is the rubric, here are the evidence types, here are the scoring bands, mark each CV with a band and supply two quotes per must-have, output as a table. You spend two minutes running the model on the CV set. You spend two hours, with the engineering lead, reading the top thirty and the sampled twenty from the bottom and producing a shortlist of eight. You spend twenty minutes on the workpaper. Total: just over three hours, against an estimated eight to ten hours for the same shortlist by hand. The two slots get filled; the 206 unsuccessful candidates get a rejection note from a human-reviewed list; the workpaper goes into the folder.
Six months later, a rejected candidate writes to ask why she did not progress. You pull the workpaper. The rubric says English fluency was a must-have, evidenced by the CV being writable in unambiguous English. Her CV was in mixed English-Nepali with grammatical issues that two reviewers flagged. You can answer her honestly, in writing, with the rubric attached. The conversation closes. That is the value of the workflow.
Check your understanding
Quick check
—You are screening 200 entry-level engineering CVs for a Lalitpur fintech with two open junior slots and four days to shortlist. Which workflow is the most defensible?
Quick check
—Why is auto-rejection — using an AI score to send rejection emails without any human eye on the CV — particularly dangerous for entry-level hiring in Nepal?
What comes next
A defensible workflow at the individual cycle level is necessary but not sufficient. Over months and quarters, even a well-prompted model can produce a pipeline that is silently unfair to particular groups — by gender, by region, by university tier — without any single cycle looking wrong. The next section is about how to audit your own pipeline for those patterns, what counts as a real signal versus statistical noise at Nepali SME volumes, and the one-page quarterly memo that turns the audit into a record.