Chapter 03 · Section II · 17 min read
Auditing your own pipeline for bias
The recurring outcome audit — pass rates by gender, region, and university tier every cycle — is the single habit that separates HR teams who can defend their AI use from HR teams who only think they can.
The last section gave you a workflow that holds up at a single hiring cycle. This section is about the harder problem behind it: a workflow can be clean on Tuesday and your pipeline can still be quietly unfair on a six-month view. The model that performed well on one role with one rubric may, applied across five roles and four quarters, be producing a shortlist that systematically passes one group through faster than another. No single cycle looked wrong. The aggregate is wrong. The only way to know is to look — deliberately, on a schedule, with the same questions every time. This is the outcome audit, and it is the single most important habit AI-using HR teams in Nepal need to build right now, while the practice is young and the regulator has not yet written rules you will have to retrofit to.
The recurring breakdown
The audit is a table. Once per hiring cycle, or once per quarter if cycles are short, you produce a single sheet with one row per stage of the funnel — applied, model-shortlisted, human-shortlisted, interviewed, offered, hired — and three or four columns of breakdown for each row. The columns are the axes Nepal-specific bias is most likely to travel along.
Gender is the first column, because it is the cleanest signal. Most CVs in Nepal carry an unambiguously gendered name; you can mark each CV at the application stage and carry the marking through. The question the table answers is whether the pass rate from one stage to the next differs sharply by gender, and whether the gap widens or narrows as the funnel narrows. A pipeline that is 40 percent women at application and 12 percent women at offer is telling you something. What it is telling you depends on the role and the pool, but the table makes the conversation possible.
Region of origin is the second column. The seven provinces — Koshi, Madhesh, Bagmati, Gandaki, Lumbini, Karnali, Sudurpaschim — are the practical bins for a Nepali HR audit. You will not have clean data on every CV; many will be ambiguous. Do not let the missing data stop the audit. Mark what you can mark — usually from the address, the school, or the candidate’s stated hometown — and run the breakdown on the marked portion. A pattern where Madhesh-origin candidates pass the model-shortlist stage at half the rate of Bagmati-origin candidates with similar credentials is the kind of signal you want to catch in quarter two, not in year three when a journalist catches it for you.
University tier is the third column. The natural bins for Nepal are Tribhuvan University and its campuses, Kathmandu University, Pokhara University, named private universities, Indian universities, and overseas. These are not a quality ranking; they are a representation grouping. The question is whether your funnel passes graduates of Kathmandu University and overseas universities through at sharply different rates than equally credentialed Tribhuvan graduates. If yes, the model — or the rubric, or the reviewers — has acquired a preference that is not in your written criteria. That preference is the audit’s job to find.
A fourth column some teams add is first-language or language of CV: English-only, Nepali-medium, mixed. This is often the most diagnostic column for AI-assisted screening because language-pattern proxies are exactly what large language models latch onto. If your model-shortlist rate for English-only CVs is twice your rate for mixed-language CVs at the same scoring band, the model is reading fluency as a competence signal in a way your rubric did not authorise.
What “a gap” means at Nepali SME volumes
The standard objection to small-organisation audits is statistical. “We hire twenty people a year — we cannot get significant breakdowns across seven provinces.” This is true and it is not a reason to skip the audit. Two things to hold together.
First, statistical significance is not the only threshold that matters. If you hire twenty people and zero are from Karnali across two consecutive years, you do not need a p-value to investigate. The signal is the absence. The same is true for a fifteen-percentage-point gap in pass rates between two groups that are otherwise similar on the rubric criteria — same university tier, same evidence types, same scoring band. The gap is not proof of bias. It is a flag that something in the pipeline is treating two similar groups differently, and the response is not “is this statistically significant?” but “what is the most likely lever, and can we fix it?”
Second, you can aggregate across cycles. A single twenty-person cycle is too small for most breakdowns. Four cycles across a year is eighty hires, several hundred applications, and enough volume to spot pattern-level issues even if any individual cell stays small. The audit memo each quarter should reference the rolling annual numbers as well as the current quarter’s. The rolling number is where the signal usually appears.
What to do when a gap shows up
A gap is a starting point, not a verdict. The audit’s job is to surface it; the investigation’s job is to find the cause. There are four levers to check, roughly in the order of how often they turn out to be the real cause.
Lever one: the rubric. Read the rubric again with the gap in mind. Does it require something — “communication skills demonstrated in English,” “experience at a recognised tech employer” — that is itself a proxy for the group that is passing through faster? Rubric criteria written innocently can encode preferences nobody intended. A rubric that says “prior experience at a tech-forward employer” will quietly favour candidates from Bagmati, where most such employers are located. Rewrite the criterion to focus on the underlying skill the role actually needs, not the typical signal of that skill, and re-audit next cycle.
Lever two: the prompt. The prompt you give the model shapes what it weights. A prompt that says “find the strongest candidates” without further constraint is an invitation for the model to apply its own priors, which over-weight things like elite school names and English fluency. A prompt that walks the model through the rubric criterion by criterion, asking for evidence on each, narrows the model’s discretion sharply. The prompt is rewritable in five minutes; the gap may close in one cycle.
Lever three: the pool. Sometimes the gap is upstream of the model. If your application pool is 12 percent women and your hire rate is 12 percent women, the model is faithfully passing the pool through. The model is not the lever; the sourcing is. This sends you back to Chapter 2 — the JD and the channels and the language of the outreach. The audit is what tells you the sourcing is the problem; without the audit you would be tuning the wrong dial.
Lever four: the reviewers. AI is not the only source of bias in the funnel. Human reviewers carry their own preferences, often unexamined. If two reviewers consistently give the same group lower scores than two other reviewers at the same scoring band, the human-review stage is where the gap is being introduced. The fix is calibration — periodic exercises where reviewers score the same CVs independently and compare — not blaming the model.
Fix the most likely lever, document the fix in the audit memo, and re-audit next cycle. If the gap closes, you have learnt something about your process. If it does not, you have ruled out a lever and proceed to the next. This is what professional improvement looks like in HR; it is what regulators and complainants will ask you to show.
The quarterly audit memo
The artefact at the end of the audit is a memo, one page, filed and dated. Five things on it.
One. The period covered and the volumes — applications, shortlist, offers, hires, by month if useful. Two. The breakdown table — pass rates by gender, region, university tier, and any fourth axis your team has chosen. Current period numbers and rolling annual numbers, side by side. Three. Gaps flagged, with the threshold you used. Four. The investigation for each flagged gap — which lever was checked, what was found, what was changed. Five. The audit date for the next cycle, and the responsible owner.
Why this is the highest-leverage habit in the whole course
Across every section so far the message has been the same: AI in HR is useful when verification is cheap and the downside is small, and dangerous when verification is slow and the downside is somebody’s career. The pipeline audit is the mechanism that detects when the danger has crept into your process without anyone noticing. It is not glamorous. It is twenty minutes a quarter with a spreadsheet, plus the time to fix what it finds. That twenty minutes is the difference between an HR function that can answer for its decisions and one that cannot.
The teams that will get into trouble with AI in HR over the next three years are not the ones who used AI most aggressively or most cautiously. They are the ones who used AI without auditing — who let the tools run cycle after cycle without ever pulling the breakdown sheet and asking the obvious questions. The audit is the discipline that keeps the rest of the course’s practices honest. Build the habit before you need it.
Check your understanding
Quick check
—A Nepali firm runs an AI-assisted hiring pipeline. At the application stage the pool is 40 percent women; at the offer stage it is 25 percent women. The HR head says, 'But the model does not see gender — we never pass it in.' Which response is strongest?
Quick check
—What is the single most important habit for keeping AI-assisted screening defensible over time at a Nepali SME?
What comes next
This section assumed that you cannot eliminate bias by hiding the inputs — that the audit on outputs is what matters. The next section explains why, in mechanism-level detail. The “we removed the name” defence is the most popular and most wrong response to bias concerns in AI screening, and understanding exactly how it fails is what makes the case for the outcome audit airtight. It also names the specific Nepal-context proxies — surname for caste, neighbourhood for class, school for both — that quietly carry the information through whatever you redact.