Chapter 06 · Section II · 17 min read
The "verify everything" workflow
A single default rule — every AI-produced number is unverified until you have checked it against a system of record — and the tiered workflow that lets that rule survive March without collapsing.
The most expensive AI errors in accounting do not come from the questions where you know to be careful. They come from the moments when the output looks ordinary — a total that is roughly the size you expected, a rate that sounds about right, a reconciliation in which everything appears to tie — and you let it pass because nothing in the screen is shouting. The verification workflow this section describes exists precisely for that moment. It is built around one default rule, simple enough to remember at 10 p.m. on a Thursday in March, and disciplined enough to catch the silent errors that hide inside reasonable-looking numbers.
The default rule
Read this once and write it on the inside cover of your workpaper:
Every figure an AI produces is unverified until you have re-computed it from source or cross-checked it against a system of record. There are no exceptions for “obvious” cases. Obvious is where errors hide.
That last sentence is the load-bearing one. The figures that trick you are not the ones that look strange — those you check by reflex. They are the ones that look exactly as you expected, because the model has produced something plausible and you have unconsciously decided the work is done. A 2024 study of accountants reviewing AI-generated reconciliations found that error-catch rates dropped by more than half when the AI-produced figure was within five per cent of the expected value, even when the figure was wrong. Plausibility is the trap. The rule above exists because human attention does not.
A tiered verification approach
“Verify everything” sounds expensive. It is not, if you match the tier of verification to the cost of being wrong. There are three tiers, and a sensible workflow picks the lowest one that still catches errors at the rate the work requires.
Tier 1 — Automated re-check. The model produces a number; a formula in Excel, or a query in your accounting software, produces the same number from source. If they match, the figure is verified in seconds. If they do not, the formula is the truth and the model was wrong. This is the cheapest and most reliable tier, and it is the one to use whenever the source is structured: sums, sub-totals, depreciation across a schedule, VAT across a vendor list, withholding across a payroll. The thirty seconds you spend writing the SUMIF is the verification.
Tier 2 — Manual spot check. The model produces a categorised list — 200 transactions assigned to ledger accounts, say — and you pick a random ten, plus the five largest by value, and check each against the underlying document. If your sample is clean, you accept the rest under the standard sampling logic you already use in audit work. If your sample shows errors, you reject the whole output and either redo it manually or feed back to the model with the errors flagged. This is the tier for work where automated re-check is not possible but human eyes can adjudicate at the level of the individual item.
Tier 3 — Visual inspection. The model has summarised, translated, or restructured prose. You read what it produced, slowly, against the source. There is no shortcut; reading is the verification. This is the tier for management commentary, regulatory summaries, and client letters — work where the question is not arithmetic correctness but whether the meaning carried over.
The instinct to reach for is always the lowest tier that does the job. If a Tier 1 check is possible, use it; you do not need to read the schedule line-by-line if a SUMIF will catch any error in two seconds. If only Tier 3 is possible — because the output is prose — then accept that the work was prose-review work, and price it accordingly.
Concrete habits — never, never, never
The abstract rule turns into discipline only when it cashes out as specific habits that an honest practitioner will not break. Three of them are worth stating bluntly.
1. Never paste an AI-produced figure straight into a tax return, a financial statement, or any filing. The figure goes into a working first, where it can be tied to source. The filing pulls from the working, not from the chat window.
2. Never quote an AI-produced rate, threshold, or statutory limit to a client without checking the source. The model is fluent in tax law it has half-read. A wrong TDS rate quoted in writing is your professional exposure, not the model’s.
3. Never accept an AI-produced reconciliation without independently checking the unmatched items and at least sampling the matched ones. A reconciliation that “ties” is the easiest output for a model to fake convincingly, because the model knows the totals are supposed to agree and can produce a list that appears to make them agree.
The thirty-second test
For day-to-day decisions about whether to verify a given piece of output, a rough rule helps: if you can verify it in thirty seconds, do it always; if it would take thirty minutes, decide based on the cost of being silently wrong.
The thirty-second checks are cheap enough that there is no honest reason to skip them. Re-summing a column. Re-checking a single rate against the IRD site. Reading a one-paragraph summary against its source. These take less time than the cup of tea you are drinking while you do them.
The thirty-minute checks are where judgement enters. A full re-performance of a model’s depreciation schedule across 300 fixed assets takes real time, and the question becomes: if this is silently wrong, what does it cost? For an internal management report, perhaps the cost is small and the check can be a sample. For the depreciation figure in a statutory audit report, the cost is your name and the firm’s licence, and the check is not optional regardless of how long it takes.
The corollary, and this is the part most firms get wrong: when verification becomes too expensive to justify, that is a signal that AI is the wrong tool for this task — not a signal to skip verification. If checking the model’s work costs more than doing the work manually would have, you have not saved time; you have transferred risk to a place where you cannot see it. The right response is to do this kind of work without AI, not to do it with AI and pretend the check happened.
What this looks like in a real day
Consider the morning of a typical small-firm bookkeeper, three clients on her desk. Client one: a categorised expense schedule from 180 transactions, produced by the model. She runs a SUMIF in Excel by category — Tier 1 — and the totals tie to the bank export. Total verification time: under a minute. She accepts the output and moves on. Client two: a draft management commentary from the previous month’s P&L. She reads it against the trial balance — Tier 3 — and catches one variance the model attributed to the wrong driver. She fixes it and notes the correction in the workpaper. Time: ten minutes, against the forty it would have taken to draft from scratch. Client three: a depreciation re-computation across 230 assets following the new Finance Act rates. She asks the model to draft the schedule; she also asks it to show the formula it used per asset class. Then she rebuilds the schedule herself in Excel using the IRD-published rates and compares totals. Tier 1, but careful — the cost of silent error here is too high for a sample. She finds two assets where the model used the old pool rate. She corrects them. Total time saved versus building from scratch: about ninety minutes, because the structure was already there even if the numbers were not.
That is what the workflow looks like in practice. It is not a counsel of paranoia; it is the routine application of a single rule, tiered by cost, applied with no exceptions for outputs that happen to look ordinary.
Check your understanding
Quick check
—Which of the following best states the default rule that should govern every piece of AI-produced output you use in client work?
What comes next
Confidentiality and verification are the two technical disciplines. The third question is commercial, and it is the one that will define the next decade of accounting practice in Nepal: if AI cuts a six-hour task to two, what do you charge? The next section is a short closing essay on the ethics of charging for AI-assisted work — and why your answer is, eventually, part of how the profession recovers what it was always supposed to be about.