I’ve been thinking a lot lately about the gap between using AI and trusting AI.
Most of us using AI in our accounting work have quietly crossed a line we haven’t named yet. We’ve stopped being the person who writes the memo. We’ve become the person who decides whether the memo is good enough. That’s a real change — and most of us are doing it badly.
Here’s what I mean. Enterprise AI failure rates hit 80% in 2025 — twice the rate of traditional IT projects. In most of those cases, the AI itself wasn’t the problem. The failures came from vague instructions, poor integrations, and no real way to check whether the output was any good. That last one is what we need to talk about.
The “Vibes Check” Isn’t Quality Control
If you’re using AI to draft client memos, audit checklists, or disclosure language — and your review process is “skim it, nothing looks wrong, send it” — you’re doing what the engineering world calls vibes-based evaluation. It’s the equivalent of accepting a first-year associate’s “looks good” and forwarding it straight to the client.
Think about it this way: you’d never accept that from a staff associate. You check the workpaper against the program. You verify the numbers tie. You make sure the conclusion is supported. You have a quality bar, and you apply it. The AI is your associate. You’re still the reviewer.
What an AI Eval Actually Is
The structured alternative is called an AI eval — short for evaluation. The process is five steps: define what good looks like for this specific output type, generate ten samples (not one), score each against binary pass/fail criteria, diagnose what went wrong, and fix it.
The ten-sample rule matters. One good output tells you what the AI can do on its best day. Ten outputs tell you what it does on average — and averages are what your clients experience.
A Real Example: The §1231 Memo
A CPA uses AI to draft a memo summarizing a client’s Section 1231 gain position. The prompt: “Summarize the tax treatment for Section 1231 gains.” The AI produces a clean, well-organized memo that correctly describes the general rule. It looks great.
What it doesn’t mention: the Section 1231 lookback rule — the provision that recharacterizes gains as ordinary income if the taxpayer had §1231 losses in the prior five years. That rule was central to the client’s situation. The client gets the memo. The memo is wrong in a way that sounds completely right.
When you run this through an eval with four pass/fail criteria, one check fails 8 out of 10 times: does the output address the lookback provision? That pattern is a diagnosis.
Two Failure Types, Two Fixes
There are two types of AI failures, and knowing which one you have tells you exactly what to fix.
A specification failure means the prompt was too vague — you never asked for the lookback rule, so the AI never included it. Fix: rewrite the prompt to be more specific. The more context and detail you give the AI about all the rules that could apply, the better it will perform.
A generalization failure means the AI understood your intent but applied it inconsistently, especially on complex fact patterns. You asked for the right thing, got it right 6 times out of 10, and the other 4 had problems. Fix: add examples to the prompt — show the AI what a complete, correct memo looks like.
What This Means for How You Work
Two things, practically speaking.
First, before you use AI for any client-facing output, you need a written quality bar. Not a gut feeling — a short pass/fail checklist specific to each output type. If you can’t define what good looks like before you start, you’re not in a position to evaluate whether AI produced it. And honestly, that’s a signal: if you can’t articulate the quality bar, you probably shouldn’t be using AI for that output in production yet.
Second, “human in the loop” has to mean something specific. A partner skimming an AI-drafted memo is not quality control. Real human-in-the-loop means the person who holds the technical intent of the engagement systematically checking whether the AI output actually serves that purpose. The profession’s quality standards don’t change because AI wrote the first draft.
Key Takeaways
- When you use AI to produce work product, your role has shifted from creator to evaluator. That takes a different skill than writing better prompts.
- An AI eval is five steps: define criteria → generate 10 samples → score pass/fail → diagnose failure type → fix and retest.
- Risk checks often matter more than capability checks — what could this output be missing? What could a client misread?
- Specification failure = tighten the prompt. Generalization failure = add examples.
- Your judgment is the asset going up in value. AI can draft the §1231 memo. You’re the one who knows whether it’s right.
Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: [lesson link]


Leave a Reply