A few weeks ago, GPTZero , a company built specifically to hunt down AI-generated content, went looking through PwC’s published research. What they found should make every accountant who uses AI pause: four “thought leadership” reports from PwC Middle East, at least one of them flagged with a 100% probability of being fully AI-generated once you exclude the references section. The report promoted a PwC framework called “Citizen Pulse” and claimed it was already in use by the governments of Denmark, Saudi Arabia, the United States, and Australia. GPTZero looked for public evidence any of that was true. There wasn’t any.
I use AI in nearly every deliverable I put in front of a client, so this story hit close to home. It’s not really a story about PwC being careless. It’s a story about what happens once AI adoption outpaces the controls built to check it, and I think every accountant needs to understand exactly what changed here, because it isn’t PwC-specific.
Only 10 of the report’s 17 citations even matched footnotes in the body of the report. It wasn’t caught internally, but an outside firm that built its entire business model around finding exactly this. One of GPTZero’s investigators pointed out that the report couldn’t even decide whether “Citizen Pulse” was a real product or a hypothetical one, flipping between the two throughout the document. That inconsistency, they said, is a classic symptom of text generated by a poorly-prompted AI model.
And PwC isn’t alone. In the past year, all four of the Big Four firms have been caught in some version of this. Deloitte issued a partial refund to the Australian federal government after a report it delivered turned out to have AI-fabricated citations. EY had to withdraw an entire study on loyalty rewards programs after a probe found apparent hallucinations and fake footnotes running through it. KPMG’s report on agentic AI was found riddled with citations that didn’t hold up. Four for four, in under twelve months, at firms with entire risk and quality functions built to prevent exactly this.
So what actually changed? Two things, and they’re stacking on top of each other. First, AI adoption inside firms moved faster than firms built review processes for it, the drafting got a new tool, but the review step didn’t get one to match. Second, and this is the part I think most people are missing: AI errors are now the story reporters and skeptics actively want to write. Five years ago, a hallucinated footnote in an obscure report would have gone unnoticed by everyone except the one person who clicked the link. Today it gets a dedicated investigation and a headline in half a dozen outlets within the same week. The mistake didn’t get worse. The scrutiny aimed at that one category of mistake got dramatically higher, and it’s not going back down.
That’s the environment every firm, not just the Big Four, is now publishing into. Smaller firms are more exposed, not less. A name-brand firm has a communications team ready to respond within the hour. Most firms don’t. For everyone else, the first time you hear about the error might be the same moment your client does.
Here’s the shift I’d make: stop treating AI-assisted review as a proofread. Proofreading catches typos, awkward phrasing, and tone. It does not catch a confidently written, completely fabricated fact. Which is exactly what a hallucination is. AI-generated text is specifically good at sounding right while being wrong, and nothing about how it reads will tip you off.
What catches it is a narrow, specific check: pull every named entity, every statistic, every citation, and every claimed relationship out of the document, and confirm each one traces back to something real. A twenty-minute checklist, run by someone who owns it every single time, the same way a tie-out happens every time on a financial statement, not just when someone remembers.
Key Takeaways
- The reputational bar for AI-assisted content is asymmetric. One hallucinated citation draws more attention than a hundred accurate reports, because the story writes itself.
- Four of four Big Four firms have now been caught in under a year. This is systemic, not a rare miss — “we’re careful” isn’t a control.
- Proofreading for tone is not the same thing as verifying a fact is real. Build the second check in, separately, as a named step.
- The tell is always specificity: invented products, invented client or government names, citation counts that don’t match the reference list.
- The real question isn’t “does this look right.” It’s “would this survive a hostile AI-detection scan” — because that scan is already running on published work across the profession.
Want the CPE credit? Take the full lesson on EverydayCPE and earn 0.2 CPE credits: When AI Errors Become the Headline


Leave a Reply