Issue #29: My AI passed its own work. So I built one to check it.
My review agent gave a clean bill of health. Here's why I didn't take its word for it.
Welcome back to The Customer Continuum. Issue #29.
Here’s an uncomfortable confession to open with. While I was building last week’s issue, an issue about an AI agent trained to protect customer confidences, my own AI assistant shipped two mistakes into my production line. A download link that pointed nowhere, and a paragraph confidently describing files from the wrong week. I caught both, but I caught them the way you catch things at 11pm, by luck and irritation rather than by process. The tool did impressive work all day and then quietly handed me two errors wearing the same confident voice as everything it got right.
That’s the problem this issue is about, and it isn’t bad AI work. It’s good AI work, which is more dangerous, because good work earns the trust that stops you from checking.
So this week I ran the experiment the last two issues were building toward. I took the agent I built last week, the one that reads a pile of customer reviews and turns them into themes, signals, and quotable customer lines, and I put its work in front of a second agent whose only job is to read another agent’s output and say whether it’s safe to ship. One AI does the work. A different AI checks it. This is the build report from the first time they ran as a pair.
What I ran
I built a fresh test batch: sixty synthetic customer reviews for a fictional analytics company, written to read like a real export, with the same uneven voices and lengths you’d pull off a review site on any Tuesday. Buried inside were six reviews that should never leave the building. Two mentioned unreleased product plans a vendor had shared privately, one paired a genuine growth story with a discount the reviewer said outright not to quote, one asked the reader to keep a competitor switch secret, and two, written as separate reviews from the same company, only became dangerous when you connected them: one hinted at big changes the reviewer couldn’t discuss, the other casually mentioned a merger closing in Q3. Read together with their attribution, they disclose a deal that isn’t public.
I also planted the opposite trap. Several reviews used dangerous-sounding words in completely safe ways: a customer talking about their own roadmap, their own beta program, a discount published on the vendor’s public pricing page. An agent that panics at keywords would bury those and starve the quote bank. An agent with judgment keeps them.
Last week’s test proved the agent protects its quote bank, the list of customer lines cleared for marketing and sales use. This batch attacked everywhere else: the theme summaries where the agent paraphrases at volume, the signal lists where it extracts growth and competitive intel, and the attribution that connects a quote to a company. Those are the surfaces where a confidential detail can slip through in the agent’s own words, with nobody watching.
The agent’s run
It went six for six. Every planted review was withheld, nothing from them bled into the summaries, every safe decoy survived, and it connected the two-review trap on its own, pulling both reviews because together they disclosed a deal that wasn’t public. If you read last week’s issue, none of that should surprise you, since protecting confidences is the exact job this agent was built to do. The point this week isn’t that the agent is good. It’s really about what happens next, because a perfect run is calibration evidence, not proof, and the whole reason checks get skipped is that the work keeps looking this good right up until it doesn’t.
The checker’s run
The checking agent got a fresh, empty workspace, its instructions, and nothing else. No knowledge of my test, no answer key, no history.
Its first response was a refusal. I’d assembled its input wrong and left out the one thing it exists to review, the agent’s actual output, and instead of improvising it returned a single line saying validation couldn’t proceed, named exactly what was missing, and declined to produce the synthesis itself when a stray instruction in my paste invited it to. Then it did something I didn’t ask for. Working from the raw reviews alone, it listed the six reviews it would be watching when the output arrived, by ID, with the reason for each. That’s the same six I’d buried in the data. It found every landmine blind, before ever seeing the work it was supposed to check, and its first real catch of the day was me, sending it a malformed request.
On the rerun, with all three pieces in place, it returned a pass with notes. It confirmed all six sensitive reviews were withheld with zero reproduced language, and it showed its work, naming the specific leak paths it searched: unreleased feature names, the discount figure, the merger details, whether the hidden competitor mention could be traced. It raised no false alarms, which matters more than it sounds. A checker that invents problems to justify itself trains you to ignore it, and an ignored checker is more dangerous than no checker, because it lets you believe the work was reviewed.
Then it flagged something neither the agent nor I had thought to look for. One quote in the approved bank named a competitor in a head-to-head comparison. The customer said it publicly, so it’s perfectly clean on confidentiality, and it still carries risk, because a named competitive claim in external collateral is a legal question, not a privacy one. The working agent was never built to see that category. The checker saw it anyway, along with a recommendation to mark the withholding notes as internal-only so the description of what got held back doesn’t travel into a deck either.
Read that sequence again as a team lesson. The worker was excellent. The checker refused a broken request, found every trap without an answer key, confirmed genuinely clean work without crying wolf, and then caught a risk in a category nobody assigned it. Neither one alone gives you that. The pair does.
What this means for your team
Here’s the takeaway in one sentence you can say in a meeting tomorrow: build the agent, build the checker, and never trust one without the other.
The checker isn’t there because the worker is bad. Mine just ran clean under real pressure. I created the checker because you can’t know the worker was good until something independent tells you, and “it looked fine to me” is not independent, as I proved personally two weeks ago at 11pm.
Every team adopting AI right now is one impressive week away from dropping the review step, and the review step is the entire difference between a tool you trust and a tool you got lucky with.
If your team runs even one AI agent on customer-facing work today, reviews, case studies, campaign copy, community replies, this pattern is the cheapest insurance you’ll buy this year. Yes, it’s an extra step. It costs one extra prompt and ninety seconds per artifact. But what’s actually at stake isn’t the artifact but the trust you’ve built with your customer that you’ll represent them accurately every single time; and the trust you’ve built with your boss, your leadership, and your peers that work coming out of your team is safe to put their names behind. Both kinds take years to earn and one shipped mistake to lose, and you never get to choose which piece of work is the one that loses it.
What’s next
Next week, the checking agent and its full testing track become a published part of the kit, with an upgraded system prompt folding in everything these runs taught me, so your version starts where mine finished.
This week’s free starter: the Paired Run Protocol
The free starter is the exact protocol I used for this test, written so you can run it against any AI agent your team already uses. It covers the setup that keeps the check honest (a fresh workspace with no shared history between worker and checker), the three things every checking request needs (a short brief of what the work was supposed to be, the source material, and the finished work itself), and a compact starter checker prompt you can adapt this afternoon. Run it once against something your team shipped last week. What it finds, or doesn’t, will tell you more about your AI workflow than any vendor demo.
For paid subscribers
You already have the whole library. The Copilot and all 135 agents landed in your account on day one, and the full Chief of Staff validator system prompt, the same one that ran everything above, has been in your kit for two weeks. This week’s job is pairing it with what you built last week.
If you built the two Voice of Customer agents from the Week 6 doc, your fastest path is one message: give the validator the task brief, your source reviews, and your agent’s output, in that order, and ask if it’s safe to ship. The Week 7 build doc in your paid folder walks the pairing step by step, including the input contract and what a pass, a pass with notes, and a revise verdict each mean for what you do next.
For the readers following the practical test suite doc: this run delivered the Test 9 scenario, quiet reproduction across an agent’s output surfaces, against the live synthesis operator, and the Test 8 attribution-drift design folds into the v3 prompt shipping with the next issue.
Access the resources here:
Next week, the validator goes into the kit.
— Kevin
P.S. If this issue changed how you think about trusting an AI’s clean-looking work, forward it to one customer marketing leader whose team started using AI this year. That’s how this newsletter grows.





