Twenty-five product launch pages went through an AI auditor, and the first pass came back looking clean, organized, and completely wrong in five different spots.
The maker behind AllyOtter, a tool that checks WarriorPlus and JVZoo launch pages against network review rules, ran this exact test. u/investigatormaker built the checker to catch compliance problems before a page goes to a reviewer, so a vendor doesn’t get bounced for something small and fixable. The first audit run looked confident. It was also wrong in ways nobody caught until a human read the actual drafts.
Here’s what the auditor got wrong, and every one of these looks reasonable until you check the source text:
- It flagged “get rich quick” when that phrase sat inside a “This is NOT for you if…” list, meaning the page was disclaiming it, not making the claim.
- It read a “30-Day Action Roadmap” as a second refund period, which it wasn’t.
- It flagged countdown code that belonged to a site-wide plugin, with no actual timer showing on that page.
- It marked a disclaimer as missing when a script filled it in after the page loaded.
- It called a plain Gmail address “no contact,” as if a working support email doesn’t count as contact information.
Each mistake alone sounds minor. Stack five of them across 25 pages and you’ve got a tool that erodes trust fast, because vendors start ignoring every warning once a few turn out to be false.
Why This Actually Matters 🎯
An audit agent that’s wrong 5 ways isn’t just annoying, it’s actively dangerous to the workflow it’s supposed to protect. If a vendor gets a note claiming their page violates rules it doesn’t actually violate, one of two things happens. They either panic and rewrite content that was fine, or they start distrusting every future audit, including the ones that catch real problems.
That second outcome is the expensive one. The whole point of an automated audit is to save reviewer time and build trust with the vendors submitting pages. A tool with a reputation for crying wolf does the opposite of both. The original poster’s fix wasn’t a smarter regex or a bigger rule list. It was a second AI pass built specifically to question the first pass’s findings before a human ever saw them.
How To Steal This For Your Own Pipeline 🛠️
The core move here works for almost any AI classifier, not just page audits. Anywhere your agent flags, tags, or scores something, add a verification step that forces it to justify the flag or drop it. Here’s the exact prompt the author used as that second pass:
“For each finding, quote the exact text or element from the page and say where it is. If the quote is inside a disclaimer, a ‘not for you’ list, a testimonial or example content, drop the finding. Say which source the rule comes from: something the reviewer asked us for, or the network’s published guideline. If you can’t quote it, drop it.”
Break down why this works:
- Quote-or-drop forces evidence. The model can’t flag something it can’t point to directly on the page. This alone kills hallucinated violations, since there’s no room for “it seems like” reasoning.
- Context exclusions catch the sneakiest false positives. Disclaimers, “not for you” lists, and testimonials are exactly where risky-sounding phrases hide on purpose. Telling the model to check location before flagging stops it from treating a denial as a claim.
- Source attribution prevents made-up rules. Asking the model to name where a rule came from, a specific reviewer request or a published guideline, stops it from inventing a policy that sounds plausible but doesn’t exist anywhere.
- A human still reads every draft. Of 20 notes that made it through this second pass, 2 got dropped entirely, 6 got rewritten because they credited the wrong rule source, and 6 more would have shown the vendor a false warning if the audit link had gone out unchecked.
That last number is the one worth sitting with. Even after adding a verification layer, human review still caught problems in 14 out of 20 drafts. The lesson isn’t “add one prompt and trust it.” It’s “add the prompt, then keep the human in the loop anyway.”
Tips And Tricks Worth Stealing 💡
- Apply quote-or-drop anywhere you classify content. Moderation bots, compliance checkers, resume screeners, ad reviewers, all of them benefit from forcing the model to cite its evidence instead of pattern-matching on vibes. One commenter, u/Advanced-Physics-977, mentioned running into the same issue with a moderation bot that kept flagging people over nothing, which tracks.
- Separate “what triggered this” from “who said this rule exists.” Those are two different questions, and conflating them is exactly how a made-up policy gets attached to a real page.
- Track your override rate. The author knew the fix worked partly because they measured it: 2 dropped, 6 rewritten, 6 saved from a false flag. Without that count, you’re guessing whether the second pass is pulling its weight.
- Don’t skip the human pass just because accuracy improved. A better prompt reduces errors. It doesn’t eliminate the need for someone to read the output before it reaches a real person.
If you’re running any kind of AI audit, moderation, or review pipeline, steal this exact structure: first pass flags, second pass verifies with quotes and sources, human reads every draft before it ships. Check out the original discussion for the full breakdown, and if you want to see the fixed checker in action, it’s worth a look. 🚀
Frequently Asked Questions
Q: Why does requiring exact quotes help reduce false positives?
When an AI has to quote the exact text and location, it forces verification instead of inference. The model can’t assume a “30-Day Action Roadmap” is a refund period or miss a dynamically-loaded disclaimer, it either finds the text or drops the finding. This simple constraint surfaces where the model is guessing.
Q: How do you prevent the model from confusing your internal rules with the target’s actual guidelines?
Make source attribution explicit in your prompt. Ask the model to name which guideline each finding comes from, a client requirement or a published network rule? This forces the model to distinguish between your policies and external requirements, and reviewers catch misattribution immediately. As one commenter noted, source confusion accounted for half their false positives.
Q: Do you still need human review after a stricter verification prompt?
Yes. The original auditor still required human review, 14 of 20 drafts needed changes even with the stricter prompt. The verification prompt eliminates obvious hallucinations, but humans catch nuanced issues like misattributed sources and false audit links that the prompt logic alone misses. Think of it as filtering 80% of noise so reviewers can focus on judgment calls.
Q: Does this approach work beyond page audits, like moderation or support?
Absolutely. Any AI system that flags or makes claims should use this pattern: require exact quotes, verify context, force source attribution. Commenters reported the same problem with moderation bots flagging quoted text, the underlying issue is identical, and so is the fix.
Our AI page auditor was confidently wrong 5 ways. The verification prompt that fixed our pipeline
by u/investigatormaker in PromptEngineering