40 out of 40. That’s how many support agent conversations a blind judge labeled as AI in a test run by a team on r/PromptEngineering. The same judge correctly spotted 5 out of 5 human-written replies, so the test wasn’t rigged against the bots.
The agents were accurate. They still felt like bots. Here’s what the test found.
The setup
The team built a fake car dealership as the knowledge base, so any fact outside it was provably invented. A made-up price, a wrong range figure, or an invented warranty term could be caught by simple lookup, with no guessing about what the “real” answer should be. They tested two models (Gemini Flash-Lite and a self-hosted Qwen3-14B). The judge only saw transcripts and had to guess human or AI. It never saw the prompt, the model names, or the knowledge base, so it could only judge how the replies read.
The insight breakdown
In 34 of the 40 conversations, the judge named the same tell: the ending. Between 92% and 96% of replies finished with “Would you like to…” or “Let me know if…”. A single closer is harmless. Repeated on every turn, it becomes a fingerprint. Real support staff answer the question and stop, and the silence is part of how a human sounds.
The cause was in their own prompt. Every example reply ended in a question. The model copied the examples, not the rules. If you’ve ever written a rule and watched the model ignore it while mimicking your sample, this is that. Examples act as the strongest signal in a prompt, because they show the model exactly what a finished answer looks like. A rule is one sentence. An example is a pattern repeated several times.
Pressure made it worse:
Customer: how much for the ranger
Agent: 14.2 lakh. Would you like to know more about its specs or book a test drive?
Customer: you already said that
Agent: You’re right, I apologize. Would you like to hear about the 402 km range?
Customer: just give me a number
Agent: The Ranger has a 402 km range. Would you like to schedule a test drive?
The customer asked for a price three times and got a pitch three times. Asked again later, it gave a different price. Notice that the apology didn’t change the behavior at all. The agent acknowledged the complaint and then did the exact thing the customer was complaining about. A human would have just repeated the number.
The fixes, with numbers
- Replies ending on a question or offer dropped from 74% to 30%.
- Average reply length fell from 29 to 22 words.
- The judge picked “human” 3 times out of 9.
A modest result, honestly. But it moved, and that’s the point. Shorter replies and fewer closers both pushed in the same direction, and neither required a bigger model or a fine-tune. It was prompt edits, measured one at a time.
3 practical applications
🧭 Audit your examples before your rules. Count how your sample replies end. If every one closes with a question, your agent will too. Add a plain rule that most messages should just end, then show examples that actually stop. A good mix is roughly two out of three examples ending on the answer itself, with the odd follow-up question only where the customer’s request was truly unclear.
🧭 Use placeholders in examples. The team found that a believable fake price in an example got reused in real answers. They switched to <price>. Same trick works for names, dates, and policies. If your example says “our return window is 30 days,” some of your customers will hear that, whatever your actual policy says.
🧭 Give numbers their own rule. Qwen said 350 km while “402 km” sat right there in its context. A dedicated rule such as “quote figures exactly as they appear in the knowledge base” is cheap insurance. Pair it with a check in your test set: pick five facts, ask for each one several ways, and flag any answer that drifts from the source.
Tips and pitfalls
- Pitfall: a fix can overshoot. Their first patch made the agent deny a price it had given one message earlier. They added a separate rule for being challenged: confirm what was said, don’t flip.
- Pitfall: tidying the prompt is not fixing the prompt. Rewriting the same rules into cleaner sections scored 36% before and 36% after. Only the specific fixes moved anything.
- Tip: test with a blind judge and a human control. If the judge can’t spot the humans, you can’t trust the verdict on the bots.
- Tip: measure one tell at a time. “Ending on an offer” is countable. “Sounds robotic” is not. Once you can count it, you can run the same transcripts before and after each edit and see whether the number moved.
One commenter noted the same pattern in long-form drafts. Models trained to be helpful want to hand something back every turn, so questions and offers pile up at the end. Banning the closer outright is the blunt version of what this team did more carefully. The blunt ban works for a quick patch, but it risks the overshoot problem above, where the agent goes cold when a follow-up would actually help.
Call to action
The team published the prompt, the chat, voice and WhatsApp variants, and all the numbers on GitHub (duvi-ai/agent-prompts). Pull your last 20 agent replies today and count how many end with an offer. If it’s more than a third, check your examples first. Then tell us what tell you keep seeing in your own agents.
Frequently Asked Questions
Q: Why do AI agents always end with questions like ‘Would you like to…’?
AI models trained to be helpful default to “handing something back” at the end of every response. The fix: teach agents that some replies should just end once the answer is complete. In this test, removing the closing question and setting hard response-length limits dropped question-endings from 74% to 30%, and ironically made conversations sound more human.
Q: Do better-written prompts actually make AI sound less like AI?
Not necessarily. Rewriting rules into cleaner sections didn’t budge detection rates (36% before and after). What worked: very specific structural fixes, placeholder values instead of real numbers, separate rules for edge cases, and rules for handling corrections. Generic “improve the prompt” doesn’t move the needle; surgical fixes to specific behaviors do.
Q: Why do models use wrong information or repeat examples instead of reasoning through answers?
Models tend to copy patterns and values directly from example replies in your prompt rather than apply the rules. Fix: use placeholders like <price> in examples so the model learns the structure, not the data. This prevents the agent from reusing old specs or made-up numbers in new conversations.
Q: How do you stop AI from contradicting itself when a customer pushes back?
Add a specific rule for handling corrections and challenges, separate from your general helpfulness rules. The team’s first fix made agents deny information they’d given moments earlier, so they added a dedicated rule that fixed the contradiction without breaking natural-sounding replies.
Q: What other patterns reveal AI besides the closing question?
Symmetrical structure, every paragraph getting exactly the same number of sentences, is another tell. Similarly, responses that are too uniform in length stand out as mechanical. Even the first two lines of a reply can signal AI if you know what to look for.
We had a blind judge guess human or AI on our support agent’s replies. 40 out of 40 said AI
by u/Outrageous_Mark9761 in PromptEngineering