Two Customers Cancelled, and Your AI Summary Blamed the Invoices. Here Is the Fix

Short version: when an AI shortens a transcript, it tends to turn “two people complained” into “this causes that.” A prompt that makes the model keep case counts, limits and who said what stops that from happening.

The problem in one example

A Reddit user in r/PromptEngineering tested a summary prompt on a short invented transcript excerpt:

“Two customers cancelled their subscriptions. Both told our support team that their first invoices were confusing. We haven’t checked how common that complaint is among other customers.”

The tempting summary is “confusing invoices cause cancellations.” That sentence adds a cause nobody established. It also drops the fact that there were only two cases. The speaker reported two cancellations and two complaints. They never said invoices caused anything, and they said outright that they hadn’t checked how common the complaint was.

The poster treated that causal wording as a failure condition. I like that choice. A prompt is much easier to improve when you define in advance what a bad output looks like.

What the prompt does

The prompt works because each instruction blocks a specific way a summary goes wrong:

  • Keep counts and limits. Every claim carries its stated number of cases and any uncertainty the speaker voiced.
  • Attribute to the speaker. If the speaker is unknown, the summary says so instead of guessing.
  • Preserve nested attribution. A guest reporting what customers told support must stay a guest reporting what customers told support. Flatten it to “customers said” and you lose the chain of custody. A commenter in the thread made this exact point.
  • Support each claim with a short exact excerpt.
  • Never invent a locator. Keep a supplied timestamp or line number. If none exists, say it wasn’t supplied.
  • Separate complaints from causal findings.
  • Hold back recommendations. Proposals only appear if requested, in a section labeled Model suggestion, so they never pass as the speaker’s advice.

How the poster tested it

The checks were simple and concrete. Did the output keep “two”? Did it show the speaker was relaying customer complaints? Did it keep the note that prevalence hadn’t been checked? Did it avoid invented speakers, timestamps and unrequested advice? If a survey was proposed, was it clearly labeled as the model’s idea? For the real podcast transcript, they also checked the quoted excerpts against the recording.

That last step matters. The excerpts are only evidence if someone confirms they match the source.

One pass or two?

The open question was whether to handle nested attribution in a single pass, or to extract claims and excerpts first and summarize after. The poster is also clear that an extra pass doesn’t make a summary reliable on its own. Either version has to pass the same checks.

One commenter pointed at the real weak spot: the shortening instruction. Compression is where attribution tends to get lost. Their fix was to tag every claim with its evidence level, such as observed in 2 cases, inferred, or unsupported. That fits well with this prompt.

My take is that a one-pass summary is fine for short, simple transcripts. For long transcripts with several speakers, extracting claims and excerpts first gives you something to audit before the summary exists. Whichever you pick, run the checks.

Use cases

  • Summarizing customer interviews and support call recordings
  • Turning podcast or webinar transcripts into notes
  • Pulling findings out of user research without inflating anecdotes
  • Meeting summaries where “someone said” must not become “the team decided”

Prompt of the Day

This is the prompt as the poster shared it:

“Summarize the main claims using only the supplied transcript. For each claim you include, retain any stated number of cases and any uncertainty or limitation. Attribute it to the speaker if identified; otherwise say the speaker is unidentified. Preserve nested attribution, such as a guest reporting what customers told support. Include a short exact excerpt as support. Preserve a supplied timestamp or line identifier; if neither exists, say the source locator was not supplied. Do not invent one. Keep reported complaints separate from causal findings. Do not add follow-up recommendations unless requested. If requested, put your proposals in a separate section labeled Model suggestion and distinguish them from the speaker’s recommendations. Do not state a conclusion that the transcript does not support.”

To try the evidence-tag idea from the comments, add one line: “Label each claim as observed (with case count), inferred, or unsupported.”

Call to action

Take the two-customer excerpt above, run it through your current summary prompt, and see whether “two” and “we haven’t checked” survive. If they don’t, paste in the prompt above and run it again. Then tell me which pass setup worked better for you, one pass or two.

Frequently Asked Questions

Q: How do I keep nested attribution (like “customer told support, support told speaker”) intact when summarizing?

Treat attribution as a structural requirement in your prompt, not just a style note. Use source-depth labeling: tag each claim as “speaker direct,” “speaker reported,” or “speaker inferred.” This forces the model to preserve the chain of custody, so when a customer complained to support and support reported it to the speaker, that distinction survives the summary.

Q: Why does compression and shortening make important details like sample sizes disappear?

Compression is where attribution dies silently. When you ask the model to “make it shorter” in a second pass, that’s where “two customers” becomes just “customers.” Fix this by requiring every claim to carry an evidence tag (“observed 2 cases,” “inferred,” “unsupported”) and explicitly rule that any sample size from the source must survive compression. Then test the shortening pass separately to catch where numbers vanish.

Q: How do I prevent causal language from appearing when the source only reports a correlation?

Ban causal verbs (“caused,” “led to”) unless the source itself uses causal language. Pair this with your source-depth labeling: “inferred” and “reported” claims should never get causal framing. If customers complained about invoices but didn’t say invoices *caused* cancellations, your summary shouldn’t either.

Q: Is a two-pass prompt approach riskier than a single pass?

Yes. Multiple passes give the model more chances to drop details between steps. A single well-structured pass with clear source-depth labeling tends to preserve attribution better than a two-pass approach where details slip away in the middle step.

Prompt review for keeping two customer anecdotes from becoming a causal claim
by u/Pleasant-Step-2157 in PromptEngineering

Scroll to Top