Your ChatGPT Score Is Rigged

Try this test right now. Hand ChatGPT four versions of the same page. Keep three normal, and wreck the fourth on purpose with typos and broken logic.

Then ask it to score all four out of 10. No cheating!

I ran across a Redditor who did exactly that and posted the results, and they should worry you a little. Every version came back a 7 or an 8, this Redditor found, even the one wrecked on purpose. That one still landed a 7, with a note about “minor tightening.”

Here’s why that happens. A single number score has nowhere to go. A 4 feels rude. A 9 feels like flattery. So the model parks everything in the same narrow band between 7 and 8. It’s not evaluating your work, it’s picking the answer that won’t upset you.

The fix isn’t a smarter scoring prompt. It’s dropping the score entirely and forcing the model to make an actual decision.

🧪 The Test: Four Prompts That Won’t Let It Hedge

This Redditor built four prompts that trade “rate this” for “choose, rank, or cut this.” Run your own draft through each one and watch the compliments disappear.

Each one leans on a different trick. Forced ranking removes the tie option. Defining the scale first works like a mini example set, except the model has to build the example itself. Deletion and the gap question both swap a vague rating for a specific, checkable claim.

1. Forced ranking, no ties

Rank these [5] sections from strongest to weakest. No ties, no explanation before the ranking. Then say what the weakest one is missing.

Ranking removes the escape hatch. The model can’t call everything a 7 when it has to put one section dead last.

2. Make it define the scale first

Before scoring, write one sentence describing what a 3, a 6 and a 9 look like for [this kind of document]. Then score my draft and quote the line that put it in that band.

This is the one that actually changes behavior. Once the model commits to what a 6 looks like, it can’t quietly grade your draft a 7 anyway. You also get the exact line it used as evidence, which is where the real feedback lives.

3. Cut instead of rate

If you had to delete [20]% of this without losing meaning, which lines would go? List them, and only them.

Deleting is a decision. There’s no polite middle ground for “which lines are dead weight,” so the model has to actually pick.

4. The gap question

What would need to be true for this to be a 9? Be specific about what is missing, not what could be improved.

This one skips vague notes like “could be tightened.” It forces a concrete list of what’s actually missing.

🔍 What the Results Actually Mean

Run prompt 2, and if the model still can’t describe a real difference between a 3 and a 9, that’s your answer. It never had a scale. It had a habit.

If prompt 3 comes back empty, or with one throwaway line, the model is still hedging even after you took the number away. Push back. Tell it you expect a real list.

One commenter on the original thread put it well. A score like this gets optimized to keep you talking, not to send you back to fix your work. These four prompts get you past that.

This isn’t just for blog drafts either. Try it on a cover letter, a pull request description, or a sales email. Anywhere you’re used to getting a polite 8 out of 10, run prompt 2 first and watch the real gaps show up.

💡 Extra Tips

  • ✅ Run the “mangled on purpose” test on yourself first. Feed ChatGPT a draft you know is weak, and see if prompt 2 catches it. If it still scores a 7, don’t trust its feedback yet.
  • ✅ Save these four as reusable prompts instead of retyping them. This Redditor keeps them as saved prompts with blanks to fill in, and that’s the whole point. A real evaluation prompt should work on any draft, not just one.
  • Keep a rough log of which band your drafts land in. If prompt 2 keeps putting your work at a 6, you’ve found your actual ceiling, not the one ChatGPT was handing you for free.

Prompt of the Day: Combine two of these into one gut check. “Rank these sections from strongest to weakest, no ties. Then tell me what would need to be true for the weakest one to beat the strongest.” One prompt, two escape hatches closed.

🚀 Go Break Your Own Draft

Stop handing your work to ChatGPT and asking it to be nice about it. Ask it to rank, cut, or define the scale first, and watch the polite 7s disappear!

Head over to the original Reddit thread for the full back-and-forth. There are some sharper counterpoints in the comments worth reading too.

Frequently Asked Questions

Q: Why does ChatGPT score everything 7 or 8?

It’s playing it safe. ChatGPT is optimized to keep conversations going and not offend anyone, so without clear definitions of what a 4 or 9 looks like, it defaults to the social middle. Low scores feel rude, high scores feel like flattery, so everything lands in that 7-8 sweet spot.

Q: What other tricks work besides the four in the post?

A couple of things commenters mentioned: try descriptive grading instead of numbers (“professional-casual,” “too long/right/too short”), or reframe the context (“my intern wrote this”). Both force real decisions instead of letting the model hide in the middle with a safe score.

Q: Does telling ChatGPT someone else wrote it actually work?

Worth testing. Changing how you frame the authorship, like saying an intern or colleague wrote the draft, can shift the model’s feedback. Some people found it makes ChatGPT less polite and more critical.

Q: Why is the “delete 20%” method better than scoring?

Because there’s no middle ground. Deletion is binary, so a line either gets cut or it doesn’t. The lines ChatGPT picks first are the ones it really thinks are weak, which is way more useful than buried in a number.

Every draft I give ChatGPT comes back a 7 or an 8, including the ones I wrote badly on purpose
by u/Ok_Negotiation_2587 in ChatGPTPromptGenius

Scroll to Top