Give an optimizer a handful of failed answers and it will hand back a revised prompt that looks brilliant. Then you run it on those same examples and everything passes. That feels like progress. Often it is just a prompt that got very good at one small set of questions, the way a student who memorized last year’s exam looks great until the teacher changes the questions.
The Key Idea
A recent r/PromptEngineering post about the GEPA recipe in Reef Infra makes a point worth stealing: a prompt that fixes your examples still needs a second test. Several different checks sit inside one optimization loop, and each one answers a different question. Most teams treat “it passes now” as a single yes or no. This approach treats it as a ladder, and a prompt has to climb every rung before you trust it.
The reason this matters is simple. A rewritten prompt is shaped by the failures it was shown. It can pick up quirks that match those exact cases, like a particular phrasing, a certain output format, or a keyword that happened to appear in every failing question. Those quirks look like skill on the training examples and turn into noise everywhere else.
Old Way vs. New Way
The old way: collect failures, ask a model to rewrite the prompt, re-run the failures, see green, ship it. It is fast, and it feels rigorous because there is a test at the end. But the test is the same material that produced the fix.
The better way, based on how the GEPA recipe is described:
- Pick a parent prompt from an archive of candidates, so you are never stuck with a single line of descent.
- Use execution traces to reflect on a small training minibatch. The traces show where the model went wrong, not just that it did.
- Rewrite one prompt component. Changing one piece at a time makes it clear what helped.
- The child has to beat its parent on that minibatch before it goes any further.
- A child that passes gets a full validation run. The prompt you actually serve is chosen by mean validation score.
- A separate held-out test tells you how it performs outside the search process entirely.
The old way has one check, and it uses the same examples that inspired the fix. The new way has three, and each one is harder to fool than the last. 🧪
What Each Check Actually Tells You
- Minibatch: Did the proposed fix help locally? This is a cheap, quick filter. It throws out rewrites that make things worse or change nothing, so you do not waste a full run on them.
- Validation: Does this child deserve to replace the currently preferred prompt across a broader set? A fix can win on five examples and lose on fifty. Averaging over the whole validation set exposes that.
- Held-out test: Does it hold up on data the search never touched? This is the only number that estimates how the prompt behaves in the real world.
Skip the third one and “the optimizer fixed the failures” can quietly turn into “the optimizer memorized my review set.” Imagine a support bot prompt tuned on ten refund questions. It might score 100 percent on them, then stumble on a shipping question because the rewrite leaned too hard on the word “refund.” Only a set it has never seen would show you that.
How to Apply This to Your Own Prompts
- Split your examples before you start. Make three piles: train, validation, test. Do it on day one, not after you have seen the results. A rough starting point is half for train, a quarter each for validation and test, but even 20 examples split three ways beats no split at all. Shuffle first so each pile has a similar mix of easy and hard cases.
- Only show the optimizer the train pile. Failures it learns from should never come from validation or test. If a good failing example turns up in the wrong pile, resist the urge to move it.
- Make every new version beat its parent. A rewrite that cannot beat the prompt it came from does not move forward. Keep the old versions in a list with their scores, so you can go back if a later idea goes nowhere.
- Pick the winner on validation, not on the examples that inspired the fix. Use the mean score across the whole set. Watch for a winner that gains a little on average but collapses on one type of question, and look at the per-example scores before you commit.
- Run the test pile once, at the end. If you keep peeking and tweaking, it becomes another validation set. Write down the score the moment you see it, and do not rerun to see if it improves.
The Sneaky Part
One commenter nailed the real-world version of this. Separate splits sound obvious until you are three iterations deep and realize your validation set has been steering the search the whole time. Every time you look at validation results and change something, a little of that set leaks into your prompt. It rarely feels like cheating. It feels like being careful. The fix is boring discipline: keep the test pile locked away and treat it like a sealed envelope. 🔒
If you suspect a leak already, the cure is cheap. Write ten fresh examples by hand, from real user questions if you have them, and score your current best prompt on those. A big drop from your validation score is the sign that the search has been learning your set instead of your task.
Your Next Move
Open your current prompt project and check how many of your “proof” examples are also the examples you used to write the fix. If the answer is most of them, carve out a fresh set today and run the prompt against it. You might be pleasantly surprised. You might also find out the prompt was only ever good at your homework. Either way, you will know, and that is the whole point.
Frequently Asked Questions
Q: Why do I need a separate test set if validation already proved the prompt works better?
Validation shows improvement over your current prompt *on that specific batch*, but if you’ve optimized multiple times against the same validation set, the optimizer might just be fitting quirks in that data. A held-out test set, one the optimizer never saw, proves the improvement actually generalizes to new problems, not just your review examples.
Q: How do I catch eval drift before my validation set starts steering the search?
This creeps up gradually over multiple optimization rounds. Watch for a gap: if your validation scores keep climbing but real-world performance plateaus, your validation set has become part of the optimization loop. Set your splits before you start and treat the test set as read-only, don’t even look at it until optimization is done.
Q: What’s a practical way to split my data in real projects?
Roughly 60% minibatch examples to guide the optimizer, 20% validation to pick the “winning” prompt, and 20% held-out test to verify it actually works. The key: finalize these splits upfront and keep the test set completely separate. This removes the temptation to tweak based on test performance.
The prompt that fixes your examples still needs a second test
by u/siddhantfuture in PromptEngineering