Three models, three rounds, one price tag: $1.0554. That’s the all-in cost of running one campaign brief through GPT Image 2.5 Sunburst, GPT Image 2.5 Flare, and GPT Image 2, eight images per round, 24 images total. The original poster, u/woulatte, ran the test in r/PromptEngineering, and the takeaway wasn’t which model won. It was what changed in the prompts between rounds.
The insight breakdown
Most people run one model, pick the best output, move on. This creator sent the same brief to three frontier models in parallel instead, with the estimated cost shown before every run, so there was never a moment of “wait, what did that just cost me.” Round 1 kept things simple: a pure photographic brief, no text layer, just composition, lighting, and product placement. Round 2 added a full typography layer that wasn’t there before, headline, subhead, badge copy, and the results shifted direction hard. Suddenly the model wasn’t just rendering a scene, it was balancing legibility against the photo underneath it, and that tradeoff showed up everywhere: cropped product shots, busier backgrounds, less confident composition. Round 3 was the smart move: freeze what’s already working, surgically replace only the parts that aren’t, and land on the actual catalogue image. That’s the round most people skip because it feels like less work than starting over, but it’s the one that actually gets you to a usable asset.
The costs tell their own story: $0.3384 for round 1, $0.357 for round 2, $0.36 for round 3. The price barely moved, three cents between the cheapest and the most expensive round, but the outputs moved a lot, which is the real finding here. Small edits to a brief redirect a model harder than most people assume, and that’s visible only when you compare rounds side by side instead of judging each round in isolation. A single typography line added in round 2 changed how the model weighted the entire frame. The creator also open-sourced the tool used to run these parallel batches, so the whole workflow, prompts, costs, model calls, output comparisons, is reproducible, not just described in a Reddit thread. That matters more than the specific numbers, because $1.0554 for this brief doesn’t tell you what your brief will cost, but the method for finding out does.
3 practical applications
- 📊 Run competing models in parallel on one brief before picking a favorite. A single model’s output tells you nothing about direction, only about one guess. If GPT Image 2.5 Sunburst gives you something close but not quite right, you have no way of knowing whether a different model would’ve nailed it on the first try, or whether the brief itself is the problem, unless you’re comparing outputs side by side from the same prompt.
- 📊 Log the cost per round next to the images. Once you see $0.34 buys you a full 8-image batch, you stop treating generation cost as a mystery and start treating it as a budget line you can plan around. That’s the difference between guessing at your monthly image spend and actually forecasting it, which matters the moment you’re running this across a dozen campaigns instead of one.
- 📊 Treat each refinement round as a controlled experiment. Change one layer, like typography, and note exactly what shifted, instead of rewriting the whole brief and losing the signal. If you change three things at once between rounds, you’ll never know which one caused the output to improve or degrade, and you’ll end up repeating the same trial-and-error loop on the next brief.
Tips and pitfalls
Keep an unchanged-prompt run in every round as a control. One commenter, Inevitable_Salary871, pointed out that without a baseline, you can’t tell if a change in output came from your edit or from model randomness. Image models aren’t deterministic between calls, so a shift in composition could just be noise, not signal, and you won’t know which one you’re looking at unless you have something to compare it against.
Watch for interaction effects. Another reply, from driftpath53, raised a sharp question: does the “freeze and surgically replace” approach break down when the new section interacts heavily with the part you froze? Adding a text overlay on top of a “frozen” product shot can still force the model to rebalance lighting or shadow underneath it, even if you didn’t touch that layer directly. Test that interaction directly instead of assuming isolation holds, especially on briefs where the added layer overlaps the existing composition.
Don’t optimize only for the best final image. u/-DirtNerd- made the case that a model landing 80% there in round 1 can be more useful in production than one that needs three rounds to get there, since consistency saves you rounds later. A model that’s slightly worse on peak quality but reliably close on the first attempt will cost you less time and money across a hundred campaign briefs than one that occasionally nails it but needs constant re-rolling to get there.
Call to action
The full brief text, all 24 images, and a short walkthrough video are sitting in the original discussion, along with the open-source repo the creator built for running these parallel comparisons. Worth a look if you’re iterating on image prompts and want to see the actual deltas, not just the finished shots.
Frequently Asked Questions
Q: Should I prioritize a consistent model or one with higher peak quality?
It depends what you value. If you want predictable results you can count on, a model hitting 80% consistently sounds better than one that’s all over the place. But if you like having options and can pick the best one from each round, maybe variance is actually an advantage. Either way, tracking both consistency and peak quality helps you decide.
Q: When does the surgical refinement approach (freezing parts) break down?
When typography or layout changes cascade through the rest of the composition. If you’re just swapping text that’s already positioned, surgical edits work great , but if adding or repositioning text changes your overall balance or color relationships, you might need to rethink more than just that one section. Running a full rewrite in parallel as a sanity check isn’t a bad idea when you’re unsure.
Q: How do I know if an improvement came from my prompt change or just random luck?
Keep an unchanged version running in each round as a control, then compare them directly. Otherwise you can’t tell if your edit actually helped. For campaign work, some people lock product details + copy and change just one variable per round , makes it way easier to see what actually moved the needle.
Q: What’s the real cost metric beyond generation price?
Count how many outputs still need manual fixes before they’re usable as ads. A cheap model might generate images you all have to touch up, while an expensive one gets you closer to done. That hidden labor cost can flip which model is actually cheapest overall.
Same creative brief, 3 models, 3 refinement rounds — what changed in the prompts?
by u/woulatte in PromptEngineering