Yesterday a World of Warcraft harness showed up on r/PromptEngineering. Step 2 is the twist: the most important options in the menu are the ones that make the model do nothing.
What’s new
u/Individual_Cold_4119 built SageCraft, a harness that plays WoW with a decision model called Sage. They say up front that they co-founded Levanto, the company behind Sage, so read it with that in mind.
The model gets a screenshot, the game state, and an explicit menu of actions. That’s all. It doesn’t type free-form commands. It picks from a list. That matters because a closed list is easy to validate. You can check in one line of code whether the answer is on the menu, and you never have to parse a paragraph of model prose to work out what it meant.
Here’s the menu from one of the final decisions of the run:
- attack_mob_level_1
- reject_selected_target
- change_search_strategy
- ui_blocked
- recover_now
- cannot_assess
Sage picked attack_mob_level_1 in about 0.9 seconds. The harness checked the target, then ran a guarded burst of three Smite casts. About 35 seconds later it read level 5 twice, and the dwarf priest was there. Reading the level twice is a small detail worth copying: the harness confirms the result before it believes it.
The whole campaign was 4,573 decisions over 32 sessions. The author supervised it and fixed the harness between sessions, but the last 18 minutes ran with no human input.
The twist
Only one of those six options moves the character forward. The other five are ways to say “not yet.” That ratio is the whole idea. Most of the menu exists to give the model a legal way to decline.
And there’s a rule underneath them. If the model’s choice is missing, late, or outside the menu, the harness does nothing and asks again. It never invents a fallback action.
A lot of agent setups do the opposite. When the model is confused, something downstream guesses, and the guess wrecks your state. Picture a script that defaults to “attack the nearest thing” whenever the model times out. Now you’re fighting a mob you never chose, at a bad moment, because a timeout got treated as a decision. As u/ExplanationSome9004 put it in the comments, baking in “I literally cannot tell what’s happening” as a first-class option is a smart move. I agree.
Steal the pattern: a mini-workflow
- 🧭 Give the model a closed menu of actions. One action per line, named so a human can read the log. A name like attack_mob_level_1 tells you what happened without opening a single screenshot.
- 🛑 Add at least one “don’t act” option. In this build that’s reject_selected_target, ui_blocked, and cannot_assess. Each one covers a different kind of “no,” so the log tells you why the model held back.
- 🔁 Add a recovery option, like recover_now or change_search_strategy, so the model can say “my approach is stuck” and not only “this one step is bad.” Without it, a model stuck in a loop can only keep rejecting targets forever.
- 🛡️ Put a guard between the choice and the execution. Here the harness verifies the target before the Smite burst fires. The model proposes, the harness checks, and only then does anything happen.
- ⏱️ If the answer is late, missing, or off-menu, do nothing and re-ask. No hidden defaults. A retry costs you a second or two. A wrong guess can cost you the whole run.
Pro tips
- Keep the “can’t tell” option in the menu even when the model seems confident. It’s cheap, and it keeps a confused model from guessing.
- Log every decision with the menu it saw. 4,573 decisions across 32 sessions is the kind of record that lets you fix a harness between runs. When something goes wrong, you can see whether the model chose badly or the menu offered the wrong choices.
- Name your options for the person reading the log at 2 a.m., not for the model. If you can’t tell what happened from the option names alone, rename them.
- Treat the author’s open question as a design test for your own menus: should “cannot_assess” stay one escape hatch, or split into specific missing information (no target visible, health unreadable, and so on)? Splitting gives you better debugging data. One hatch keeps the menu short and quick to choose from. The author hasn’t compared menus in a controlled test and says so, so nobody knows the answer yet.
Honest caveats
This is one campaign, supervised, with fixes between sessions. The author doesn’t claim this is the best menu, and neither should we. The post has 0 upvotes as I write this, so treat it as an interesting design note and not a proven result. Level 5 is also early in the game, so we don’t know how the pattern holds up against harder fights or messier screens.
Call to action
The harness is open source and there’s a video in the comments of the original post. Go watch the dwarf level up, then look at your own agent and ask what it does when it can’t tell what’s happening. If you have a view on splitting the escape hatch, drop it in the thread. The author is asking. ⚓
Frequently Asked Questions
Q: Should ‘cannot assess’ be split into specific missing-information options?
It depends on how noisy the game state gets. If the model often mixes up a UI block with a missing health bar, separate those cases so the harness knows what to fix. If the menu grows too long, though, you risk slower answers and analysis paralysis, so one generic ‘confused’ option may be the better call.
Q: How many options is too many for the model?
Commenters pointed out that more choices slow the model down, so keep the menu as short as you can while still covering the outcomes that really happen. Add a new option only when your logs show the model repeatedly confusing two specific actions. Small menus with options that are easy to tell apart tend to work best.
Q: What should the harness do when the model picks the generic confused option?
One suggestion was to re-prompt before accepting the answer. Send the model a tighter crop of the screen or a different camera angle, then ask again. Only treat the result as a real ‘cannot assess’ if the second read is still unclear, so the escape hatch stays useful instead of becoming a shortcut.
The six choices I gave a decision model right before it reached level 5 in WoW
by u/Individual_Cold_4119 in PromptEngineering