TL;DR: A Reddit user shared part 1 of a long spec prompt that tells an LLM to build a Prompt Engineering Agent Workbench. It’s an app where you create, optimize, validate and version prompts inside projects. The core idea is simple: keep the system complex inside and the interface simple outside.
What the prompt actually asks for
The post (written in Portuguese, by u/Ornery-Dark-5844 in r/PromptEngineering) isn’t a prompt for writing prompts. It’s a product spec. You hand it to a coding-capable LLM and it describes what to build, not how to code it. That’s a smart split. The model already knows how to program. What it needs is clear priorities.
The priority stack is the best part. User experience comes first, then workflow, then functionality, then agent behavior. Implementation details come last. If internal complexity ever conflicts with a clear interface, the spec says to keep the capability and simplify how it’s shown.
The main building blocks
- Project Launcher: the first screen answers “what can I do here and how do I start?” instead of dumping a dashboard on you.
- Isolated projects: each one holds its own objective, platform, prompts, versions, decisions, validations and memory. Nothing leaks between projects.
- A workspace with four parts: project context, an agent chat area, a result area, and contextual actions like Optimize, Validate, Compare and Restore.
- Prompt Studio: a real editor with generate, optimize, validate, duplicate, compare, save and restore.
- Version comparison: previous vs current, what changed, possible regressions, and a short justification.
- Validation screen: shows status, risks, recommendations and confidence, so you can actually decide something.
The cognitive pipeline
The agent works through eight stages: Intent, Requirements, Domain, Architecture, Implementation, Verification, Operation and Evolution. Each stage has its own job and hands off to the next.
What makes this useful is that depth scales with the task:
- Quick: Intent, Requirements, Implementation, Verification. For small fixes.
- Standard: adds Domain and Architecture. For most prompt creation and optimization.
- Engineering: runs all eight, including Operation and Evolution. For production systems.
The spec insists these are real differences in agent behavior, not just a slider in the UI. It also says the target platform (ChatGPT, Claude, Gemini, Grok, APIs) should change how the prompt gets optimized, and that the agent must not invent capabilities a platform doesn’t have.
Memory without the hype
One line I like: learning is described as the project’s context and knowledge evolving, not the model being retrained. Interaction becomes observation, analysis, extraction, validation, consolidation, and finally project knowledge. That keeps expectations honest.
The look and feel
Dark mode by default, cold palette (graphite, navy, steel blue, soft cyan), and warm colors only when they mean something. Red for errors, orange for risk, yellow for pending. No decorative neon, no badge soup. Every screen gets one clear primary action, and empty states tell you what to do next.
Use cases
- Solo builders who want a personal prompt library with real version history.
- Teams that need to compare prompt versions and keep a decision log.
- Anyone prototyping an internal tool and wanting a solid spec to hand to a coding agent.
- Prompt writers who are tired of losing good iterations in scattered chat tabs.
What to watch for
This is only part 1 of 2, so the second half presumably covers the rest of the build. The post had a single upvote and no comments when I looked, so nobody has stress-tested it yet. Also, 49 sections is a lot of instruction. Expect the model to drift or skim in places, so build in stages and check each one against the completeness checklist in sections 47 and 48.
Prompt of the Day 🧭
You don’t need all 49 sections to borrow the pattern. Here’s a compact version:
“Build a [type of app]. Prioritize in this order: user experience, user workflow, functionality, agent behavior, internal state. Hide internal complexity from the main interface but keep the capability available in an advanced area. Use isolated projects, each with its own context, versions and memory. Give every screen one clear primary action and give every empty state a next step. The app is complete only when a user can go from creating a project to saving a new version entirely through the interface.”
That last line is the trick. A completion criterion written as a user journey is much harder for a model to fake than a feature list.
Call to action
Try the compact version on your next app idea, then tell me where the model cut corners. And if you want the full text, find part 1 on r/PromptEngineering and wait for part 2 before you start building.
Parte 1/2: PROMPT ENGINEERING AGENT WORKBENCH
by u/Ornery-Dark-5844 in PromptEngineering