Faceless explainer channels usually work backward. They write the script first, then scramble to find or generate a visual for every sentence. The original poster behind a new r/PromptEngineering thread, Ok_Astronomer_526, flips that order and sorts shots into five buckets before writing a single line of narration.
Quick start: before you script anything, label every shot idea as place, hands, movement, object detail, or transition. Build any shot with readable text (interfaces, screens, numbers) inside an editor instead of an AI prompt. After the rough cut, cut any line that needs a unique visual but adds little to the explanation.
🗂️ The Five Buckets
The creator’s template isn’t a report on a finished channel or a measured result. It’s a planning pass you run before writing narration, meant to separate shots that carry information from shots that only set a mood. The value isn’t in any single prompt. It’s in the sorting pass that happens before you write narration. That pass keeps a mood shot from carrying the same weight as a shot that has to convey specific information.
Here’s how the five buckets break down for a hypothetical episode about loading screens:
- Place: a wide, establishing shot. A phone sitting on a desk.
- Hands: a human action that grounds the scene. Someone setting the phone down.
- Movement: motion that carries the eye from one beat to the next.
- Object detail: anything with readable text. This one gets built in an editor, not generated, so the timing and wording stay under your control.
- Transitions: abstract AI-generated bridges between sections, built to support narration rather than compete with it.
One tip worth stealing: the object-detail bucket is the one creators most often outsource to AI by mistake. If a shot needs to display real text, numbers, or timing, build it in your editor so you control exactly what appears on screen.
Sorting shots this way before scripting gives every visual a job. Nothing gets added just because it looks cool.
Old Way vs New Way
The old way treats every shot as equally important. You write the narration, then chase down or generate a visual for each sentence. Mood shots end up competing with information shots for the viewer’s attention.
The new way separates the two jobs early. Information-carrying shots (place, hands, object detail) get built with control, often inside an editor so text and timing stay exact. Mood shots (movement, transitions) get generated loosely, since they only need to support the pacing, not explain anything.
That separation also hands you a built-in test. After you cut a rough version of the episode, remove any sentence that needs its own unique visual but barely helps the explanation. If the video still makes sense, the sentence (and its shot) wasn’t pulling weight.
Practical Steps
- List every beat in your explainer before writing narration, and tag each one with a bucket: place, hands, movement, object detail, or transition.
- Build any object-detail shot with readable text (interfaces, numbers, UI) inside a video editor. Don’t ask an AI model to render text you need to trust.
- Write a tight, restrictive prompt for transition shots, since these are the ones you’ll generate loosely. The original poster’s transition prompt for a loading-screen episode looked like this:
An abstract field of softly moving light between two dark sections, one slow camera push, no phone, no interface, no letters, no numbers, no logos, no cuts. Keep the motion restrained enough to sit behind narration.
- Notice what that prompt excludes as much as what it includes. Every “no” in that list (no phone, no interface, no letters, no numbers, no logos, no cuts) is doing a job. Each one stops the AI model from generating a shot that competes with the narration.
- Assemble a rough cut, then apply the removal test. Cut any sentence that needs a unique visual but adds little value, and see if the explanation still holds.
A quick warning: don’t skip the removal test just because the rough cut looks finished. It’s the step that catches shots you kept out of habit instead of need, and it’s the fastest way to trim runtime without losing clarity.
Try It on Your Next Script
This bucket system works because it separates two different jobs: explaining and setting a mood. Once you stop asking one shot to do both, your prompts get simpler and your edits get faster!
Before your next explainer, sort your shot list into these five buckets and watch how much easier the AI prompts become to write. It also works outside video: the same five buckets can sort visuals for a slide deck or a long-form article.
And if your transition prompt keeps sneaking in a stray letter or logo, the original thread is worth a look. Other creators are already comparing notes on the same problem.
A five-bucket prompt plan for faceless explainer visuals
by u/Ok_Astronomer_526 in PromptEngineering