Most People Prompt AI Scenes With Adjectives. This 3-Step Workflow Uses Motivation Instead

Most people write an AI character scene like this: “A shy woman smiles at a man, romantic mood, cinematic.” Then they wonder why the result feels like a mannequin doing an impression of a person.

A Redditor on r/PromptEngineering (u/Practical_Low29) shared a different approach for 30-second couple scenes in Seedance 2.5. It skips the mood words and treats the prompt like a tiny script with a director’s notes. It’s faster to debug, too, because you can see exactly which beat broke.

The key idea

Robotic characters usually come from reactions with no reason behind them. If a smile just appears, it looks pasted on. If the smile follows a glance, a pause and a half-second of hesitation, it looks like a person.

So the workflow is: Conflict, emotional progression, performance details, timing, duration check, final prompt. You build the scene from the cause of each reaction, not from the label of the emotion.

Old way vs. new way

The old way: Describe the mood, list the emotions you want, stack a bunch of simultaneous instructions (“blushing, smiling, looking away, leaning closer, laughing softly”), and hope the model sorts it out. The usual result is every reaction firing at once.

The new way: Give each reaction a trigger, write out how the expression unfolds over time, and check that the whole thing actually fits the clip length.

The difference is clarity. The model gets a sequence, not a pile.

The 3 steps

1. Give each reaction a reason

Sketch the emotional progression first. The example from the post:

Trust, then searching, then doubt, then asking for help, then realizing the trick.

Then connect each change to something visible: a glance, a line, a pause or an action.

Physical closeness needs a reason too. Don’t just write “they lean in.” Write that she reaches for his phone, catches his wrist, then looks up and notices how close they’ve become. The moment grows out of what they’re doing.

2. Describe how the expression unfolds

Don’t name the emotion. Walk through it. For a shy reaction, the author writes something like:

“She pauses. Her eyes shift away; her head follows a moment later. She presses her lips together, then one corner lifts before the smile fully appears.”

Notice the order. Eyes first, head second, lips third, smile last. That lag between body parts is what makes it feel human.

Small contradictions help as well. She says “I’m not jealous,” but hesitates and looks away before answering. What she says and what she does don’t match, and that gap is what reads as real. One commenter put it well: that contradiction is the secret sauce of any performance, whether you’re directing a person or a prompt.

One rule to keep this clean: one main action plus one subtle emotional cue per beat. More than that and everything happens at once.

3. Check what fits into 30 seconds

This is the step most people skip, and it’s where scenes fall apart.

Twelve lines at two seconds each is 24 seconds of speech. That leaves about 6 seconds for everything else. Movement can overlap with dialogue, but silent reactions and pauses still need their own room.

The fix is simple:

  • 🎤 Read the lines aloud and time them
  • ⏱️ Time the actions separately
  • ✂️ Trim anything that rushes the scene

If you have to speed up a gesture to make it fit, cut a line instead.

A quick template you can copy

For each beat in your scene, fill in:

  • Trigger: what just happened (a line, a touch, a silence)
  • Main action: the one thing the character does
  • Subtle cue: one small tell (a look away, a pressed lip)
  • Time: how many seconds it gets

Add up the times before you write the final prompt. If the total is over your clip length, trim first and generate second.

Why this works beyond video

The same logic applies to any generative tool where you control a performance: video, voice, even written dialogue. Give every change a cause, describe it as a sequence, and respect the clock. It’s basically what a director does with an actor, just compressed into text.

Try it this week

Pick one short scene you’ve already generated and found stiff. Rewrite it using the three steps, keep it to one action and one cue per beat, and time it out loud before you hit generate. Compare the two versions side by side and see which one feels more alive.

If you try it, tell us which beat changed the most. I’m curious whether the pause or the contradiction does more of the work for you.

Frequently Asked Questions

Q: Should I write ‘pause’ explicitly, or just describe what the character does during the silence?

Some creators swear by explicitly calling out the silence: ‘She pauses; her eyes shift,’ because it forces Seedance to respect the timing. Others find they get better results just describing the physical action, like a hesitation before speaking. Since both work, try each on the same scene and see what feels more natural for your character.

Q: How much detail is too much detail in a single beat?

If you’re writing more than one main action plus one subtle cue per beat, you’re probably piling on too much. It’s way easier to throw five emotional instructions at the model than to trust a single well-timed glance to do the work. When in doubt, cut it in half and see if the moment still lands.

Q: How do I make subtle contradictions between words and body language work?

Have the physical action come first: the hesitation, the gaze shift, the flinch. Then comes the dialogue. So when your character says ‘I’m not jealous,’ the audience already saw them look away or tense up. That gap between action and dialogue is what makes reactions feel real.

Making AI characters feel less robotic: my 3-step workflow for 30-second scenes
by u/Practical_Low29 in PromptEngineering

Scroll to Top