Quick Start: Instead of writing a MiniMax H3 prompt from scratch, feed a reference image to ChatGPT, extract its visual language, then rebuild that language into a fresh scene. Steal the mood, not the picture.
Staring at an empty prompt box is how most people burn twenty minutes before generating a single frame. This Redditor, u/Practical_Low29, does the opposite: find a photo first, describe it, then write. The post breaks down a workflow tested across several MiniMax H3 experiments on Atlas Cloud, and it flips the usual order of operations.
The old way vs. the new way
The old way starts with a blank text field and a vague idea: “cinematic sunset scene, moody lighting.” You type, generate, get something close but not quite right, then spend ten rounds tweaking adjectives. The problem isn’t your vocabulary. It’s that you’re inventing atmosphere from nothing.
The new way starts with an image that already nails the mood you want: a YouTube thumbnail, a movie still, an old family photo, a Pinterest find. You’re not copying the subject. You’re extracting the structure underneath it: composition, lighting, pacing, emotional temperature. Then you rebuild a new scene using that structure as scaffolding.
That’s the whole shift. One approach invents mood from zero. The other reverse-engineers mood from something that already works, then swaps in your own characters and story.
How to actually do it
- Find a reference image. Don’t hunt for the exact scene you want. Hunt for the feeling: is it intimate or wide? Warm or cold? Documentary or dreamy? An image of two people by a quiet road at sunset might not give you the road you need, but it hands you a wide composition, small human figures, and a melancholic warm backlight you absolutely can reuse.
- Upload it to ChatGPT and ask for a scene description, not an image analysis. The author’s exact prompt:
Describe the scene in this image in English, focusing primarily on what is happening, the characters, their actions and body language, the setting and the overall atmosphere. Also briefly describe the composition, framing, camera angle, approximate lens choice, lighting, color palette and cinematic aesthetic. Keep it concise and scene-focused rather than overly technical.
The key word is “scene-focused.” You don’t need a catalog of every object in the frame. You need subject, action, environment, composition, camera, lighting, atmosphere: the exact ingredients MiniMax H3 responds to.
- Turn the description into your first prompt, then edit it once. Swap the character, the clothing, the location, the weather, the time of day. This is the actual trick 🎯: you’ve separated visual language from visual content, so you can keep the mood and throw out everything else.
- Add motion, because a still description is not a video prompt. The author layers three kinds of movement:
- Character movement: “slowly turning her head toward the approaching bus”
- Environmental movement: “tall grass sways gently, distant tree branches move subtly”
- Camera movement: “the camera slowly pushes forward with subtle handheld movement”
Stack all three and the model has real information to work with instead of a static caption with a video label slapped on it.
- Keep the important stuff first. The author’s priority order is worth memorizing: main subject, main action, environment, composition, camera movement, lighting, atmosphere, texture. Bury the action under three paragraphs of styling notes and the model quietly deprioritizes it. Longer isn’t better here. Clearer is better.
Why this matters beyond MiniMax H3
I’ve tested plenty of “just describe what you want” advice that falls apart the second you need something specific. This workflow works because it removes the hardest part of prompting: inventing atmosphere out of thin air. You’re not asking yourself “what should this feel like.” You already found the feeling. You’re just translating it.
It also explains why so many AI videos look flat even with detailed prompts. The detail is there, but the structure (where subjects sit in frame, how much environment shows, whether the light is soft or harsh) is missing, and that structure is exactly what a reference image hands you for free.
Try it this week
Grab one image that’s been sitting in your camera roll or a screenshot folder for no good reason. Run it through the five steps above, generate one MiniMax H3 clip, then swap the character and generate a second one from the same bones. Compare them. You’ll see the same mood wearing two different outfits, and that consistency is the whole point.
The full breakdown, including more example prompts, is worth a read if you want to steal the workflow wholesale. Check out the original discussion for the complete rundown!
How to Create Better MiniMax H3 Text-to-Video Prompts from Reference Images
by u/Practical_Low29 in PromptEngineering