I have a soft spot for stories about explorers, so this one grabbed me instantly. An AI creator just shared how they compressed the entire life of Marco Polo into a single 30-second video, and the way they built it is a proper step-by-step playbook. No editing tricks, no massive team. Just a smart workflow and the right models stacked in the right order.
The person who posted it built the whole thing with GPT-Image-2 for the visuals and Seedance 2.5 for the motion, running it all through Runway. What I love is that they didn’t just push a button and hope. They treated it like a real production pipeline. So let’s walk through exactly how the creator did it, step by step, with the reasoning behind each move.
The workflow, step by step
- Start with real historical research and reference material
- Build the full storyboard using GPT-Image-2 on Codex
- Hand the images to Seedance 2.5 for video generation
- Run the final generation on the Runway platform
Now here’s why each step matters, because the rationale is where the real lessons live.
Step 1: Ground everything in real history
Before touching any model, the author dug into actual historical sources. They pulled real paintings, including portraits of Kublai Khan, and even used their own photos for specific scenes like the Gobi Desert and certain artifacts.
Why it matters: AI models invent details when you leave gaps. By feeding in accurate references for characters, geography, and objects, the creator anchored the output to reality instead of letting the model guess. This is the difference between a video that feels authentic and one that feels like generic fantasy.
Step 2: Build the storyboard with GPT-Image-2
Next, the expert constructed the entire storyboard on Codex using GPT-Image-2. This produced the set of reference images that would define every scene in the final video.
The original poster calls this the most crucial step of the whole process.
That callout is worth sitting with. The video model can only work with what you give it. If your images are strong, consistent, and well composed, the animation stage has a solid foundation. If they’re weak, no amount of motion magic will save the result. Nail the images first, and everything downstream gets easier.
Step 3: Let Seedance 2.5 do the heavy lifting
With the storyboard locked, the creator handed everything to Seedance 2.5 for the actual video generation. The setup was surprisingly lean:
- One simple text prompt
- 30 reference images fed in as guidance
- A single 30-second output, no editing at all
The only thing added afterward was music. That’s it. According to the person who shared it, Seedance used nearly every image uploaded. It occasionally misses one or two, or shuffles the order, but the overall capability and prompt adherence impressed them a lot.
Why it matters: A lot of AI video still needs heavy stitching and manual cleanup. Getting a coherent 30-second sequence straight out of the model, with that many references honored, is a real signal of how far these tools have come. Fewer manual steps means faster iteration for anyone trying this at home.
Step 4: Run it all on Runway
For this project, the final video was generated on the Runway platform. The creator’s reasoning is practical: Runway gives you a full set of tools and models in one place, so you’re not bouncing between five different apps.
They also flagged something useful for the technical crowd. The Runway MCP is one of the most helpful tools in their Claude Code and Codex workflows. If you’re building automated pipelines, that integration lets you call these models directly from your coding environment.
What you can take from this
Even if you never make a Marco Polo film, the structure here is gold for any AI video project. Here’s the pattern to steal:
- Research first: Gather real references before generating anything
- Images before motion: Perfect your storyboard, then animate it
- Keep prompts simple: Let strong reference images carry the weight
- Consolidate your tools: Work on one platform to cut friction
I think this is a big deal because it shows the bottleneck has quietly moved. The magic isn’t a secret prompt anymore. It’s the preparation. The creators who win are the ones who do the homework, build clean references, and understand that each model has a specific job in the chain.
It also hints at something bigger for anyone into history, education, or storytelling. Ancient scenes that only existed in paintings and imagination can now move, breathe, and play out in seconds. That opens doors for teachers, museums, content makers, and curious minds everywhere.
The creator shared the full breakdown and the finished video on their LinkedIn post. If you want to see the actual 30 seconds of Marco Polo’s life brought to motion, go check out the original post for all the campaign details. It’s genuinely worth the watch.