Four Models, One Rocket, Your Pick

Choosing a model for heavy 3D coding work? Here’s the fast way to narrow it down. Let someone else run the bake-off first, then read the results before you burn your own tokens.

The original poster, active in r/PromptEngineering, ran exactly that test. Four models went in: Opus 5.5, Opus 5, GPT-6 Astra, and GPT-6 Sol. All four got one job: build a full Falcon 9 Starlink launch in Three.js, no build step, no shortcuts, one shot each.

What Actually Separates a Good Model Here

Before you copy anyone’s approach, know what the prompt actually demanded. Six things at once:

  • Procedural rocket geometry built from primitives (booster, legs, grid fins, fairing)
  • A launch pad environment with a strongback and a sky gradient
  • Physics based liftoff using real thrust, mass, and gravity math
  • A custom particle system for exhaust and pad smoke
  • A cinematic camera that shakes with thrust and dynamic pressure
  • Directional lighting with real shadows, plus a flickering point light on the engine

Any model can nail one or two of those. Very few hold all six together without the frame rate falling apart.

That is the real test here, not can it write Three.js code. It is whether a model can juggle procedural geometry, a physics loop, a particle system, and a camera rig at once. All in a single file, still running smooth. That is a fair proxy for how a model handles any complex, multi-system app, not just rockets.

The single-file, no-build-step constraint is doing quiet work too. It rules out the usual crutch of splitting logic across ten clean modules. Every model has to hold the whole state machine in its head at once: launch, hold, clamp release, ascent. That is exactly where weaker models start dropping details.

How They Stacked Up

The full run lives in the video, but a couple of results stood out enough that viewers called them out by name.

Model What Stood Out
Opus 5.5 🥇 Plume particles at stage separation, praised for smooth color interpolation from bright ignition white to expanding gray smoke
GPT-6 Astra 📉 Camera work came out janky during the shake sequence, according to viewers

Opus 5 and GPT-6 Sol’s runs are in the video too, but the reaction so far is thin on both. Worth watching before you commit to anything.

Going by what has actually been reported, Opus 5.5 is the safer pick for particle heavy work. Think plumes, smoke, fire, anything needing smooth color and opacity transitions over time. That is a narrow, specific strength, not a blanket best model claim. If your priority is camera work or physics accuracy instead, run your own comparison on those exact criteria before picking a winner.

Run the Same Test Yourself

Start with the prompt, word for word:

Write a complete, single-file HTML/JavaScript application that simulates a SpaceX Falcon 9 (Starlink mission) rocket launch in 3D using the Three.js library. The application must run locally in a web browser without requiring a build step by loading Three.js via CDN. Push your rendering and mathematical capabilities to the limit.

Then build in this order:

  1. Geometry first. Get the booster, interstage, second stage, and fairing rendering as static primitives before you animate anything. If the shape is wrong, nothing downstream matters.
  2. Launch pad and lighting second. Strongback, ground plane, sky gradient, directional light with shadows. This is your baseline scene, test it in isolation before adding motion.
  3. Physics loop third. Wire up the hold, the clamp release, and the thrust versus mass versus gravity calculation on its own, before touching particles or camera.
  4. Particle system fourth, and keep it decoupled. This is where Opus 5.5 reportedly pulled ahead. Give the model room to tune scaling, velocity, opacity, and color interpolation without fighting the rest of the scene.
  5. Camera last. Shake tied to thrust and dynamic pressure is the hardest thing to get right. It is also the piece most likely to break if you bolt it on too early.
  6. Benchmark frame rate across whatever models you test. A gorgeous plume that tanks your FPS is not a win.

One tip: if you are testing multiple models on a prompt this dense, run each one twice. A single generation can get lucky or unlucky on any one of the six requirements, and you want to know which.

Second tip: score each model against the six requirements separately instead of judging the whole scene at a glance. A launch that looks impressive on first watch can still be hiding a weak particle system or a camera that never reacts to thrust.

Go watch the full showcase, see all four launches side by side, and judge the camera work and lighting for yourself. That is where the real decision gets made, not in a comment section.

Same prompt, 4 frontier models (Opus 5.5, Opus 5, GPT-6 Astra, GPT-6 Sol): Coding a multi-stage Falcon 9 launch in Three.js in one try
by u/davchi1 in PromptEngineering

Scroll to Top