I still remember the first time I ran a free model on my own laptop and thought, “there’s no way this is real.” That feeling came flooding back watching this breakdown. Matthew Berman, the AI creator behind Forward Future, just walked through a brand new open-source model out of China, and I couldn’t look away.
The model is Qwen 3.8 Max from Alibaba. It’s open weights, free to download, and you can run it on your own servers right now. The creator points out it clocks in at a massive 2.4 trillion parameters, putting it right in the same league as Kimi K3. In plain terms: this is a frontier-grade model that regular people can actually get their hands on.
What’s new
According to the analyst, Qwen 3.8 Max goes toe to toe with the best closed models on the planet. A few numbers he highlighted:
- Terminal Bench (agentic coding): 86.6, edging out Fable at 84.6
- SWE Pro: 67.7, trailing Fable’s 80
- Multimodal reasoning, document intelligence, spatial understanding, visual perception: Qwen came out on top
He’s quick to add a smart caveat: treat every benchmark with a grain of salt. Scores can be gamed, and a model that crushes one test isn’t always great in the real world.
The twist
Here’s the part that made me sit up. The expert explains that Alibaba showcased Qwen reproducing research papers from scratch. No starter code, no ready-made pipeline, just the paper and a set of GPUs. Then it went further and invented and tested 18 of its own improvement ideas across four rounds. That’s recursive self-improvement, the step right before a model starts making its own discoveries. They even showed it running an autonomous silicon design flow, which matters because chip design is exactly where China has lagged.
The pricing catch
The original poster breaks down the cost, and this is where it gets spicy:
- Qwen 3.8 Max: $2 per million input, $6 per million output
- GPT 5.6 Soul: $5 input, $30 output
- Fable: $10 input, $50 output
But he makes a point I think more people need to hear. Price per token is only half the story. What actually matters is the total cost to finish a task. If a cheaper model burns three or four times the tokens, the savings vanish. Smart way to look at it.
Why I think this matters
The creator lays out both sides honestly. In the short term, these free Chinese models are fantastic for everyone. They give businesses real options, less platform risk, and freedom to self-host and fine-tune. For a tinkerer, downloading a frontier model and running it yourself is just plain fun.
But he doesn’t dodge the harder question. If US companies get 95% of the capability for a fraction of the price, many will drift toward these open-source models. Over time, models get co-designed with specific chips, and that could quietly build dependence on Chinese hardware. He frames it as geopolitical risk, not a personal grudge.
His closing thought is the one I keep chewing on: does cheap open-source pressure commoditize intelligence, or does it simply not matter because whoever wins recursive self-improvement first wins everything? He’s genuinely torn, and honestly, so am I.
Want the full picture, including the rebuilt benchmark charts he made with Codex? Watch the full video for the details and decide where you land.