The Bet That Your GPUs Are Half Empty

A French startup thinks the AI industry is leaving speed on the table. Kog, according to TechCrunch AI, is going low, down to assembly and binary code, to pull far more inference out of the datacenter GPUs enterprises already own. The pitch is blunt: your Nvidia H200s and AMD MI300Xs aren’t running anywhere near their limit, and the fix is software, not new silicon.

That’s a pointed counter-narrative. Markets just handed Cerebras a warm IPO welcome in May for its purpose-built inference chips. Kog is arguing the opposite. CEO Gaël Delalleau told TechCrunch that “GPUs have a bright future,” and that the idea they’re bad at decoding has become a misconception as newer cards pack more memory bandwidth that “only begs to be unlocked.”

What’s actually being claimed

Kog hit the front page of Hacker News in May with a tech preview showing 3,000 tokens per second on a single request. Impressive number, but read the fine print: that was Laneformer 2B, a purpose-built model with roughly 2 billion parameters, now open sourced. The headline promise is “30x faster LLM inference,” and the gap between a tiny model and a real LLM is exactly what Kog still has to close.

Delalleau says the same approach scales. Skeptics aren’t sold. The honest read is that Kog has a compelling demo and an unproven thesis, and it knows it.

Why speed is suddenly the whole game

Inference cost and latency are now the bottleneck, not model quality. TechCrunch AI notes Kog is targeting users burned by slow AI workflows. Anyone running Claude Code knows the feeling of waiting on a long job, and Anthropic already charges a premium for a Fast Mode. That’s the tell. When a vendor puts a price multiple on speed, speed has become a product.

Kog says it pulled in 200 business leads off the preview. Software engineering looks like the first real use case, with design partners building prompt-to-app and prompt-to-game tools where faster output means more revenue. One lesson from early demand shaped the roadmap: customers won’t fine-tune small models, so Kog pivoted to accelerating larger ones instead.

The moat, and its ceiling

Delalleau’s edge is his background. Solid-state physics at École Polytechnique, then offensive cybersecurity, a four-time DEFCON CTF finalist. He frames the work as reverse-engineering hardware “down to assembly language and binary code” to make a GPU do things it wasn’t designed to do.

That depth is also the constraint. Each new GPU takes weeks or months of hands-on research, and the team is 11 people. So Kog can only support a handful of chips for now. The plan is to fold its methodology into agent-based pipelines that scale across more hardware. Until that lands, this is artisanal work, not a platform.

Kog isn’t the only one chasing this. Fellow French startup ZML built hardware-agnostic software that bypasses Nvidia’s CUDA. Delalleau positions Kog closer to Stanford’s Hazy Research lab, deeper into the metal.

The European angle

There’s a sovereignty story here too. Kog is backed by Scaleway, Bpifrance, and French Tech 2030, and its seed was co-led by Varsity VC. As Europe pushes to build its own AI capability across models and chips, a startup squeezing more from existing hardware fits the political moment. That could mean friendlier capital and customers at home.

What to watch, and what to do

The date that matters is September. Delalleau expects a first major model running at 10x speed, which he says unlocks customer traction and a Series A. Treat that as the proof point. If it slips or lands only on small models, the 30x claim stays a demo.

Practical takeaways:

  • If you run heavy inference, don’t assume a hardware refresh is the only lever. Software optimization on cards you already own is a live option worth pricing.
  • Benchmark real workloads, not vendor demos. A 2B model at 3,000 TPS says little about your 70B production model.
  • If you’re a buyer, ask which specific GPUs are supported today. With an 11-person team, the list is short.

Kog is a wager that the industry over-indexed on new chips and under-invested in wringing out the old ones. September will tell us whether that bet holds. TechCrunch AI has the full interview and demo details.

Scroll to Top