Homa Wants to Retire TCP Inside AI Data Centers

A transport protocol called Homa is making a direct pitch: TCP has had its run inside AI clusters, and it’s time to replace it. According to Hacker News, a video presentation titled “Homa: The end of TCP for AI clusters” reached a score of 170 in the Research category. That’s a lot of attention for a networking talk.

The Hacker News listing gives only the title and a link to the video, so it doesn’t show the talk’s specific claims or numbers. The research behind Homa is public, though, and it explains why the idea keeps coming back.

🧠 What Homa Actually Is

Homa started at Stanford in John Ousterhout’s research group. It’s a transport protocol built specifically for data centers. TCP was built for the open internet, where links are slow and unreliable and connections are long-lived.

Inside a modern data center, those assumptions break. Machines sit microseconds apart and swap huge numbers of short messages, and every microsecond of waiting costs money. Ousterhout’s argument, set out in his paper “It’s Time to Replace TCP in the Datacenter,” is that TCP’s core design choices hurt this workload. Tuning won’t fix them.

Homa changes several things at once:

  • Messages instead of byte streams. Applications send discrete messages, which is how RPCs and AI workloads already think.
  • No connections. There’s no per-connection state to set up or maintain across thousands of peers.
  • Receiver-driven flow control. The receiver decides who gets to send next. That cuts the queue buildup that drives latency spikes.
  • Shortest-message-first priority. Small messages skip ahead of bulk transfers, so they don’t wait behind large ones.

📊 Why the Numbers Matter

In earlier published evaluations, Homa researchers reported big drops in tail latency compared with TCP and DCTCP, a data center variant of TCP. Tail latency means the slowest requests, the ones that hold up everything else. The gains were largest for short messages under load, where some measurements showed improvements of more than an order of magnitude.

That’s the metric AI clusters care about most. In distributed training, thousands of GPUs often have to finish a communication step before any of them can move on. One slow message stalls the whole group, and those GPUs cost a lot to leave idle. Fixing tail latency matters more than raising average throughput.

🔍 Why This Is Coming Up Now

What stands out here is the timing. The AI buildout has made networking a first-class bottleneck. Big operators already work around TCP with RDMA, a technique that lets machines read each other’s memory directly over fabrics like InfiniBand and RoCE. Industry groups such as the Ultra Ethernet Consortium are also designing new transports for AI traffic.

Homa fits into that same push. The real question isn’t whether TCP struggles at AI scale. Most people building these clusters already agree it does. The question is which replacement wins.

⚠️ Limitations Worth Keeping in Mind

Homa faces real obstacles, and its own backers have been open about them:

  • Ecosystem inertia. Decades of software, tools and operational know-how assume TCP. Switching means changing applications, not just flipping a kernel setting.
  • Competition. RDMA-based fabrics and Ultra Ethernet already have strong vendor support in AI deployments.
  • Inside the data center only. Homa isn’t trying to replace TCP on the public internet. Its case covers controlled data center environments.
  • Uncertain production evidence. Research benchmarks are encouraging, but large-scale production results on frontier AI clusters still need independent confirmation.

🛠️ What Practitioners Should Do

If you run or design AI infrastructure, here’s how to use this:

  1. Measure tail latency, not just bandwidth. Your p99 communication times may be quietly wasting GPU hours.
  2. Track the transport debate. Homa, RDMA and Ultra Ethernet all go after the same pain point, and the winner will shape hardware choices over the next few years.
  3. Experiment where it’s cheap. Homa has a Linux kernel implementation. Teams with RPC-heavy internal services can test it on non-critical workloads.

🔭 What Comes Next

TCP isn’t going away overnight, and it won’t disappear from the wider internet at all. Inside AI clusters, though, the case against it keeps getting stronger as clusters grow. My advice is to treat the transport layer as a real design choice, not a fixed default. The full talk and the Hacker News discussion around it are available at the original source.

Scroll to Top