Homa Transport Protocol
Last updated 2026-09-19Key points
What it is
- Homa is a new network protocol (rules for sending data between computers) designed for AI clusters (groups of machines working together on AI tasks).
- Unlike TCP (the internet's standard transport protocol), Homa is message-based, not stream-based, allowing short messages to bypass long ones.
- It handles both large and small messages well, keeping short messages at very low latency (delay).
- Homa measures under 100 microseconds for tail latency (worst-case delays), roughly 13 times faster than TCP.
How to use it
- Understand that Homa replaces the stream-based model of TCP with messages, where the fundamental unit is a remote procedure call (RPC, a request-and-response exchange).
- Use the Linux kernel module (loadable code that extends the operating system) available on GitHub for experimentation.
- Contact John Ousterhout for help if needed.
Watch out for
- Don't assume legacy protocols like TCP and RDMA (a method for direct memory access over a network) still fit; they were built for older environments.
- Don't assume big transfers are all that matter; AI workloads increasingly involve many small transfers where latency is crucial.
Tools named
- Homa (a network protocol for AI clusters), Linux kernel module (loadable code that extends the operating system).
Lesson 1: What is Homa Transport Protocol and why it matters
Homa is a network protocol (rules for sending data between computers) developed at Stanford by John Ousterhout's team as a clean-slate redesign for data centers. Unlike TCP (the internet's standard transport protocol), Homa is message-based, not stream-based. Its fundamental unit is a remote procedure call (a request message plus a response), and every message is independent rather than serialized into a stream. Short messages can bypass long ones instead of getting stuck behind them.
Two design choices matter. First, Homa knows message lengths, buried in the transport layer itself, so a receiver knows exactly how much data is coming once it sees the first packet. Second, congestion control (slowing senders when the network is busy) is handled by the receiver rather than the sender, because congestion happens at the last link. This lets Homa respond faster and more precisely.
Why does this matter for AI development? Historically, AI workloads were enormous transfers of data like weight gradients, where throughput mattered most. That is changing. The role of short messages in AI is increasing, and Homa handles the mix of large and small messages well, keeping short messages at very low latency. On tail latency (worst-case delays), Homa measured under 100 microseconds versus TCP's more than one millisecond, roughly 13 times faster, and even long messages were almost twice as fast. Ousterhout has released a Linux kernel module (loadable code that extends the operating system) on GitHub for experimentation.
Lesson 2: How to use Homa Transport Protocol: step-by-step
To use Homa (a transport protocol designed for data centers), start by understanding that it replaces the stream-based model of TCP (Transmission Control Protocol, the standard internet transport) with messages. Homa is message-based, not stream-based. The fundamental unit is a remote procedure call (RPC, a request-and-response exchange), which consists of a request message from a client and a response message from a server. Homa knows message lengths because they are carried in the transport. This lets a receiver predict the future: as soon as it gets the first packet of a message, it knows exactly how much more is coming. That gives the receiver complete information about congestion, so Homa controls congestion from the receiver rather than the sender. Step by step, a sender with a message breaks it into packets but transmits only the first few. The receiver responds with grants (permissions to send more) and uses them to favor short messages, implementing SRPT (shortest-remaining-processing-time scheduling) by preferring short messages. Homa also uses priority queues in modern switches (hardware ports with multiple egress queues), dynamically choosing queues to give shorter messages priority. Short messages bypass long ones instead of being serialized. On benchmarks, Homa’s P99 tail latency (99th percentile, the slowest 1%) for short messages is less than 100 microseconds versus more than a millisecond for TCP, roughly 13 times faster. Even long messages are almost twice as fast as TCP. Try the Linux kernel module, available on GitHub. Contact John Ousterhout for help.
Lesson 3: Best practices and pitfalls
Homa is a new transport protocol (rules for sending data across a network) designed from scratch for AI clusters — groups of machines working together on AI tasks. Its biggest pitfall is assuming legacy protocols like TCP and RDMA still fit. They don't: they were built for older environments and suffer very high tail latency (worst-case delay) when small messages mix with large ones. With TCP, tail latency exceeds a millisecond; Homa cuts it to under 100 microseconds, roughly 13 times faster.
Another mistake is assuming big transfers are all that matter. AI workloads increasingly involve many small transfers where latency is crucial, and short messages must not queue behind long ones.
Homa's best practices follow from three design choices. First, it is message-based, not stream-based: every message is independent, and short messages can bypass long ones. Second, congestion control (deciding when to slow sending) runs at the receiver, not the sender, because congestion happens at the last downlink and the receiver has better information. Third, Homa uses switch priority queues (separate output lanes for packets) to favor short messages, implementing SRPT (shortest-remaining-processing-time scheduling). The sender transmits only the first few packets until the receiver grants permission.
If you run AI clusters, try Homa — you can likely reduce tail latency by an order of magnitude or more.