SGLang’s RadixAttention slashed warm-up time-to-first-token to 68 ms on a dual-RTX 3090 rig, while vLLM’s PagedAttention lingered at 184 ms—an 82.9 % latency gap that makes multi-turn agent interactions feel noticeably smoother.
Both projects try to speed up large language model (LLM) inference, but the benchmark targets workloads that power autonomous assistants: long system prompts, tool-calling schemas and a rolling history of user-assistant turns. Those repeated tokens sit in GPU memory; if the cache isn’t managed efficiently, they become a compute bottleneck that stalls the conversation.
The workload that tipped the scales
Traditional LLM servers optimise for a single exchange—user prompt, model reply, done. Modern agents, however, keep a “conversation state” that can span dozens of turns, each adding hundreds of tokens to the cache. The test suite replayed such a multi-turn loop on two RTX 3090 cards, measuring:
- Warm time-to-first-token (TTFT) – delay before the first token appears after a new turn starts.
- Overall latency – average time per token across the whole loop.
- Sustained throughput – tokens processed per second when the loop runs continuously.
- Cache-hit ratio – proportion of tokens reused from the key-value (KV) cache instead of being recomputed.
SGLang beat vLLM on every metric: 68 ms vs 184 ms TTFT, an 82.9 % latency cut, 39.8 % higher throughput and a cache-hit ratio of 96.8 % compared with vLLM’s 84.2 %.
PagedAttention vs RadixAttention
Both engines store intermediate KV pairs in GPU memory, but they organise that memory differently.
- PagedAttention (vLLM) chops the cache into fixed-size blocks. If a prompt ends mid-block, the trailing tokens must be recomputed each turn because the block can’t be partially reused. The approach is simple and works when prompts line up with block boundaries, but it wastes cycles on the “edge” tokens that change most often in agent loops.
- RadixAttention (SGLang) treats the cache as a tree. It finds the longest common prefix between the current prompt and what’s already cached, regardless of block limits, and prunes unused leaves. This lets a massive system prompt stay pinned in GPU memory while only the newest turn is processed, boosting the cache-hit ratio and cutting redundant work.
When each engine shines
| Scenario | Preferred engine |
|---|---|
| Multi-turn autonomous agents, heavy tool calling, tree-of-thought reasoning | SGLang (RadixAttention) |
| Broad hardware support (including AMD and Gaudi accelerators), speculative decoding, vision-language models | vLLM (PagedAttention) |
The distinction isn’t only performance; it’s also about ecosystem. vLLM’s broader hardware compatibility makes it a safer default for organisations with heterogeneous compute clusters. Its speculative decoding feature—generating multiple token candidates in parallel—can accelerate single-turn generation, a niche where RadixAttention’s tree-based caching offers less advantage.
Bottom line
For developers building multi-turn agents that juggle long prompts and incremental updates, SGLang’s RadixAttention delivers a markedly faster, more cache-efficient experience on consumer GPUs. Teams that need wide hardware coverage or specialise in single-turn generation may still lean on vLLM. The choice now hinges on whether the workload is “agent-heavy” or “hardware-diverse.”
