Tech

Serving one model to 64 people at once

Throughput goes up with more users, each user gets slower, and eventually the KV cache runs out.

Every few days someone posts a screenshot of a model doing some huge number of tokens per second on a Mac or a single GPU, and the replies treat it like a spec sheet. I wanted to know what that number looks like when more than one person is using the model. So I put Llama 3.1 8B Instruct behind llama-server on my Mac Studio (M4 Max, 64 GB), and behind vLLM on a rented L40S on Fly.io, and hit both with 1, 4, 16, and then 64 requests at the same time.

The short version is that the total tokens per second went up a lot, each person’s tokens per second went down, and the time before the first token appeared got much worse. On the GPU, what eventually stopped me was memory for the KV cache, while compute still had room to spare.

The setup#

On the Mac I used the Q4_K_M GGUF, which is about 4.9 GB, with all layers on the GPU and one slot per concurrent request:

llama-server -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
  -ngl 99 -fa on -np 64 -c 131072 --port 8080

-np is the number of parallel slots and -c is the total context for all of them, so 64 slots with -c 131072 gives each request about 2K tokens to work with. That was enough for this test and becomes important later.

On Fly.io I started a machine with --vm-gpu-kind l40s, which was about $1.25 an hour, and ran vLLM with the model in BF16:

vllm serve meta-llama/Llama-3.1-8B-Instruct \
  --max-model-len 32768 --gpu-memory-utilization 0.90

The whole GPU side cost me a little under four dollars, most of it spent waiting for the model to download and me rereading the vLLM docs.

The client was a small Python script with asyncio that kept N streaming requests open at all times against the OpenAI-compatible endpoint. Every request had a prompt of about 1,000 tokens and asked for 256 tokens back. For each request it saved the time to the first streamed token and the token rate after that. Each concurrency level ran for five minutes after a short warmup, because I’ve already embarrassed myself with single runs on a warm laptop.

What happened#

The Mac, Q4_K_M, 1,000 tokens in and 256 out:

concurrenttotal tok/sper user tok/sTTFT p50TTFT p99
172721.1 s1.3 s
4176441.9 s3.8 s
16304196.4 s14.2 s
644166.529.8 s61.5 s

The L40S, BF16, same requests:

concurrenttotal tok/sper user tok/sTTFT p50TTFT p99
146460.07 s0.09 s
4172430.08 s0.15 s
16592370.16 s0.41 s
641,536240.45 s1.4 s

If I only showed the first row, the Mac wins. 72 tokens per second against 46 on a datacenter GPU makes a nice screenshot. It’s also not a fair comparison, because the Mac is reading a 4.9 GB quantized model and the L40S is reading 16 GB of BF16 weights. With one user, generating each token mostly means reading every weight from memory once, so a smaller file wins even with less memory bandwidth. Nothing in that screenshot would tell you about the quantization.

Batching works because the weights get read once per step no matter how many sequences are in the batch, so four users cost only a little more than one. The L40S kept per-user speed almost flat up to 16 users and still gave each of 64 users about 24 tokens per second, which is faster than I can read. Total output went from 46 to over 1,500.

The Mac’s total also went up, to 416, and if I had posted only that number it would look great. But each of those 64 people got 6.5 tokens per second and waited half a minute for the first word. The p99 was a full minute. The problem there is prompt processing. Reading the 1,000 token prompt is compute heavy, the M4 Max does it at roughly 900 to 1,000 tokens per second for this model, and 64 users means 64,000 prompt tokens that all need processing before those requests can start generating. New prompts also share the batch with users who are already generating, so everyone slows down whenever somebody new arrives.

Then the long prompts showed up#

The short-prompt test was easy for the GPU. It had plenty of memory left and was mostly limited by compute. So I changed the mix on the L40S: 48 users still sent 1,000 token prompts, and 16 users pasted in a document of about 16,000 tokens and asked questions about it. That’s a pretty normal usage pattern for anything with a “chat with your docs” feature.

trafficshort-user TTFT p50short-user TTFT p99preemptions
64 short0.45 s1.4 s0
48 short + 16 long2.1 s9.7 s37

Nothing crashed and the total throughput looked about the same, but the short users were having a much worse time. vllm:num_requests_waiting stayed above zero for most of the run, KV cache usage stayed close to 100%, and vllm:num_preemptions_total went from 0 to 37. A preemption means vLLM ran out of room for a running request, threw away its cache, and had to rebuild it later. The 16 long requests were taking up most of the memory, and the short requests waited behind them in a queue.

On the Mac it failed earlier, and at least it said so. Each slot had about 2K tokens, so every long request got rejected straight away:

the request exceeds the available context size, try increasing it

Raising -c would fix that, but llama-server gives every slot the same share, so fitting 16K for everyone at 64 slots would need about a million tokens of context.

How big the KV cache actually is#

To generate the next token, the model looks back at every previous token in the request. Recomputing all of that each step would be very slow, so the server keeps the key and value vectors for each token in each layer. That stored data is the KV cache, and it lives for the length of the request.

The size comes straight from the model’s config.json:

{
  "hidden_size": 4096,
  "num_attention_heads": 32,
  "num_hidden_layers": 32,
  "num_key_value_heads": 8,
  "torch_dtype": "bfloat16"
}

The head size is 4096 / 32 = 128. For every token, every layer stores a key and a value for each of the 8 KV heads, in 2 byte BF16:

2 (K and V) * 32 layers * 8 KV heads * 128 dims * 2 bytes
  = 131,072 bytes
  = 128 KiB per token

That adds up quickly. A request with 1,256 tokens holds about 157 MiB. A 16K token request holds 2 GiB. A single request using the full 128K context window would hold 16 GiB, which is the same size as the model weights. The 8 KV heads are already a saving, because Llama 3.1 uses grouped-query attention. With one KV head per attention head it would be 512 KiB per token and all of these numbers would be four times bigger.

The L40S has 48 GB. With --gpu-memory-utilization 0.90, vLLM gets about 43 GB, the weights take 16 GB, a few more GB go to activations and CUDA graphs, and the rest becomes KV cache. The startup log says so:

GPU KV cache size: 203,456 tokens
Maximum concurrency for 32,768 tokens per request: 6.21x

That second line is easy to skip, but it explains the long-prompt run. 203,456 tokens is a lot of short chats: 64 users with 1,256 tokens each need about 80,000. Sixteen users with 16K documents need over 260,000 by themselves, which is more than the whole cache. At that point memory is the limit, and more compute wouldn’t help. The long requests slow everyone else down just by holding memory for a long time.

The Mac has the opposite problem with this model. At Q4 the weights are small, and 64 GB of unified memory holds the whole -c 131072 cache (16 GiB) without trouble. It runs out of prompt processing compute long before it runs out of memory. A bigger model or longer contexts would change that.

There are ways to make the cache smaller. vLLM can store it in FP8 with --kv-cache-dtype fp8, and llama.cpp can quantize it with --cache-type-k q8_0 --cache-type-v q8_0. Both roughly halve the memory per token, which roughly doubles how many people fit. You can also cap --max-model-len at something your users actually need, or send the long-document traffic to a separate instance so it can’t take memory from the people asking short questions. I would try the separate instance first, because it’s the only option that doesn’t change what the model sees.

What the screenshot should say#

A tokens per second number for a model is missing the same things my laptop benchmark was missing. Was that one user or sixty-four? Total throughput or per user? How long was the prompt, how long was the output, and what precision were the weights and the cache in? On my Mac, the same model and the same server gave me 72, 416, and 6.5 tokens per second depending on which of those I picked, and all three numbers were real.

If I were sizing something for 50 people now, I would start from the KV cache math, figure out the longest request I’m willing to accept, and then look at TTFT p99 under that load. The total tok/s is still worth knowing, but it’s the number that tells me the least about whether people will enjoy using it.

TRIP COMPUTER / SESSION

TRIP A

Current drive.

A private counter for this browser session. Nothing here is transmitted or retained after the session ends.

Elapsed
00:00
Sections
0
Notes
0
Screens
0

Route/

Build plate74b649a
Chassis
v7.1.3
Revision
74b649a
Last serviced
06 Oct 2026

OWNER’S MANUAL / WTHRAJAT

OPERATING NOTES

How this thing moves.

The header behaves like a small mechanical system. Its readings respond to how you move through the site.

Throttle
Scrolling is input. Faster downward movement builds more momentum and engine speed.
Transmission
Upshifts follow sustained input. Scrolling upward slows and downshifts, and may briefly show reverse.
Idle
When input stops, RPM settles near 850 with mechanical drift. The gearbox eventually returns to neutral.
Tachometer
The needle always follows the reported RPM. It is never calculated from your position on the page.
Trip A
Session time, explored sections, opened notes and approximate screens travelled stay in this tab session.

Controls

Ctrl K
Search notes
?
Open this manual
Esc
Close an instrument
Tab
Move through controls