I bought the 9800X3D for games, and the 96 MB of L3 cache is the whole reason it costs more than the normal 9700X. It works. Some games that used to stutter in busy scenes just stopped doing it. So when I started running models on the CPU, I wondered if the big cache would help there too. 96 MB is a lot of cache, and LLM inference is mostly about moving numbers from memory to the cores as fast as possible.
Before testing I expected not much, and that turned out right. What limited my CPU was the link between the cores and the RAM, which I hadn’t thought about at all.
Why the cache should not matter#
When a model writes a token, it uses every weight once. Llama 3.1 8B at Q4_K_M is about 4.9 GB, so each token means reading 4.9 GB from memory. A cache helps when you read the same data again before it gets pushed out. 4.9 GB pushes out 96 MB about fifty times per token, so by the time the next token starts, the first layers are long gone from L3. Every token comes from RAM.
So for anything except a tiny model, the speed ceiling should be simple:
max tokens/sec = memory bandwidth / model size
My RAM is two sticks of DDR5-6200. Two channels, 8 bytes each, 6,200 million transfers per second:
6200 MT/s * 8 bytes * 2 channels = 99.2 GB/s
99.2 GB/s divided by 4.9 GB gives about 20 tokens per second on the CPU. I didn’t get anywhere near that.
The number I actually got#
All of this is a CPU-only build of llama.cpp on Windows, with the 5070 Ti not
involved at all, and llama-bench doing five repetitions per setting:
llama-bench -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf \
-t 1,2,4,6,8,16 -p 512 -n 128 -r 5
| threads | prompt tok/s (pp512) | generate tok/s (tg128) |
|---|---|---|
| 1 | 22 | 3.1 |
| 2 | 43 | 5.8 |
| 4 | 82 | 9.6 |
| 6 | 118 | 10.5 |
| 8 | 152 | 10.6 |
| 16 | 158 | 9.9 |
Generation stops improving at around 6 threads and peaks at 10.6 tokens per second. That’s 10.6 times 4.9 GB, so about 52 GB/s of weights read per second. Barely half of the 99.2 GB/s from the math above. I assumed llama.cpp was wasting something, so I ran AIDA64’s memory benchmark, and it said reads top out at about 64 GB/s. Writes were even lower.
This is how Ryzen is built, and I didn’t know it when I bought it. The cores sit on a chiplet (the CCD), and the memory controller sits on a separate I/O die. They talk over Infinity Fabric. On Zen 5, one CCD can read about 32 bytes per fabric clock, and with the fabric at around 2,000 MHz:
32 bytes * 2000 MHz = 64 GB/s
The 9800X3D has only one CCD, so 64 GB/s is the limit no matter how fast the RAM is. The 9950X has two CCDs and gets two of these links, which is part of why it reads faster with the same RAM. Faster RAM still helps a little with latency, but for big sequential reads, the fabric sets the limit. llama.cpp getting 52 out of 64 GB/s isn’t bad at all.
The thread table makes sense once you see it that way. Generating is waiting on memory, and six cores are enough to keep the link full. More threads just fight over the same 64 GB/s, and with SMT (16 threads) they start getting in each other’s way, so it gets slower. Prompt processing is different. It does a lot of math for each weight it reads, so it keeps scaling up to all 8 cores. Prompt and generation have different limits, so the best thread count is different for each.
Where the cache does show up#
To actually test the cache, I needed models small enough to fit in it. SmolLM2 135M is small enough to fit at Q4. Then I used a few larger sizes and calculated effective bandwidth as model size times generation speed, all at 8 threads:
| model | file size | tok/s | effective GB/s |
|---|---|---|---|
| SmolLM2 135M Q4_0 | 88 MB | 1,050 | 92 |
| SmolLM2 135M Q8_0 | 145 MB | 410 | 59 |
| SmolLM2 135M F16 | 270 MB | 225 | 61 |
| Qwen3 0.6B Q8_0 | 640 MB | 92 | 59 |
| Llama 3.1 8B Q4_K_M | 4.9 GB | 10.6 | 52 |
| Llama 3.1 8B Q8_0 | 8.5 GB | 6.4 | 54 |
The only model that fits in L3 goes above the 64 GB/s limit, to about 92 GB/s. As soon as the file is bigger than 96 MB, it drops back to the same high 50s as everything else. So the cache does work, but only for a 135M model that I would never use for anything real.
I expected the jump to be bigger. L3 on this chip can deliver hundreds of GB/s, so 92 is disappointing. A model this small spends most of each token on overhead. Lots of tiny operations, threads waiting for each other at every layer, and the KV cache and activations taking cache space too. At 1,050 tokens per second, each token takes under a millisecond, so even small costs add up.
AVX-512 is the more useful part of the chip#
Zen 5 desktop chips run AVX-512 at full width, and llama.cpp has kernels for it. I built a second copy with it turned off to see how much it does:
cmake -B build-avx2 -DGGML_NATIVE=OFF -DGGML_AVX2=ON -DGGML_AVX512=OFF \
-DGGML_FMA=ON -DGGML_F16C=ON
| build | pp512 tok/s | tg128 tok/s |
|---|---|---|
| AVX2 | 104 | 10.4 |
| AVX-512 | 152 | 10.6 |
Prompt processing gets about 45% faster, because it was limited by math and the wider instructions do more math per cycle. Generation doesn’t change at all, because the cores were already waiting on the fabric. The thread table showed the same thing.
So for running models, the X3D part of my CPU does basically nothing, and the number that matters most is one I never looked at when buying it: how many bytes per second one chiplet can pull through Infinity Fabric. If I bought a CPU mainly for inference, I’d look at memory channels and the CCD layout before the cache, or just put the money into VRAM. The 5070 Ti in the same case runs the 8B model about fifteen times faster than this.
I still like the chip. The games are smooth, and the CPU will matter again as soon as a model doesn’t fit on the GPU and part of it has to live in system RAM. When that happens, that 64 GB/s limit decides the speed.