In a few posts now, I’ve said that writing tokens is mostly waiting on memory. If that’s true, the 5070 Ti should be wasting a lot of its 300 W while llama.cpp writes text, and a lower power limit should cost almost nothing. It would also make the case quieter for free. Then someone posted a careful RTX 5090 power test this weekend, and one of its numbers didn’t match what I expected. So I went through the public numbers to work out what a power limit should really do to a 5070 Ti.
what a power limit changes#
nvidia-smi -pl 240 tells the card to stay under 240 W. To do that, the card
lowers its core clock and voltage. As far as I know, the memory clock stays the
same. So work that is limited by math should get slower, and work that is
limited by memory speed shouldn’t care. In llama.cpp, reading the prompt is
mostly math, and writing new tokens is mostly memory. So prompt reading should
slow down and writing shouldn’t.
The llama.cpp CUDA scoreboard
already has a nice comparison for this. The 5080 and the 5070 Ti use the same
kind of memory on the same 256-bit bus. The 5080 has 20% more compute (84 SMs
against 70) and 7% more memory speed (960 GB/s against 896). On Llama 2 7B
Q4_0:
| card | prompt reading (pp512) | writing (tg128) |
|---|---|---|
| RTX 5070 Ti | 8,420 tok/s | 182.4 tok/s |
| RTX 5080 | 9,488 tok/s | 184.7 tok/s |
| difference | +12.7% | +1.2% |
Prompt reading went up a lot with the extra compute. Writing went up only 1.2%, even less than the extra memory speed would suggest. These numbers come from two different people with different setups, so the small details don’t mean much, but the big picture is clear. More compute did almost nothing for writing.
the test that made me look closer#
The 5090 test used one card and logged power every 250 ms, so I trust it more than most. Two results matter here. With SGLang, an MoE model and one user, the card only drew about 233 W. The limits went from 575 W down to 400 W, so the limit never kicked in and nothing changed. With llama.cpp, a dense 31B model and one user, the card drew 482 W. At a 400 W limit, writing dropped from 75.9 to 71.9 tokens per second, about 5%. Prompt reading dropped about 8%.
The 482 W surprised me. I thought a card waiting on memory would be half idle, and it’s the other way around. 75.9 tokens per second from a file of about 17.5 GB means reading around 1.3 TB/s, about 74% of what the 5090’s memory can do. Moving that much data uses a lot of power by itself. So the first thing to check is whether the card even draws more power than the limit.
two numbers that don’t fit#
Some people in the scoreboard comments also tested with lower limits, and their numbers look worse. One 4090 at a 300 W limit wrote 167.5 tokens per second, against 186.2 for the normal 4090 on the scoreboard. That’s 10% lower. One 5090 at 400 W wrote 236.7 against 290.0, which is 18% lower. Prompt reading dropped about 10% in both.
These are different machines and setups from the normal entries, so they’re weak proof, but I can’t ignore them. My best guess is this. At 180 to 290 tokens per second, each token takes only 3 to 6 ms, and not all of that time is spent reading weights. Some of it is small setup work between steps, and that part runs at the core clock. A lower clock makes that part slower, even when memory speed stays the same. The 5070 Ti uses about 78% of its memory speed on the scoreboard, which leaves about a fifth of each token for this kind of extra work.
working it out for llama 3.1 8b#
For Llama 3.1 8B Q4_K_M, the same 78% gives about 142 tokens per second at the
normal 300 W. The 5090 drew about 84% of its limit while writing, so I’m
assuming the 5070 Ti draws around 250 W while writing. Prompt reading should
use the full limit. Scaling the scoreboard numbers for the bigger model, I’d
expect prompt reading at around 6,500 tokens per second.
| limit | writing tok/s | prompt tok/s | energy per written token | energy per prompt token |
|---|---|---|---|---|
| 300 W (normal) | ~142 | ~6,500 | ~1.77 J | ~46 mJ |
| 270 W | ~142 | ~6,270 | ~1.77 J | ~43 mJ |
| 240 W | ~139 | ~6,040 | ~1.72 J | ~40 mJ |
| 210 W | ~133 | ~5,820 | ~1.57 J | ~36 mJ |
The writing column is the uncertain one. If it holds, 270 W is free, 240 W costs about 2%, and 210 W costs about 6% but saves about 11% of the energy per token. Prompt reading gets slower, but it saves even more energy per token, because it was already using the full limit. If the 4090 and 5090 comments are closer to the truth, writing at 210 W could be 10% slower, not 6%.
Not every card can go down to 210 W. nvidia-smi -q -d POWER shows the lowest
and highest limits for your card. On Windows, -pl needs an admin terminal,
and the limit resets when you restart.
If you want to check this yourself, the normal test is too short. At 142 tokens
per second, tg128 is less than a second of work. Logging power every 250 ms
gives only three or four readings, and most of them are taken while the card is
still speeding up. Longer runs work better:
nvidia-smi -pl 240
llama-bench -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -ngl 99 -fa 1 \
-p 4096 -n 1024 -r 5
And keep this running in a second terminal:
nvidia-smi --query-gpu=timestamp,power.draw,clocks.sm,clocks.mem,temperature.gpu \
--format=csv,noheader -lms 250 > power-240.csv
Energy per token is the average power while writing, divided by tokens per
second. The most important column is clocks.mem, because everything above
assumes the memory clock doesn’t change. If it drops at 210 W, the table is
wrong for a boring reason, and the setup-work idea never gets tested.
If writing really ignores the limit, the best setting is the lowest limit that costs nothing, because one person chatting with a model spends most of the time waiting for writing, not prompt reading. If writing drops like it did for that 4090, then those 7 ms per token have more clock-bound work in them than I thought, and the 64 users post needs a footnote.