Tech

What a power limit should do to a 5070 Ti running llama.cpp

Generation mostly waits on memory, so a lower power limit should barely slow it down.

In a few posts now, I’ve said that writing tokens is mostly waiting on memory. If that’s true, the 5070 Ti should be wasting a lot of its 300 W while llama.cpp writes text, and a lower power limit should cost almost nothing. It would also make the case quieter for free. Then someone posted a careful RTX 5090 power test this weekend, and one of its numbers didn’t match what I expected. So I went through the public numbers to work out what a power limit should really do to a 5070 Ti.

what a power limit changes#

nvidia-smi -pl 240 tells the card to stay under 240 W. To do that, the card lowers its core clock and voltage. As far as I know, the memory clock stays the same. So work that is limited by math should get slower, and work that is limited by memory speed shouldn’t care. In llama.cpp, reading the prompt is mostly math, and writing new tokens is mostly memory. So prompt reading should slow down and writing shouldn’t.

The llama.cpp CUDA scoreboard already has a nice comparison for this. The 5080 and the 5070 Ti use the same kind of memory on the same 256-bit bus. The 5080 has 20% more compute (84 SMs against 70) and 7% more memory speed (960 GB/s against 896). On Llama 2 7B Q4_0:

cardprompt reading (pp512)writing (tg128)
RTX 5070 Ti8,420 tok/s182.4 tok/s
RTX 50809,488 tok/s184.7 tok/s
difference+12.7%+1.2%

Prompt reading went up a lot with the extra compute. Writing went up only 1.2%, even less than the extra memory speed would suggest. These numbers come from two different people with different setups, so the small details don’t mean much, but the big picture is clear. More compute did almost nothing for writing.

the test that made me look closer#

The 5090 test used one card and logged power every 250 ms, so I trust it more than most. Two results matter here. With SGLang, an MoE model and one user, the card only drew about 233 W. The limits went from 575 W down to 400 W, so the limit never kicked in and nothing changed. With llama.cpp, a dense 31B model and one user, the card drew 482 W. At a 400 W limit, writing dropped from 75.9 to 71.9 tokens per second, about 5%. Prompt reading dropped about 8%.

The 482 W surprised me. I thought a card waiting on memory would be half idle, and it’s the other way around. 75.9 tokens per second from a file of about 17.5 GB means reading around 1.3 TB/s, about 74% of what the 5090’s memory can do. Moving that much data uses a lot of power by itself. So the first thing to check is whether the card even draws more power than the limit.

two numbers that don’t fit#

Some people in the scoreboard comments also tested with lower limits, and their numbers look worse. One 4090 at a 300 W limit wrote 167.5 tokens per second, against 186.2 for the normal 4090 on the scoreboard. That’s 10% lower. One 5090 at 400 W wrote 236.7 against 290.0, which is 18% lower. Prompt reading dropped about 10% in both.

These are different machines and setups from the normal entries, so they’re weak proof, but I can’t ignore them. My best guess is this. At 180 to 290 tokens per second, each token takes only 3 to 6 ms, and not all of that time is spent reading weights. Some of it is small setup work between steps, and that part runs at the core clock. A lower clock makes that part slower, even when memory speed stays the same. The 5070 Ti uses about 78% of its memory speed on the scoreboard, which leaves about a fifth of each token for this kind of extra work.

working it out for llama 3.1 8b#

For Llama 3.1 8B Q4_K_M, the same 78% gives about 142 tokens per second at the normal 300 W. The 5090 drew about 84% of its limit while writing, so I’m assuming the 5070 Ti draws around 250 W while writing. Prompt reading should use the full limit. Scaling the scoreboard numbers for the bigger model, I’d expect prompt reading at around 6,500 tokens per second.

limitwriting tok/sprompt tok/senergy per written tokenenergy per prompt token
300 W (normal)~142~6,500~1.77 J~46 mJ
270 W~142~6,270~1.77 J~43 mJ
240 W~139~6,040~1.72 J~40 mJ
210 W~133~5,820~1.57 J~36 mJ

The writing column is the uncertain one. If it holds, 270 W is free, 240 W costs about 2%, and 210 W costs about 6% but saves about 11% of the energy per token. Prompt reading gets slower, but it saves even more energy per token, because it was already using the full limit. If the 4090 and 5090 comments are closer to the truth, writing at 210 W could be 10% slower, not 6%.

Not every card can go down to 210 W. nvidia-smi -q -d POWER shows the lowest and highest limits for your card. On Windows, -pl needs an admin terminal, and the limit resets when you restart.

If you want to check this yourself, the normal test is too short. At 142 tokens per second, tg128 is less than a second of work. Logging power every 250 ms gives only three or four readings, and most of them are taken while the card is still speeding up. Longer runs work better:

nvidia-smi -pl 240
llama-bench -m Meta-Llama-3.1-8B-Instruct-Q4_K_M.gguf -ngl 99 -fa 1 \
  -p 4096 -n 1024 -r 5

And keep this running in a second terminal:

nvidia-smi --query-gpu=timestamp,power.draw,clocks.sm,clocks.mem,temperature.gpu \
  --format=csv,noheader -lms 250 > power-240.csv

Energy per token is the average power while writing, divided by tokens per second. The most important column is clocks.mem, because everything above assumes the memory clock doesn’t change. If it drops at 210 W, the table is wrong for a boring reason, and the setup-work idea never gets tested.

If writing really ignores the limit, the best setting is the lowest limit that costs nothing, because one person chatting with a model spends most of the time waiting for writing, not prompt reading. If writing drops like it did for that 4090, then those 7 ms per token have more clock-bound work in them than I thought, and the 64 users post needs a footnote.

TRIP COMPUTER / SESSION

TRIP A

Current drive.

A private counter for this browser session. Nothing here is transmitted or retained after the session ends.

Elapsed
00:00
Sections
0
Notes
0
Screens
0

Route/

Build plate11bded9
Chassis
v7.1.3
Revision
11bded9
Last serviced
07 Oct 2026

OWNER’S MANUAL / WTHRAJAT

OPERATING NOTES

How this thing moves.

The header behaves like a small mechanical system. Its readings respond to how you move through the site.

Throttle
Scrolling is input. Faster downward movement builds more momentum and engine speed.
Transmission
Upshifts follow sustained input. Scrolling upward slows and downshifts, and may briefly show reverse.
Idle
When input stops, RPM settles near 850 with mechanical drift. The gearbox eventually returns to neutral.
Tachometer
The needle always follows the reported RPM. It is never calculated from your position on the page.
Trip A
Session time, explored sections, opened notes and approximate screens travelled stay in this tab session.

Controls

Ctrl K
Search notes
?
Open this manual
Esc
Close an instrument
Tab
Move through controls