r/LocalLLaMA · · 2 min read

How to estimate tokens/sec for your hardware

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

How to estimate tokens/sec for your hardware

We all want more tokens per second but I keep seeing confusion on what to expect for given hardware.

For the decoding phase (TG/s) to produce one token all the model weights and KV cache needs to be read from VRAM. The compute isn't the bottleneck, only memory bandwidth.

This means we can estimate the maximum TG/s we can ever achieve given the model weights and memory bandwidth.

If we ignore the KV cache for now, the formula is:

 VRAM GB/s TG/s = ------------------ model weights GB 

The math is more complicated for mixture of expert (MoE) models, but easy for dense models.

For Qwen3.8 27B Q4_K_XL, we have model weights of 16.8 GB (we exclude things not read every token; MTP layer and input embedding table)

For AMD Radeon AI PRO R9700, we have a memory bandwidth of 637 GB/s.

Therefore the theorical maximum for this model & hardware is:

637 / 16.8 = 38 TG/s 

In the real-world it only goes down from here due to inefficiencies in the software/hardware stack. On my system running that model and hardware with llama.cpp, I get 29 TG/s, so 29 / 38 = 76% of ideal.

Also as the KV cache grows, those bytes are read for every token. Continuing the example with Qwen3.8 27B, the KV cache BF16 it costs 64 KB per token read.

The full formula becomes:

 VRAM GB/s TG/s = ------------------------------------------------------------ model weights GB + KV cache GB/token * context size tokens 

We can make that formula more useful by moving VRAM GB/s over to the left. This allows us to plot TG/s per VRAM GS/s vs context size for a particular model.

Continuing our example:

https://preview.redd.it/9r2lw1jy9knh1.png?width=1508&format=png&auto=webp&s=6fd6ba8991c3e091ec7261e72a527478a6b89d24

This allows you to plug in your own VRAM GB/s.

For a 5090 with 1.8 TB/s memory bandwidth

1,800 * 0.0590 = 106 TG/s maximum 1,800 * 0.0293 = 53 TG/s maximum at 256k context window 

Caveats

  • Assumes entire model and context is in VRAM
  • Simplified formula is only for dense models
  • Speculative decoding is added on these base numbers
  • These are theorical maximums. Real-world numbers are lower due to inefficiencies in the software/hardware

AI was used to draw the plot. Everything else is written by me.

submitted by /u/Pyrolistical
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA