How to estimate tokens/sec for your hardware
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| We all want more tokens per second but I keep seeing confusion on what to expect for given hardware. For the decoding phase (TG/s) to produce one token all the model weights and KV cache needs to be read from VRAM. The compute isn't the bottleneck, only memory bandwidth. This means we can estimate the maximum TG/s we can ever achieve given the model weights and memory bandwidth. If we ignore the KV cache for now, the formula is: The math is more complicated for mixture of expert (MoE) models, but easy for dense models. For Qwen3.8 27B Q4_K_XL, we have model weights of 16.8 GB (we exclude things not read every token; MTP layer and input embedding table) For AMD Radeon AI PRO R9700, we have a memory bandwidth of 637 GB/s. Therefore the theorical maximum for this model & hardware is: In the real-world it only goes down from here due to inefficiencies in the software/hardware stack. On my system running that model and hardware with llama.cpp, I get 29 TG/s, so Also as the KV cache grows, those bytes are read for every token. Continuing the example with Qwen3.8 27B, the KV cache BF16 it costs 64 KB per token read. The full formula becomes: We can make that formula more useful by moving Continuing our example: This allows you to plug in your own VRAM GB/s. For a 5090 with 1.8 TB/s memory bandwidth Caveats
AI was used to draw the plot. Everything else is written by me. [link] [comments] |
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.