r/LocalLLaMA · · 1 min read

Qwen3.8 Flash on 12GB VRAM - 15 tokens/s

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Achieved steady 15 tokens/second output and 100-120 promp processing per second with 12GB GPU (RTX 5070 SFF) with Qwen3.8-Flash-Next-GSQ-RCO-GGUF at 3 bpw (IQ3_XXS which maches BF16 on AIME25). Around 76GB full gguf - while only 47GB needs to be sharded (loaded into VRAM+RAM) rest is done in SSD.

Output quality is really good - under my tests it matches Unsloth’s Q6-Q8 level.

Hardware:

64GB DDR5 (5600)
12GB - RTX 5070 SFF Gigabyte
1TB NVME
Ryzen 5 7600 CPU

Outputs and results:

| Server startup | 4.9s |
| Cold launch → 1K prompt | 31.7s** |
| **1K generation
| 13.8 tokens/s |
| 8K generation | 14.6 tokens/s |
| 20K prompt processing | 104.7 tokens/s |
| 20K wait before output | 3,11min |
| 20K generation | 14.0 tokens/s |
| 90–100K generation | 11.3 tokens/s |

128K max context length in my testing.

Use FreeToken CLI with llama.

Testing 2.40bpw version soon which should be equivalent to NVFP4 - Q4 quants and keep you posted.

If anyone runs similar things let me know your experiences.

submitted by /u/KnownAd4832
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA