r/LocalLLaMA · · 1 min read

NVFP4 kv cache quantization on sm120 will make 32GB VRAM systems very capable

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

The best i can get from Qwen3.6-27B on my 32GB VRAM (2 x 5060) is ~60 tok/sec gen speed at context size 196608. (sakamakismile text nvfp4). Fp8 kv quantization. NVFP4 kv cache quantization can’t get here fast enough.

Reminds me of the time there was this game i couldn’t play on my first pc, because it needed 640KB of RAM minimum.

submitted by /u/Gray_wolf_2904
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA