Qwen3.6 27B on a 5090, 6.4k sample tok/s distribution after tuning MTP/cache settings
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Spent a while tuning llama.cpp for Qwen3.6 27B on a 9800X3D / 64GB / 5090 box and wanted to share the real distribution instead of just a headline number, since averages hide a lot. Ran with q8 KV cache, 192k context, MTP draft=10, spec-draft-p-min=0.5, batch/ubatch 512. Logged 6,454 samples across a mixed agentic coding + debugging + doc session over 20 hour ish. Peak bucket sits at 120-130 tok/s, mean 140.7, median 134.9, with a long tail up to 233. Worth noting the hybrid attention/SWA cache handling in llama.cpp still isn't perfect for this model if you see prompt reprocessing warnings in your logs that's why. Happy to share launch flags if anyone wants to compare setups. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.