Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and configurations on each to find the following:
I can get full 262k context (can nearly fit 2 full contexts in kv cache simultaneously), >70 tok/sec single stream decode, close to 200 tok/sec decode at 3 to 4 concurrency (saturates there), over 10k tok/sec prefill (when kv cache is light).
Stock vLLM v0.29 on Linux serving https://huggingface.co/RedHatAI/Qwen3.8-27B-INT4
It is int4 which is ideal for ampere. It has weighed fp8 kv cache and externally benchmarked with near identical performance to bf16 weights.
vllm serve RedHatAI/Qwen3.8-27B-INT4 --max-model-len auto --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 --enable-prefix-caching --kv-cache-dtype fp8 --trust-remote-code --max-num-seqs 8 --gpu_memory_utilization 0.9 --max-num-batched-tokens 8192 --mm-encoder-tp-mode data --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --tensor-parallel-size 2
It's pretty much the standard vllm recipe for qwen3.8-27b which I didn't think was special. I measured performance using the metrics API endpoint from vllm and Prometheus+perses on real workflows over the past month. I also used vllm bench serve while tuning the settings. It's nothing special, and a pretty standard approach, which is why I haven't shared it before! I don't know why the redhat quant is overlooked. It's pretty much the ideal quant for ampere... And it has weighted kv cache which very few have.
Edit: answering questions from comments
I've tried several hardware boxes (rented) as well with similar results. Two example ones I'm currently renting:
CPU: AMD EPYC 7532 32-Core Processor Effective Cores: 64 CPU RAM 64.3 GB Disk: KINGSTON SNV3S4000G Disk Bandwidth: 1918.0 MB/s Motherboard: MZ32-AR0-00 PCIE Bandwidth: 12.8 GB/s PCIE: 4.0/8x
CPU: AMD Ryzen Threadripper PRO 3945WX 12-Cores Effective Cores: 24 CPU RAM 128.8 GB Disk: CT2000P510SSD5 Disk Bandwidth: 6536.0 MB/s Motherboard: MC62-G40-00 PCIE Bandwidth: 12.7 GB/s PCIE: 3.0/16x
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.