r/LocalLLaMA · · 2 min read

Qwen3.8-27B: >70 tok/s (>160 tok/s concurrent), 10k tok/s prefill, full context on 2x3090 (or and 48GB or larger on ampere or higher), vanilla vllm

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I didn't know my set up was outperforming nearly everyone until reading another discussion where people were struggling getting half of that speed with half the context on the same hardware. I benchmarked a couple dozen quants, vllm, sglang, llama.cp and benchmarked settings and configurations on each to find the following:

I can get full 262k context (can nearly fit 2 full contexts in kv cache simultaneously), >70 tok/sec single stream decode, close to 200 tok/sec decode at 3 to 4 concurrency (saturates there), over 10k tok/sec prefill (when kv cache is light).

Stock vLLM v0.29 on Linux serving https://huggingface.co/RedHatAI/Qwen3.8-27B-INT4

It is int4 which is ideal for ampere. It has weighed fp8 kv cache and externally benchmarked with near identical performance to bf16 weights.

vllm serve RedHatAI/Qwen3.8-27B-INT4 --max-model-len auto --enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 --enable-prefix-caching --kv-cache-dtype fp8 --trust-remote-code --max-num-seqs 8 --gpu_memory_utilization 0.9 --max-num-batched-tokens 8192 --mm-encoder-tp-mode data --speculative-config '{"method":"mtp","num_speculative_tokens":3}' --tensor-parallel-size 2

It's pretty much the standard vllm recipe for qwen3.8-27b which I didn't think was special. I measured performance using the metrics API endpoint from vllm and Prometheus+perses on real workflows over the past month. I also used vllm bench serve while tuning the settings. It's nothing special, and a pretty standard approach, which is why I haven't shared it before! I don't know why the redhat quant is overlooked. It's pretty much the ideal quant for ampere... And it has weighted kv cache which very few have.

Edit: answering questions from comments

I've tried several hardware boxes (rented) as well with similar results. Two example ones I'm currently renting:

CPU: AMD EPYC 7532 32-Core Processor Effective Cores: 64 CPU RAM 64.3 GB Disk: KINGSTON SNV3S4000G Disk Bandwidth: 1918.0 MB/s Motherboard: MZ32-AR0-00 PCIE Bandwidth: 12.8 GB/s PCIE: 4.0/8x

CPU: AMD Ryzen Threadripper PRO 3945WX 12-Cores Effective Cores: 24 CPU RAM 128.8 GB Disk: CT2000P510SSD5 Disk Bandwidth: 6536.0 MB/s Motherboard: MC62-G40-00 PCIE Bandwidth: 12.7 GB/s PCIE: 3.0/16x

submitted by /u/maqifrnswa
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA