Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model.
Hardware
| Item | Details |
|---|---|
| CPU | AMD Ryzen AI MAX+ 395 (16C/32T) |
| GPU | Radeon 8060S (integrated, RADV STRIX_HALO) |
| RAM | 128GB unified memory |
Software
| Item | Details |
|---|---|
| OS | Ubuntu 26.04.1 LTS / Kernel 7.0.0-30 |
| Vulkan | Mesa 26.0.8 / Vulkan API 1.4.335 |
Kernel Boot Parameters (Excerpt) amdgpu.gttsize=126976 amdgpu.noretry=0 ttm.pages_limit=28835840 ttm.page_pool_size=14417920 iommu=off
llama.cpp
Using the fork by Laurent Zuijdwijk:
git clone https://github.com/LaurentZuijdwijk/llama.cpp cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx Models
- Main:
Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64(AtomicChat) - MTP Draft:
Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf(dzannotti)
Launch Command
./build/bin/llama-server \ --host 0.0.0.0 --port 8080 \ --model ./models/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64-00001-of-00033.gguf \ -c 262144 --n-predict 32768 \ -t 2 --threads-batch 8 \ -ngl 999 --parallel 1 \ -b 8192 --ubatch-size 512 \ --load-mode mlock \ -fa on -cb \ -ctk f16 -ctv f16 \ --cache-reuse 1024 \ --jinja --reasoning on --reasoning-preserve \ --cache-prompt \ --chat-template-kwargs '{"reasoning_effort":"medium"}' \ -md ./models/Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf \ --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.75 Benchmark Results ($n=85$)
| Metric | Max | Min | Avg | Median |
|---|---|---|---|---|
| PP (tokens/s) | 305.45 | 19.90 | 138.61 | 131.66 |
| TG (tokens/s) | 46.76 | 17.11 | 26.67 | 26.69 |
Observations
When compared with Qwen3.8-27B, the quality of output for architectural and design tasks (such as OpenSpec proposals) appears noticeably superior. Depending on the instructions provided, it occasionally strays while attempting to "improve" the output, yet overall the performance is entirely satisfactory.
I generated the Japanese text using Claude(Opus4.6) and then translated it using Gemini(Flash 3.6).
[link] [comments]
More from r/LocalLLaMA
-
Alibaba plans AI model with 5 trillion to 10 trillion parameters, unveils new chip
Sep 22
-
Qwen 4 Announced at Apsara Conference
Sep 22
-
MiMo-V2.6 distilled themselves into Qwen 9B!
Sep 21
-
Uncensor an LLM without touching weights: inject a tiny trained KV-cache bank (~18MB) and unload it anytime
Sep 21
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.