r/LocalLLaMA · · 2 min read

Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Basalt: Flash-Next at 665 tok/s structured, 354 prose on a 5090 + 5060 Ti (2.6x Strata)

Basalt: Blackwell inference engine

Hi all! So over the last week I've been working on a Strata fork that's heavily tuned and can achieve throughput up to 2.6x what Strata usually does on the same weights. It's designed as a specialized engine that only supports Qwen3.8 Flash-Next and Blackwell architecture, including dual GPUs like my current hardware (5090 + 5060 Ti - 9950X 32GB DDR5 RAM). Basalt features include:

- SPEED: 665 struct, 354 prose, 7,317 prefill at 64k, IQ3_XXS, 400 W, speed is the highlight
- Single 5090 works too (no second card): 585 struct, 316 prose on IQ3_XXS, ~12% slower than with the 5060 Ti, prefill unchanged
- Real concurrency for up to 8 users: shared KV or per slot, MTP enabled - 623 tok/s total at 8 streams (I don't have Strata's numbers to compare)
- Fine-tuned MTP for lower quants, for increased throughput
- Custom vision encoder designed from scratch, up to 3x faster than llama.cpp's on the GPU, 4x on the CPU
- A simple server UI that shows current throughput (including concurrency stats), expert distribution and hardware statistics. No chat, BYOH (bring your own harness)
- OpenAI + Anthropic compatible
- Linux support (No Windows or Mac)

Basalt uses a similar format to NInfer, where weights are re-packed (not re-quantized) into a .basalt file, including all the metadata, vision and MTP, so you only have to keep a single file per quant. Initial support includes ISTA-DASLab's GSQ-RCO for Q2, IQ3_XXS and IQ3_S and UD-Q4_K_XL and Q8 from Unsloth, so you can pick the weights depending on your VRAM/RAM budget and quant preference.

Quick Q&A:

+ Is it open source?
- Yes, fully open source, MIT license: https://github.com/jesdga95/basalt, fork it, improve it, share it with friends and foes.

+ Where are the weights?
- Here: https://huggingface.co/jesdga/Qwen3.8-Flash-Next-Basalt pick your poison, fast and dumb or smart and slow. IQ3_S is a good middle ground (~89% top 1 agreement, 300 tok/s prose on my setup).

+ Why didn't you just contribute upstream to Strata?
- This is not a single feature that can be easily merged into Strata, it basically rewrites most of the decode and part of the prefill kernels and strips support for non Blackwell cards including AMD, Intel and older Nvidia generations. I have however contributed patches to Strata and llama.cpp and any critical findings will be pushed upstream.

+ Will you support my AMD 98123X?
- Sure, send one my way. For now I can only support what I can personally test and I intend to keep it that way for the time being.

+ Why not 400 tok/s?
- I'm still trying!

+ This is vibe coded slop
- Yes, but it's fast slop. Nobody is hand-writing cuda kernels anymore.

submitted by /u/jesdga95
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA