Qwen3.8 flash next ISTA-DASLab GGUF 50t/s TG and 1500t/s PP with 12GB VRAM and 64GB RAM Laptop on 'Strata' engine
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| I think most people are sleeping on this inference engine. I tried multiple llama.cpp forks and none of them comes close to the inference speed of Strata. Initial version had some bugs with kv cache, cpu throttling and the developer fixed them. Inference engine (only runs on Nvidia for now; AMD support is experimental): https://github.com/Niko1221/Strata Here are some metrics with screenshots. My laptop has 5070ti 12GB VRAM, 64GB ddr5 RAM, Intel 275HX CPU, gen4 SSD. Aquarium test (unsloth studio connected via local API) The model I used was https://huggingface.co/ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF/tree/main/IQ3_XXS which has a good quality for its size. Above, the model generated the aquarium test. At 43k context depth, it was running at 51 t/s. Stock llama.cpp reached only 23t/s with the same quant. 32k context read at 1500t/s (unsloth studio via local API) This quant could only reach 100t/s PP with stock llama.cpp using the same quant. Strata was reading 32k context text at 1500t/s. This is way above my expectation. This quant can load with up to 200k context at 8bit. However, I was only using 131k context. As you can see it is utilizing 11GB VRAM and 56GB RAM (includes system/OS programs). This engine is specifically built for one model only and only select ggufs (ISTA-DASLab) work with it. You can use IQ3_S from ISTA-DASLab which they claim recovers full model's performance on coding benchmarks. I tested IQ3_XXS for some time and I would say it is an excellent model. I never thought 12GB VRAM would be enough to run frontier models from 6 months ago locally on a laptop. What a time to be alive! [link] [comments] |
More from r/LocalLLaMA
-
PSA: ModelScope CLI is now moved to "modelscope-hub"
Sep 30
-
Qwen-family LLMs are quietly becoming the backbone of modern audio models; One chart for the architectures of 100+ audio models
Sep 29
-
Are you worried about a potential ban of Chinese open weight models?
Sep 29
-
AMD boosting AI/LLM performance for Radeon iGPUs as much as 18~23% with Linux 7.4
Sep 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.