r/LocalLLaMA · · 1 min read

Splash 1.1.0 released, GGUF quants support, MLX import and more

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

On my M5 Pro 64GB I can comfortably work in an agentic setup with the Qwen3.8 27B model in good quality (Unsloth UD-Q4_K_XL) at a decent speed of 50 t/s.

Splash combines optimized kernels, excellent speculative decoding, a well-implemented prefix cache, and mixed-weight support in a single program. To me, this is a breakthrough in local inference on Apple Silicon.

https://github.com/incoai/splash/releases/tag/1.1.0

submitted by /u/wojtek15
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA