r/LocalLLaMA · · 1 min read

Qwen3.6 35B (2 min) vs Muse Glimmer 30B (4 min) on custom Llama.cpp build (RTX 5080)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Qwen3.6 35B (2 min) vs Muse Glimmer 30B (4 min) on custom Llama.cpp build (RTX 5080)

Muse Glimmer 30B feels significantly more precise and reliable, it almost never drops the ball or breaks rules. However, its designs lack creative depth and richness.

Qwen3.6 35B, on the other hand, is prone to more occasional blunders/hallucinations, but its creative output is superior. It generates far richer, more complex voxel worlds and offers higher design quality.

LLama.ccp Build Provenance:

  • Base: llama.cpp upstream (merge 4445f8d, build 661)
  • CUDA Toolkit 13.1 + MSVC 19.44 + sm_120a-real (native Blackwell PTX)
  • Flags: GGML_CUDA=ON, GGML_CUDA_FA=ON, GGML_CUDA_FA_ALL_QUANTS=ON, GGML_CUDA_GRAPHS=ON, GGML_NATIVE=OFF
  • License: MIT (upstream llama.cpp)

Do you think Qwen3.6 is still the undisputed king here?

submitted by /u/myanimal22
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA