A caveman qwen3.6 27B
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Just saw this on huggingface: https://huggingface.co/ProCreations/grug-27b
The benchmarks claim that it's quite a bit better than qwen3.6 27B original and that they reduced the amount of necessary tokens by more than 90%. It would make 27B running on my old laptop at 3tps feel more like 30tps for the thinking part, if true.
Couldn't test it yet.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.