KAT Coder 2.5 dev: Do yourself a favor and try it!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
It is so good! I don't know why there aren't more people talking about it. Fewer tokens, faster and more accurate than Qwen 3.6 35b a3b. On my setup it's nearly as good as 27b, but 5x faster. And it completely trashes the Gemma 4 models.
At least for my use case, it feels amazing. I'd love to hear other people's experience with it. If you want an actual measure of performance, I have a GitHub repo explaining how I tested it for my type of use case with a detailed performance comparison with other models . It has the quants I used, OpenCode and llama.cpp version along with all the flags for temp, top-p, top-k etc.; and if there's some detail missing please let me know. But really I think you should just download the model and try it out yourself, because we all have different use cases and those will always be more informative than benchmarks or one person's idiosyncratic experience.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.