Gemma 4 QAT handles KV cache quantization MUCH better, KLD benchmarks show
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Link to the article: KV Cache Quantization on Gemma 4 31B: Non-QAT vs QAT KLD benchmarks with BeeLlama.cpp v0.4.3, fork of llama.cpp with more KV cache quantization options, comparing Gemma Q4_0 non-QAT vs Gemma Q4_0 QAT. Long story short: QAT is much more friendly to KV cache quantization, moving same-top agreement from "different model" to "that looks like Gemma 4?" This confirms results from previous posts on this subreddit:
Comparison of standard quants Full benchmark results, setup, method, analysis, explanations and everything else can be found in the article.
[link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.