Am I just hallucinating
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running the same prompts), it's all just vibes.
Some context, I'm running the latest build of llama-cpp, vulkan, 6900xt 16gb, 64giggles of system ram. I run gemma and qwens models; q8_0 for the moe's, q4_k_m for dense. KV at bf16. I lock a seed in to try to reduce the differences.
Any theoretical reason for the difference or am I just seeing ghosts.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.