r/LocalLLaMA · · 1 min read

Am I just hallucinating

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Or is there any reason why I feel like model output quality seems to be better when I use higher micro-batch values (ub) in llama-cpp? I don't really have any hard numbers or anything (just running the same prompts), it's all just vibes.

Some context, I'm running the latest build of llama-cpp, vulkan, 6900xt 16gb, 64giggles of system ram. I run gemma and qwens models; q8_0 for the moe's, q4_k_m for dense. KV at bf16. I lock a seed in to try to reduce the differences.

Any theoretical reason for the difference or am I just seeing ghosts.

submitted by /u/Xyklone
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA