r/LocalLLaMA · · 1 min read

What speeds are everyone getting with deepseek v4 flash 0731?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

What speeds are everyone getting with deepseek v4 flash 0731?

I’m getting~200 tps prompt processing / ~11 tps token gen, on 4x5060ti16gb with ddr4 3200 ram at 4-channel, via llamacpp, with context window of 128000, -ub/-b at 4096, “q8” unsloth’s lossless quant

submitted by /u/Ambitious_Fold_2874
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA