r/LocalLLaMA · · 1 min read

Would extremely high decode tok/s even be useful?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases?

Let’s assume that this is for actually useful models and fairly large models like Qwen 3.5 397B, GLM-5.2, etc

Or at that speed would it just make better send to load much larger models? In which case, question still applies. E.g. Kimi K3 at the high speeds

submitted by /u/LivingSwitch
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA