Would extremely high decode tok/s even be useful?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
If you were able to get an inference machine that could do decode at 1k toks/s or even 10k tok/s, would that even be helpful? Would it unlock any new use cases?
Let’s assume that this is for actually useful models and fairly large models like Qwen 3.5 397B, GLM-5.2, etc
Or at that speed would it just make better send to load much larger models? In which case, question still applies. E.g. Kimi K3 at the high speeds
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.