r/LocalLLaMA · · 1 min read

Kimi K3 full model running on 16x GB10 cluster at 20+tps

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Kimi K3 full model running on 16x GB10 cluster at 20+tps

Kimi K3 full model running on 16x GB10 cluster at 20+tps average (llama-benchy coherent corpus) 38tps peak, 750tps prefill. This is the first run of full k3 with dspark on my cluster. I will be doing some tests and try tp speed this up. As soon as it looks ready I'll publish the vllm image and instructions.
https://forums.developer.nvidia.com/t/full-kimi-k3-running-on-16x-gb10-cluster/379174

submitted by /u/ciprianveg
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA