Quantizing Kimi K3 (2.8T A50B) to GGUF ourselves - Q3_K_S works, 1.1 TB on disk
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| we're experimenting with our own dynamic GGUF quants of kimi k3, made from the original weights with our llama.cpp fork. Q3_K_S is done and works 1114.76 GiB on disk. Q1 and Q2 are in progress, results on those tomorrow rented box hardware: - AMD EPYC 9554P, 64 cores the run: - 110 threads we ran a short test for text coherence and image understanding to make sure the quant isn't lobotomized. loaded the 1969 NYT "men walk on moon" front page and asked the model to describe what's going on. it listed the masthead, the "all the news that's fit to print" slogan, the date, the 10 cent price, the headline, the sub-headline about astronauts collecting rocks and the "voice from moon" column. we haven't noticed any hallucinated text wdyt about running quants of giant models like this on cpu instead of going with smth smaller but with normal tps and zero extra costs? disclaimer: we're the team behind atomic chat ( atomic.chat ) [link] [comments] |
More from r/LocalLLaMA
-
Thomson Reuters releases Thomson-1.0-Small. A law and tax focused model
Aug 26
-
Fully quantized NVFP4 Qwen3.8-27B with QUASAR QAD
Aug 26
-
M5 Ultra 96GB vs M5 Max 128GB — is 2x bandwidth worth losing 32GB of RAM, with Qwen3.8-Flash-Next dropping tomorrow?
Aug 25
-
Qwen 3.8 27b has ThreeJs locked down.
Aug 25
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.