Gave a try to Exllamav3 and it's great!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Following this post I decided to try GLM 5.3 Flash on a 8x3090 setup and I can now run a Q4 with surprising speed; 700tk/s prefill & 42tk/s decoding! (lcp & vllm do not allow me to get that). Was afraid about quality but > 30m tokens with DSH and no issue (did not test vision yet, but looks supported). Just to say that I am really grateful to Turboderp and we should really support as much as possible others projects even if they do not comply with all our needs yet and not rely only on the big guys. [link] [comments] |
More from r/LocalLLaMA
-
Apple A20 Pro debuts with 7-core GPU, 32-core Neural Engine and 50% more memory bandwidth (~115 GB/s)
Sep 9
-
Surveillance plagiarism by OpenAI
Sep 9
-
Don't let FOMO win if you're interested in local llm from a hobby/learning aspect
Sep 9
-
Server rebuild to custom loop. 2x RTX Titans 24gb, 1x 22gb 2080ti | T: 70GB VRAM.
Sep 9
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.