If you have a 3090, or other 30xx for local LLMs, I have something for you
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I have a custom fork of llama.cpp designed around the ampere architecture specifically (though many of the upgrades also translate to faster performance of blackwell + lovelace). The recommended config supports 90+ TPS (for agentic/coding, at temp 1; greedy will of course be faster) through 100K tokens, with context of up to 240K.
If you want the repo, it is here:
https://github.com/JakeATX/llamAmpere
I recommend running with this quant, which is ~ 4 K M quality but considerably faster (technically, a 3 K XL upgrade)
https://huggingface.co/jakeatx/Qwen3.8-27B-ATX-IQ4_XS-M-GGUF
If you want the deep dive on how it is so much faster (80% vs the near comp at 200K!), at more context, there is a long form article here.
https://x.com/JakeKAllDay/status/2095646450138874095?s=20
Running faster than API speeds on my 3090 (for 27b at least) has genuinely been a step change in the utility of the model + card for me. I hope you enjoy it!
[link] [comments]
More from r/LocalLLaMA
-
NVIDIA shipped OpenShell, an open source sandbox that gives local and open agents real runtime limits instead of prompt rules. Over 100 firms joined the safety stack. OpenAI did not.
Sep 28
-
3090 for $1500???
Sep 28
-
modified qwen 3.8 27b modifies windows credential dumper to bypass EDR detection
Sep 28
-
Minisforum MS-S1 MAX-P495 @ €7.799,00
Sep 28
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.