r/LocalLLaMA · · 1 min read

Qwen3.8 flash next + exllamav3 + hermes is amazing

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I know there is nothing new with what I am saying but I recently started with hermes agent (it’s been a while I wanted to but did not have the time). Qwen3.8fn 6bpw exl3 (from turboderp) on a 6x3090 (I assume lower quants on lower number of gpus work same) gives me around 80-120t/s with good pp, and with good quality. That engine is crazy for cuda dude!

So now I control hermes from my phone securely (via Matrix) everyday and discover more and more its potential besides delegating to a coding agent/harness (opencode).

Qwen + Turboderp + Nous -> love on you

submitted by /u/takoulseum
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA