r/LocalLLaMA · · 1 min read

2 x 5070ti Qwen 27B full config / stats

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

2 x 5070ti Qwen 27B full config / stats

Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV.

Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-nightly. Which has the KV cache connector fixes and performance improvements.

Highlights - 2 concurrent threads run comfortably without generation speed loss. Decode went up to 94-87tps 0-120k context, with prefill 4.6k-2.4k. You get 170k GPU KV and extra 246k with 8GB of RAM. Which makes it very comfortable for local agentic work.

A link to full compose file with a lot of additional info on memory usage etc.

Hopefully the upcoming small Qwen 3.8 will fit into this setup as well!

submitted by /u/val_in_tech
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA