r/LocalLLaMA · · 1 min read

Best way to run Qwen3.8-27B on a system with a RTX 5090 + RTX 5070 Ti (32GB + 16GB)?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I have a system with 2 GPUs and 48GB VRAM total, a RTX 5090 + RTX 5070Ti.

What would you say is the best way to run Qwen3.8-27B on that system with the best quality and 262k context?

Would just the normal llama.cpp work with how it detects and does its own magic with dual CPU systems, or something else?

I think the RTX5090 has pcie4 x16 and the RTX5070Ti has pcie x8 if that matters.

submitted by /u/StartupTim
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA