Looking for Dual GPU Tips and tricks.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Just added a second 5060 16gb to my server for 32gb total VRAM + 80GB of ECC DDR4 It's pcie 3.0 so I think tensor parallel is not going to run well in any config but I get around 3200 tok/s prompt processing and 100 tok/s generation with qwen 3.6 35B. 27b runs at about 600 PP / 23 generation in tensor split. What settings should I be looking at to optimize my server now that I'm splitting weights across two cards? Currently running ggufs with llama.cpp but I have vllm nvfp4 models too. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.