r/LocalLLaMA · · 1 min read

Considering Buying Another RTX 3090 - Benefits?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Currently using dual RTX 3090s, and am happy with it. But never satisfied lol :)

I know I've basically maxed out my single stream TPS. (140+ on standard benchmarks now).

But I only have 48GB VRAM, So I can only do two concurrent requests @ 256k Context Length, anymore and my KV-Cache will cause OOM errors.

So I am considering adding a 3rd RTX 3090, and putting the 3rd one in Pipeline Parallel with the other two. That way I don't lose performance due to bandwidth bottlenecks, but get more room for concurrent requests.

I am likely to move to concurrent agents soon, but honestly have limited experience with them.

 INPUT │ ├─────────────────────────────────┐ │ │ ▼ ▼ ┌──────────┐ PCIe 4.0 8x ┌──────────┐ │ GPU 1 │ 16 GB/s │ GPU 2 │ │ (Node A) │◄════════════════════►│ (Node B) │ └────┬─────┘ └────┬─────┘ │ │ │ │ └─────────────────────────────────┘ │ PCIe 4.0 4x - 8GB/s ▼ ┌──────────┐ │ GPU 3 │ │ (Node C) │ └────┬─────┘ │ ▼ OUTPUT 

-> Has anyone else tried a similar setup?

-> What kind of results did you get in single stream / concurrent stream performance?

submitted by /u/Civil_Fee_7862
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA