What GPUs will give me GOOD speeds and on DSV4 Flash and similar models, and not have to run a mega quantized version? Budget around $15k-ish.
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
I wish I could spend $15k on my own homelab hardware, but no this is for work lol.
Like the title says, we're looking to run DSV4 Flash (and similar tier models) locally at good speeds, both for token gen and prompt processing.
By "good" I'm thinking in the range of 40-50+ t/s gen and at least 1000 t/s prefill at moderate context.
We also don't want to run a version that's quantized to hell, so this will need at least 128 GB of VRAM.
It'll typically be 1 user at a time, but there may be times where 2 or 3 people are trying to use it at once and it would be nice if it isn't completely painful when that happens.
A couple options I'm considering right now:
3x AMD MI210 (192 GB)
3x NVidia A40 (144 GB)
Does anyone have performance numbers for these cards for DSV4 Flash, Qwen3.8-Flash-Next or similar models?
I tried to rent these in the cloud for some performance testing, but can't find any available right now.
NVidia preferred of course because CUDA, but def open to AMD if performance is similar. 192 GB is way nicer than 144 GB on those cards above.
Trying to keep this to 3 GPUs or less because that's what'll fit in our Dell R740 and then we don't have to build a special new host.
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.