If you had a 384GB (4x Blackwell), what model would you put on it and why?
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
Hey guys,
Company I work for is actually very interested in spending the money to host our own local model for the team.
We expect probably 2-3 super users and at the worst case 10-20 concurrent users.
The LLM would be mostly used for internal company policies/data management and other various "thinking tasks". No real coding will be done by such a machine probably other than me.
I would love the communities input on what you guys think is the best fit as until now the best machine I've had access to was a 64gb Mac Studio, this is way more compute than I've ever dealt with.
We are also speccing it to potentially expand to 8 blackwells, so I would also be curious how that model math changes.
I'm leaning towards deepseek v4 flash.
Edit: 4x RTX PRO 6000s
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.