r/LocalLLaMA · · 1 min read

Those who use many layers in CPU/RAM and some in GPU - what are your specs and speeds?

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

I am trying to figure out if it's worth upgrading my RAM, but I've noticed that some MoE models don't seem to do well with many layers shared from VRAM --> CPU/RAM. This may be something on my end; a software config or perhaps my specific hardware config.

This made me curious as to how many are doing this. I'm thinking many are, especially with non-dense models, but even better if you do this with dense models; I'd like to know the results you get!

Example: You have 16gb of VRAM but you have 128gb of system DRAM (not unified - that's a separate discussion). You run a large MoE model and load some layers in GPU and the rest in CPU/RAM.

  1. Which model are you running? Include the name and quantization and

  2. What's your hw config? Just basics, like CPU type, RAM type and amount, GPU type, etc.

  3. What are your pre-fill / prompt processing and token generation speeds?

  4. How much context are you setting with KV quant type and which inference software?

submitted by /u/Jorlen
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA