Deepseek-V4-Flash-0731 Dwarfstar on Mac
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Here is the prefill performance in an M2 Ultra with 192GB of RAM. For decode, at the following depth: 45k: 23.5 t/s 192k: 18 t/s That speed is maintained with 8k token output at those depths. [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.