r/LocalLLaMA · · 1 min read

Deepseek training 2T and plans 8T model

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Quote: DeepSeek is training a 2T-parameter model and plans to eventually build an 8T-parameter model.

https://x.com/wallstengine/status/2101982843656388644

Current DeepSeek models:

  1. Flash parameter count of 552 billion
  2. Pro: 1.6T (trillion) total parameters with 49B (billion) activated weights per token

Mythos / Fable is estimated to be 10T parameter count.

submitted by /u/Terminator857
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA