r/LocalLLaMA · · 1 min read

20B Looping model (paper) matches or beats Qwen3 Coder 30B at 10% of pre-training tokens

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

No weights yet. I feel sad for them, that training run cost maybe 100s of thousands of dollars and they didn't even beat GPT-OSS 20B in every regard

But the ability to train a model from scratch on 3.5 trillion tokens instead of 35 trillion sure gives me hope. They only spent 100s of thousands of $ instead of millions, so maybe soon enough hobbyists will be able to pretrain true LLMs at home

submitted by /u/Dany0
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA