20B Looping model (paper) matches or beats Qwen3 Coder 30B at 10% of pre-training tokens
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
No weights yet. I feel sad for them, that training run cost maybe 100s of thousands of dollars and they didn't even beat GPT-OSS 20B in every regard
But the ability to train a model from scratch on 3.5 trillion tokens instead of 35 trillion sure gives me hope. They only spent 100s of thousands of $ instead of millions, so maybe soon enough hobbyists will be able to pretrain true LLMs at home
[link] [comments]
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.