[RELEASE] SupraBrain-50M-v0.1
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey there! So today we're releasing SupraBrain-50M, a hybrid language model that combines Gated DeltaNet linear recurrence with Sliding-Window Attention and Surprise-Gated update mechanisms to deliver very strong performance. Here are the benchmarks: Despite being trained on MUCH less data (5B vs 20B tokens!!), it's almost as good as Supra-Base-50M! 🔥 Some samples:
And:
Link to the HF model: https://huggingface.co/SupraLabs/SupraBrain-50M Give us a follow if you want to support us! BTW: Supra2-100M is releasing in the next few hours!! Stay tuned 🤗 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.