[BIG DATASET RELEASE] - SupraLabs/reasoning-corpus-4K-5M-v1 - Train your tiny SLMs to think!
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| Hey r/LocalLLaMA ! We are back and we have something really amazing today. Our big 5M samples Reasoning Corpus dataset. This dataset features 5 million rows of: - repo_id --> where it's from All samples are within a 5k sequence length to make it fit perfectly for SFT/finetuning a tiny model. Link to the dataset on Hugging Face 🤗: Link to the SupraLabs Hugging Face org 🤗: Also, if you want to support our work, give us a follow on Hugging Face, share and review our work, and give us as much feedback as you want ❤️🔥🤗 Already more 250 people are trusting in us and our work! We hope, this dataset is useful for you all and we'd love to see your creations upon this. This dataset has already >1k downloads and over 80 likes - be the next one to use it 🔥🎉 [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.