Zagreus-0.4B-por a small open source language model for Portuguese
Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.
| mii-llm, an open source AI lab, released Zagreus-0.4B-por, a compact bilingual Portuguese–English language model pretrained entirely from scratch. The model has approximately 400 million parameters and is part of the Zagreus family, an ongoing experiment in building small open models focused on European languages. Training was sponsored by Seeweb cloud provider and Regolo.ai, which provided the computing infrastructure used for the project. The pretraining corpus was assembled from open datasets released through Hugging Face, including:
Evaluation resultsOn the Portuguese versions of ARC, HellaSwag and MMLU, the final checkpoint achieved an average score of 0.3113. Using Eduardo Garcia’s Portuguese evaluation harness, the 483k checkpoint scored:
These results are encouraging for a model of this size, but benchmark performance is not the main reason for the release. Zagreus-0.4B-por is an open base model intended to serve as a starting point for:
It is not presented as a finished assistant or production-ready product. It is an open foundation that others can inspect, fine-tune, modify and build upon. We would be interested in feedback, independent evaluations and experiments from the Portuguese NLP and open-source communities. Some useful links: the recipe: https://github.com/mii-llm/zagreus-nesso-slm [link] [comments] |
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.