r/LocalLLaMA · · 1 min read

Built and released BetterGPT-150M – A compact 150M parameter completion model (+ live HF Space demo)

Mirrored from r/LocalLLaMA for archival readability. Support the source by reading on the original site.

Hey everyone,

​I recently finished pre-training BetterGPT-150M, a small, lightweight causal language model with ~152 million parameters.Trained on 15B tokens.

Dataset & Training: Trained across stable and annealing phases using curated datasets (including FineWeb-Edu, fine maths, cosmopedia, starcode-python), ensuring strong capability retention while maximizing token efficiency.

​Performance: Benchmark evaluations show it outperforms GPT-2 Small while remaining on par with models trained on significantly larger token budgets.

​Since many small/tiny models tend to get buried under massive LLM releases, I wanted to share it here for anyone interested in lightweight architectures, fast CPU inference, or small edge-device experimentation.

Repo: https://github.com/harikrish2727/BetterGPT

​Model Hub: https://huggingface.co/Harikrish2727/BetterGPT-150M

​Live Demo: https://huggingface.co/spaces/Harikrish2727/BetterGPT-Demo (Hosted on ZeroGPU with token streaming)

​Model Notes:

​Task: Text completion / generation (it is a standard base completion model, not instruction-tuned).

​Footprint: Very low RAM/vRAM footprint, runs instantly on standard CPUs.

​Feel free to try out prompts on the Space demo or pull the weights to run locally. Feedback, benchmark suggestions, or ideas for instruction fine-tuning are always welcome!

This is my very first serious try, hope I get genuine feedback from you guys.

submitted by /u/Rich-Title-3668
[link] [comments]

Discussion (0)

Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.

Sign in →

No comments yet. Sign in and be the first to say something.

More from r/LocalLLaMA